Text generation method and device, model training method and device and electronic equipment
By inserting the low-rank decomposition matrix in the target layer of the large language model, the first largest language model is built and the underlying calculation results are shared, the hallucination problem is solved and the inference overhead is reduced, and the model training and inference efficiency is improved.
Patent Information
- Application Number
- CN202410208439.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-02-23
- Publication Date
- 2025-08-26
AI Technical Summary
The existing large language models have hallucinations problems in practical applications, and the inference overhead is high during comparison and decoding, which affects the efficiency of the model.
By inserting the low-rank decomposition matrix into the target layer of the pre-trained second language model, the first largest language model is built, and only the low-rank decomposition matrix of the target layer is adjusted, the underlying calculation results are shared, and the comparison and decoding is performed to reduce the inference overhead.
It effectively reduces the amount of parameters for training the first largest language model, improves training efficiency, and reduces inference overhead during comparison and decoding, improves the model's inference efficiency and the reliability of generating text.
Smart Images

Figure CN120541513A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a text generation method, a model training method, a device, and an electronic device. Background Art
[0002] While existing large language models have achieved impressive results across numerous tasks, in practice they can still generate content that contradicts the input, contradicts itself, or is inconsistent with the facts. This phenomenon, commonly known as "hallucination," can have serious consequences in real-world scenarios. Contrastive decoding algorithms can be used to mitigate the hallucination problem of large language models. However, due to the large number of parameters in large language models, contrastive decoding often introduces additional inference overhead, reducing the model's inference efficiency. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail in this application. This overview is not intended to limit the scope of protection of the claims.
[0004] The embodiments of the present application provide a text generation method, a model training method, an apparatus, and an electronic device, which can reduce the inference overhead and improve the inference efficiency of the model during comparative decoding.
[0005] In one aspect, an embodiment of the present application provides a text generation method, comprising:
[0006] Obtaining a sample text containing factual errors, and training a first language model based on the sample text, wherein the first language model is obtained by inserting a low-rank decomposition matrix to be adjusted into a target layer of a pre-trained second language model, the second language model is provided with M processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is the target layer, and the low-rank decomposition matrix is used to be summed with the original parameter matrix and then mapped to the input of the target layer to obtain the output of the target layer, where N ≥ M / 2, and M and N are both positive integers;
[0007] Obtaining a target instruction, calling the trained first language model and the trained second language model to map the target instruction respectively, and obtaining a first target prediction distribution output by the first language model and a second target prediction distribution output by the second language model;
[0008] The first target prediction distribution and the second target prediction distribution are compared and decoded to obtain a target prediction text.
[0009] On the other hand, an embodiment of the present application provides a model training method, comprising:
[0010] Obtaining sample text with factual errors, and training a first language model based on the sample text;
[0011] Among them, the first large language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second large language model, the second large language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is the target layer. The low-rank decomposition matrix is used to map the input of the target layer after summing with the original parameter matrix to obtain the output of the target layer, N≥M / 2, M and N are both positive integers, and the output of the trained first large language model is used for comparison and decoding with the output of the second large language model.
[0012] On the other hand, an embodiment of the present application further provides a text generation device, comprising:
[0013] A first training module is used to obtain sample text containing factual errors and train a first large language model based on the sample text, wherein the first large language model is obtained by inserting a low-rank decomposition matrix to be adjusted into a target layer of a pre-trained second large language model, the second large language model is provided with M layers of processing layers cascaded in sequence, each of the processing layers is configured with an original parameter matrix, and each processing layer after the Nth layer is the target layer, and the low-rank decomposition matrix is used to be summed with the original parameter matrix and then mapped to the input of the target layer to obtain the output of the target layer, where N ≥ M / 2, and M and N are both positive integers;
[0014] a prediction module, configured to obtain a target instruction, call the trained first language model and the trained second language model to map the target instruction respectively, and obtain a first target prediction distribution output by the first language model and a second target prediction distribution output by the second language model;
[0015] A decoding module is used to compare and decode the first target prediction distribution and the second target prediction distribution to obtain a target prediction text.
[0016] Furthermore, in each target layer, the number of the low-rank decomposition matrices is multiple, and the first training module is specifically used for:
[0017] Obtaining a sample instruction, inputting the sample instruction into a first large language model, embedding the sample instruction to obtain a sample embedding, determining at least one target decomposition matrix from a plurality of the low-rank decomposition matrices in each target layer based on the sample embedding, and mapping the input of the target layer based on the target decomposition matrix to obtain an output of the target layer;
[0018] Determine a sample prediction distribution according to an output of the target layer of the last layer, determine a target loss based on the sample prediction distribution and the sample text, and train the first language model based on the target loss.
[0019] Furthermore, the first language model is provided with a first weight determination network, and the first training module is specifically used for:
[0020] Calling the first weight determination network to map the sample embedding to obtain a sample weight distribution, wherein the sample weight distribution includes weights corresponding to each of the low-rank decomposition matrices;
[0021] At least one target decomposition matrix is determined from the plurality of low-rank decomposition matrices according to the sample weight distribution.
[0022] Furthermore, the plurality of target layers share one first weight determination network, and the first training module is specifically used for:
[0023] Determining a target matrix combination from a plurality of decomposition matrix combinations according to the sample weight distribution, wherein the decomposition matrix combination includes one of the low-rank decomposition matrices of the target layer in each layer;
[0024] Determine the target decomposition matrix corresponding to the current target layer from the multiple low-rank decomposition matrices of the target matrix combination.
[0025] Furthermore, the first training module is specifically used to:
[0026] Determining a cross entropy loss based on the sample prediction distribution and the sample text;
[0027] Determining an entropy value of the sample weight distribution, and determining a uniformity loss according to the entropy value of the sample weight distribution;
[0028] The cross entropy loss and the uniformity loss are weighted to obtain a target loss.
[0029] Furthermore, the first language model is further provided with a second weight determination network to be adjusted, and the first prediction module is specifically configured to:
[0030] Calling the second weight determination network to map the sample weight distribution to obtain a first loss weight corresponding to the cross entropy loss and a second loss weight corresponding to the uniformity loss;
[0031] The cross entropy loss and the uniformity loss are weighted based on the first loss weight and the second loss weight to obtain a target loss.
[0032] Furthermore, the number of the target decomposition matrices is multiple, and the first training module is specifically used for:
[0033] Mapping the input of the target layer based on each of the target decomposition matrices to obtain sub-outputs corresponding to each of the target decomposition matrices;
[0034] The weights corresponding to the target decomposition matrices are determined based on the sample weight distribution, and the multiple sub-outputs are weighted based on the weights corresponding to the target decomposition matrices to obtain the output of the target layer.
[0035] Furthermore, the above prediction module is specifically used for:
[0036] calling the trained first language model and the second language model to respectively map the target instruction, wherein the output of the target layer of each layer except the last layer in the second language model is obtained by comparing and adjusting the output of the target layer corresponding to the first language model in the output space;
[0037] Determine a first target prediction distribution according to an output of the target layer of the last layer of the first language model;
[0038] Determine a second target prediction distribution according to an output of the target layer of the last layer of the second largest language model.
[0039] Furthermore, the first training module is specifically used to:
[0040] Obtaining a first reference text that is factually correct and a second reference text that is factually incorrect, wherein the second reference text is rewritten based on the first reference text;
[0041] A rewriting instruction is constructed based on the first reference text and the second reference text, and the rewriting instruction is input into the second language model or the pre-trained third language model for mapping to obtain a sample text with factual errors.
[0042] On the other hand, an embodiment of the present application further provides a model training device, comprising:
[0043] A second training module is configured to obtain sample texts with factual errors and train the first language model based on the sample texts;
[0044] Among them, the first large language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second large language model, the second large language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is the target layer. The low-rank decomposition matrix is used to map the input of the target layer after summing with the original parameter matrix to obtain the output of the target layer, N≥M / 2, M and N are both positive integers, and the output of the trained first large language model is used for comparison and decoding with the output of the second large language model.
[0045] On the other hand, an embodiment of the present application further provides an electronic device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned text generation method or the above-mentioned model training method when executing the computer program.
[0046] On the other hand, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the above-mentioned text generation method, or to implement the above-mentioned model training method.
[0047] In another aspect, embodiments of the present application further provide a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to implement the aforementioned text generation method or the aforementioned model training method.
[0048] The embodiments of the present application include at least the following beneficial effects: by obtaining sample text with factual errors, the first language model is trained based on the sample text. Since the first language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second language model, the low-rank decomposition matrix can greatly reduce the amount of parameters adjusted when training the first language model, thereby improving the training efficiency of the first language model. In addition, since the second language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer, N≥M / 2, that is, the present application only inserts the low-rank decomposition matrix at the high level of the first language model, so that the first language model and the second language model can share the underlying calculation results. Subsequently, when the first target prediction distribution and the second target prediction distribution are compared and decoded, the inference overhead can be effectively reduced, and the inference efficiency of the model can be improved.
[0049] Other features and advantages of the present application will be set forth in the following description, and in part will be apparent from the description, or may be understood by practicing the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings are used to provide a further understanding of the technical solution of the present application and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the technical solution of the present application and do not constitute a limitation on the technical solution of the present application.
[0051] Figure 1 A schematic diagram of an optional implementation environment provided for an embodiment of the present application;
[0052] Figure 2 An optional flowchart of the text generation method provided in the embodiment of the present application;
[0053] Figure 3 A schematic diagram of an optional structure of a low-rank decomposition matrix provided in an embodiment of the present application;
[0054] Figure 4 A schematic diagram of an optional training architecture for the first language model provided in an embodiment of the present application;
[0055] Figure 5 An optional structural diagram of determining the output of the target layer provided in an embodiment of the present application
[0056] Figure 6 A schematic diagram of another optional structure for determining the output of the target layer provided in an embodiment of the present application;
[0057] Figure 7 A schematic diagram of an optional architecture of the text generation method provided in an embodiment of the present application;
[0058] Figure 8 A schematic diagram of an optional structure of the text generation device provided in an embodiment of the present application;
[0059] Figure 9 A schematic diagram of an optional structure of the model training device provided in an embodiment of the present application;
[0060] Figure 10 A partial structural block diagram of a terminal provided in an embodiment of the present application;
[0061] Figure 11 A partial structural block diagram of the server provided in an embodiment of the present application. DETAILED DESCRIPTION
[0062] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0063] It should be noted that, in each specific embodiment of the present application, when it comes to the need to perform relevant processing based on data related to the characteristics of the target object such as target object attribute information or attribute information set, the permission or consent of the target object will be obtained first, and the collection, use and processing of these data will comply with relevant laws, regulations and standards. Among them, the target object can be a user. In addition, when the embodiment of the present application needs to obtain target object attribute information, the target object's separate permission or separate consent will be obtained by means of a pop-up window or jumping to a confirmation page. After clearly obtaining the target object's separate permission or separate consent, the necessary target object-related data for enabling the normal operation of the embodiment of the present application will be obtained.
[0064] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal, and can be implemented in whole or in part by using software, hardware (such as processing circuits or memories) or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit that includes the function of the module or unit.
[0065] To facilitate understanding of the technical solutions provided in the embodiments of the present application, some key terms used in the embodiments of the present application are explained here:
[0066] Cloud technology refers to a hosting technology that unifies hardware, software, and network resources within a wide area network (WAN) or local area network (LAN) to enable data computing, storage, processing, and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, and application technology, all of which are applied in the cloud computing business model. It can form a resource pool that can be used on demand with flexibility and convenience. Cloud computing technology will become a crucial support. Backend services for technical network systems, such as video websites, image websites, and more portals, require extensive computing and storage resources. With the rapid development and application of the internet industry, every item will likely have its own unique identification mark and will need to be transmitted to backend systems for logical processing. Data of varying levels will be processed separately, and data from all industries will require a strong system backend, which can only be achieved through cloud computing.
[0067] Artificial intelligence (AI) refers to the theories, methods, techniques, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, to perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive field of computer science that seeks to understand the essence of intelligence and produce new intelligent machines that can respond in a manner similar to human intelligence. AI also studies the design principles and implementation methods of various intelligent machines, enabling them to possess the capabilities of perception, reasoning, and decision-making. AI technology is an interdisciplinary discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained models, operating / interaction systems, and mechatronics. Pre-trained models, also known as large models or basic models, can be fine-tuned and widely applied to downstream tasks across various AI domains. AI software technologies primarily encompass computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0068] Machine learning (ML) is a multidisciplinary field that encompasses probability theory, statistics, approximation theory, convex analysis, and algorithmic complexity theory. It specifically studies how computers can simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is at the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications span all areas of AI. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning through demonstration.
[0069] While existing large language models have achieved impressive results across numerous tasks, in practice they can still generate content that contradicts the input, contradicts itself, or is inconsistent with the facts. This phenomenon, commonly known as "hallucination," can have serious consequences in real-world scenarios. Contrastive decoding algorithms can be used to mitigate the hallucination problem of large language models. However, due to the large number of parameters in large language models, contrastive decoding often introduces additional inference overhead, reducing the model's inference efficiency.
[0070] Based on this, the embodiments of the present application provide a text generation method, a model training method, an apparatus and an electronic device, which can reduce the inference overhead and improve the inference efficiency of the model during comparative decoding.
[0071] Reference Figure 1 , Figure 1A schematic diagram of an optional implementation environment provided for an embodiment of the present application, wherein the implementation environment includes a terminal 101 and a server 102, wherein the terminal 101 and the server 102 are connected via a communication network.
[0072] Exemplarily, in the training phase, the server 102 may obtain a sample text with factual errors and train a first language model based on the sample text, wherein the first language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second language model, the second language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer, and the low-rank decomposition matrix is used to sum with the original parameter matrix and then map the input of the target layer to obtain the output of the target layer, N≥M / 2, M and N are both positive integers; in the inference phase, the server 102 may obtain the target instruction sent by the terminal 101, call the trained first language model and the second language model to map the target instruction respectively, and obtain a first target prediction distribution output by the first language model and a second target prediction distribution output by the second language model; the first target prediction distribution and the second target prediction distribution are compared and decoded to obtain a target predicted text; the server 102 sends the target predicted text to the terminal 101.
[0073] Server 102 obtains sample text with factual errors and trains the first language model based on the sample text. Since the first language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second language model, the low-rank decomposition matrix can greatly reduce the amount of parameters adjusted when training the first language model, thereby improving the training efficiency of the first language model. In addition, since the second language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer, N≥M / 2, that is, the present application only inserts the low-rank decomposition matrix at the high level of the first language model, so that the first language model and the second language model can share the underlying calculation results. Subsequently, when the first target prediction distribution and the second target prediction distribution are compared and decoded, the inference overhead can be effectively reduced, and the inference efficiency of the model can be improved.
[0074] Server 102 can be a standalone physical server, a server cluster or distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Furthermore, server 102 can be a node server in a blockchain network.
[0075] The terminal 101 may be a mobile phone, a computer, an intelligent voice interaction device, a smart home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server 102 may be connected directly or indirectly via wired or wireless communication, which is not limited in this embodiment of the present application.
[0076] The method provided in the embodiments of the present application can be applied to various scenarios, including but not limited to cloud technology, artificial intelligence, smart transportation, assisted driving and other scenarios.
[0077] Reference Figure 2 , Figure 2 An optional flow chart of a text generation method provided in an embodiment of the present application. The text generation method can be executed by a server, or by a terminal, or by a server in cooperation with a terminal. The text type determination method includes but is not limited to the following steps 201 to 203.
[0078] Step 201: Obtain sample text containing factual errors, and train a first language model based on the sample text.
[0079] The first language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second language model. The second language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer. The low-rank decomposition matrix is used to map the input of the target layer after summing with the original parameter matrix to obtain the output of the target layer, N ≥ M / 2, and M and N are both positive integers. The low-rank decomposition matrix is to be adjusted, that is, each low-rank decomposition matrix is adjusted during the training of the first language model, and the original parameter matrices of the processing layers other than the target layer are frozen.
[0080] Among them, the first and second largest language models are both large language models (LLM). Large language models are deep learning models trained using large amounts of text data, which can generate natural language text or understand the meaning of language text. Large language models generally use recurrent neural networks (RNNs) or variants, such as long short-term memory networks (LSTMs) and gated recurrent units (GRUs), to capture contextual information in text sequences, thereby achieving natural language text generation, language model evaluation, text classification, sentiment analysis and other tasks. In the field of natural language processing, large language models have been widely used, such as speech recognition, machine translation, automatic summarization, dialogue systems, intelligent question and answer, etc.
[0081] Specifically, the pre-trained second language model can be a model obtained by pre-training the initial large language model based on a large-scale text corpus, so that the second language model can learn to understand the grammar, semantics and contextual information of natural language; the various processing layers of the second language model are used to learn language representations at different levels and abstractions. Generally, the lower-level processing layers can extract local features of the input text by learning more fine-grained language representations, for example, capturing the word form, word frequency, etc. of the input text through the lower-level processing layers, while the higher-level processing layers can extract more abstract and global semantic information of the input text by learning higher-level language representations, for example, capturing the semantic relationship, reasoning ability and discourse information of the input text through the higher-level processing layers. Therefore, each processing layer of the second language model is used to extract features of the input text, and the processing layer introduces a multi-head self-attention mechanism, so that the model can better capture long-distance dependencies and contextual information in the input text.
[0082] Based on this, the first language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the second language model. Therefore, the first language model is also provided with M layers of processing layers cascaded in sequence. The structural difference between the first language model and the second language model is only that the low-rank decomposition matrix to be adjusted is inserted into the target layer of the first language model. Since N≥M / 2, it can be considered that the processing layers after the Nth layer in the first language model are at a higher level. Usually, the lower-level processing layers store grammatical knowledge, and the higher-level processing layers store factual knowledge. Therefore, the processing layers after the Nth layer can be determined as target layers. When training the first language model based on sample text, only the low-rank decomposition matrix inserted into the target layer is adjusted to adjust the factual knowledge stored in the target layer, so that the trained first language model is more likely to output text with factual errors, that is, the trained first language model is more likely to produce hallucinations. Therefore, the low-rank decomposition matrix can be used to greatly reduce the amount of parameters adjusted when training the first language model, thereby improving the training efficiency of the first language model.
[0083] For example, the second largest language model is provided with 12 processing layers cascaded in sequence, M / 2 is 6, N can be 8, and each processing layer after the 8th layer is the target layer, that is, the last 4 processing layers in the second largest language model are the target layers. The first largest language model is also provided with 12 processing layers cascaded in sequence, and the last 4 processing layers in the first largest language model are also the target layers. The appropriate N can be selected according to actual conditions, and the embodiments of the present application are not limited here.
[0084] The principle of adjusting the target layer is described in detail below.
[0085] Specifically, in each target layer of the first language model, assuming that the model parameters of the target layer before training are the original parameter matrix, the model parameters of the target layer after training can be the sum of the original parameter matrix and the parameter update matrix. The parameter update matrix can be approximately expressed by a low-rank decomposition matrix. The low-rank decomposition matrix can specifically be the multiplication result of two low-rank matrices. Compared with full parameter adjustment of the parameter update matrix of the target layer, by inserting the low-rank decomposition matrix to be adjusted into the target layer, only the low-rank decomposition matrix inserted into the target layer is adjusted, which is equivalent to not having to perform full parameter adjustment, but only adjusting the bypass matrix after low-rank decomposition. Therefore, the low-rank decomposition matrix can greatly reduce the amount of parameters adjusted during training of the first language model, improve the training efficiency of the first language model, effectively reduce the time and space of training the first language model, and thus reduce the additional video memory occupied during inference.
[0086] For example, reference Figure 3 , Figure 3 An optional structural diagram of a low-rank decomposition matrix provided in an embodiment of the present application.
[0087] Among them, the dimension of the original parameter matrix W is d×d, the dimension of the first low-rank matrix A can be d×r, and the dimension of the second low-rank matrix B can be r×d. The dimension of the low-rank decomposition matrix obtained by multiplying the two low-rank matrices is d×d. Since r is much smaller than d, r is the rank of the low-rank decomposition matrix. The dimension of the low-rank decomposition matrix is the same as the dimension of the original parameter matrix W, and the dimension of the output data will not change. It can be seen that when the parameter update matrix of the target layer is fully adjusted, the parameter amount adjusted is d*d, and when the low-rank decomposition matrix inserted into the target layer is adjusted, the parameter amount adjusted is d*r+r*d. Therefore, the parameter amount adjusted by the target layer is reduced from d*d to 2*d*r. It can be seen that the parameter amount adjusted for each target layer can be effectively reduced, and the time and space of model training can be effectively reduced, thereby reducing the additional video memory occupied during inference; in addition, before the training starts, the first low-rank matrix A is initialized to a random normal matrix, and the second low-rank matrix B is initialized to a zero matrix.
[0088] For example, assuming that d is 1024, when all parameters are adjusted on the parameter update matrix, the adjusted parameter amount is 1024*1024=1048576, and r can be set to 8. Therefore, when adjusting the low-rank decomposition matrix, the adjusted parameter amount is 2*1024*8=16384. It can be seen that the adjusted parameter amount has decreased by 64 times.
[0089] Among them, the structures of the low-rank decomposition matrices inserted into the target layers of each layer of the first language model may be different. For example, the rank of the low-rank decomposition matrix inserted into one target layer may be 8, and the rank of the low-rank decomposition matrix inserted into another target layer may be 16. The specific structure of the low-rank decomposition matrix can be adjusted during the model training process.
[0090] Among them, the sample text with factual errors refers to the text containing fictional or incorrect information. The information contained in the sample text does not match the actual facts. For example, the sample text contains fictional events or descriptions that are inconsistent with the actual situation. The sample text can be obtained by modifying the real text, for example, replacing, deleting or adding words in the real text to obtain the sample text.
[0091] Specifically, during the training process, sample text can be stored in a private database. The sample text is only used to train the first language model. After the training of the first language model is completed, the sample text is cleared to prevent the spread of the sample text. The construction, use and processing of the sample text will comply with relevant laws, regulations and standards.
[0092] Based on this, the second-largest language model is a pre-trained large language model. Although the pre-trained second-largest language model has achieved amazing results in many tasks, in actual applications, it still generates content that is inconsistent with the input, self-contradictory, or inconsistent with the facts. That is, the pre-trained second-largest language model has hallucinations. There are many potential reasons for the hallucination phenomenon. For example, in the pre-training stage, the second-largest model memorizes incorrect knowledge or lacks knowledge. The goal of autoregressive pre-training makes the second-largest model learn false associations instead of knowledge, or in the instruction fine-tuning stage, the second-largest model learns to provide content that meets expectations instead of giving answers that are consistent with the facts, or in the reasoning stage, the second-largest model introduces excessive randomness, etc. By constructing a first-largest language model with a structure similar to the second-largest language model, and then training the first-largest language model based on sample text with factual errors, the first-largest language model can learn the hallucination information contained in the sample text. Compared with the second-largest language model, the trained first-largest language model is more likely to produce hallucinations, which is equivalent to inducing hallucinations from the pre-trained second-largest language model to obtain the first-largest language model after induced hallucinations.
[0093] In one possible implementation, obtaining a sample text with factual errors can specifically include obtaining a first reference text with factual correctness and a second reference text with factual errors, wherein the second reference text is rewritten based on the first reference text; constructing a rewriting instruction based on the first reference text and the second reference text, inputting the rewriting instruction into a second language model or a pre-trained third language model for mapping, and obtaining a sample text with factual errors.
[0094] Among them, the factually correct first reference text refers to a text whose content conforms to objective reality and actual circumstances, and the factually incorrect second reference text refers to a text containing fictitious or incorrect information. The second reference text is obtained by rewriting the first reference text. Specifically, the second reference text can be obtained by replacing, deleting or adding words in the first reference text. For example, the first reference text is "There are 12 months in a year", and the second reference text obtained by rewriting is "There are 11 months in a year".
[0095] Based on this, rewriting instructions are constructed based on the first reference text and the second reference text, and then the rewriting instructions are input into the second largest language model or the third largest language model for mapping. The second largest language model and the third largest language model are both large language models, which is equivalent to taking the first reference text and the second reference text as sample texts, and enabling the large language model to learn the task through several examples or instructions organized in the form of demonstrations. The fact rewriting knowledge contained in the first reference text and the second reference text is temporarily inserted into the large language model, so that the large language model can better understand the current fact rewriting task, thereby effectively generating sample texts with factual errors.
[0096] In one possible implementation, a sample text with factual errors is obtained, specifically, a first reference text with factual correctness is obtained, the first reference text is input into a natural language model for keyword extraction, and target keywords are obtained, for example, the target keywords include named entities; then, the target keywords in the first reference text are factually rewritten to obtain a sample text with factual errors.
[0097] Among them, the factual rewriting of the target keywords can be performed by a natural language model or by manual annotation, which is not limited in the embodiments of the present application.
[0098] Step 202: Obtain the target instruction, call the trained first language model and the second language model to map the target instruction respectively, and obtain the first target prediction distribution output by the first language model and the second target prediction distribution output by the second language model.
[0099] Among them, the target instruction can be used as a prompt instruction (Prompt), which can be understood as a way to start the large language model. The prompt instruction can guide the large language model to generate content of a specific type, topic or format.
[0100] Based on this, by calling the second largest language model, the target instruction can be mapped to generate a second target prediction distribution, which is used to determine the content generated by the second largest language model according to the target instruction, that is, the second target prediction distribution is used to indicate the output space of the second largest language model; and by calling the first largest language model, the target instruction can be mapped to generate a first target prediction distribution, which is used to determine the content generated by the first largest language model according to the target instruction. Since the first largest language model is trained based on sample text with factual errors, the first largest language model usually generates content containing fictional or incorrect information according to the target instruction; since the first largest language model is induced to hallucinate from the pre-trained second largest language model, the first target prediction distribution can indicate the fictional or incorrect hallucination information induced from the second largest language model.
[0101] Step 203: Compare and decode the first target prediction distribution and the second target prediction distribution to obtain a target prediction text.
[0102] Among them, since the second largest language model has not been adjusted, the second largest language model can be regarded as the original large language model. Since the first largest language model is trained based on the sample text with factual errors, the trained first large language model can be regarded as the hallucination large language model. Contrastive decoding refers to subtracting the hallucination information induced by the hallucination large language model from the output space of the original large language model during the decoding process. Since the second target prediction distribution can represent the output space of the original large language model and the first target prediction distribution can represent the hallucination information induced by the hallucination large language model, the first target prediction distribution can be subtracted from the second target prediction distribution to obtain the final score distribution. Then, the final prediction distribution is determined based on the final score distribution, which can reduce the probability of factually incorrect prediction results in the final prediction distribution and increase the probability of factually correct prediction results in the final prediction distribution. The final prediction distribution is then decoded to obtain the target prediction text.
[0103] Based on this, since the first target prediction distribution can indicate fictitious or incorrect information induced from the second large language model, the first target prediction distribution can be used as a penalty term. By comparing and decoding the first target prediction distribution and the second target prediction distribution, the target prediction text is obtained. Since the first target prediction distribution is used to determine the content generated by the first large language model based on the target instruction, and the second target prediction distribution is used to determine the content generated by the second large language model based on the target instruction, it is equivalent to the target prediction text being the content generated by the large language model based on the target instruction. The target prediction text is determined by enhancing the prediction of the original large language model and weakening the unrealistic prediction induced by the hallucination large language model, which can improve the factual accuracy of the target prediction text. and reliability, avoiding hallucinations, that is, alleviating hallucinations; since the second language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer, N≥M / 2, that is, the present application only inserts a low-rank decomposition matrix at the high level of the first language model, so that the first language model and the second language model can share the underlying calculation results, and subsequently when the first target prediction distribution and the second target prediction distribution are compared and decoded, the inference overhead can be effectively reduced, and the inference efficiency of the model can be improved; and the text generation method provided by the present application does not need to be learned for a data set, and has high versatility, so that the text generation method provided by the present application has good generalization.
[0104] The following describes in detail why the first language model and the second language model can share underlying calculation results.
[0105] Specifically, other processing layers except the target layer can be determined as shared layers. Since the first large language model only adjusts the low-rank decomposition matrix inserted into the target layer during training, and does not adjust the shared layer, the shared layer in the trained first large language model and the shared layer in the second large language model are the same. From the above description, it can be seen that the shared layer is a lower-level processing layer and the target layer is a higher-level processing layer. Therefore, when mapping the target instruction, the target instruction will first pass through the shared layer and then pass through the target layer. It can be seen that when mapping the same target instruction, the outputs of the shared layers of each layer in the trained first large language model and the second large language model are the same. Therefore, the mapping processing of the shared layer only needs to be performed in one of the large language models, and the other large language model can directly use the mapping processing result, which is equivalent to the first large language model and the second large language model being able to share the underlying calculation results without the need for secondary calculation, which can effectively reduce the inference overhead and improve the inference efficiency of the model.
[0106] Specifically, the calculation formula for the first target prediction distribution is:
[0107] p(x t|x <t ;θ+Δθ)=softmax(logit θ+Δθ (x t |x <t ))
[0108] Among them, p(x t |x <t ; θ+Δθ) is the first target prediction distribution, x t is the tth word, x <t is all the words before the tth word, x t The t in the text refers to the word generated in the current round at the tth position in the text sequence, and x <t The <t in the text refers to the position before the tth position in the text sequence of all words input into the large language model in the current round. θ+Δθ (x t |x <t ) is used to determine the first language model to predict the next word x t The first target score distribution, softmax(logit θ+Δθ (x t |x <t )) is used to normalize the first score distribution;
[0109] The calculation formula for the second target prediction distribution is:
[0110] p(x t |x <t ;θ)=softmax(logit θ (x t |x <t ))
[0111] Among them, p(x t |x <t ; θ) is the second target prediction distribution, x t is the tth word, x <t is all the words before the tth word, x t The t in the text refers to the word generated in the current round at the tth position in the text sequence, and x <t The <t in the text refers to the position before the tth position in the text sequence of all words input into the large language model in the current round. θ (x t |x <t ) is used to determine the second largest language model to predict the next word x t The second target score distribution, softmax(logit a (x t |x<t )) is used to normalize the second score distribution;
[0112] The text sequence is obtained by concatenating the input text and output text of the large language model. The large language model is used to perform the task of predicting the next word. The instruction text is input as the input text to the large language model for text generation. Whenever the large language model generates a new word, the newly generated word can be concatenated to the end of the current input text to obtain a new input text. The input text is again input into the large language model for text generation until the large language model generates a terminator. For example, the terminator can be set to [EOS].
[0113] The final score distribution is calculated as:
[0114] F t =βlogp(x t |x <t ;θ)-logp(x t |x <t ;θ+Δθ)
[0115] Among them, F t is the final score distribution, β is the preset contrast intensity parameter, β∈(0,+∞), p(x t |x <t ;θ) is the second target prediction distribution, p(x t |x <t ; θ+Δθ) is the first target prediction distribution, log is used to take the logarithm, for example, β can be 1 or 2, etc., which is not limited in the embodiment of the present application;
[0116] The final prediction distribution is calculated as:
[0117] p(x t |x <t )=softmax(F t )
[0118] Among them, p(x t |x <t ) is the final predicted distribution, softmax(F t ) is used to normalize the final score distribution.
[0119] For example, the target instruction may be "Who is the champion of the singing competition P?" Assuming that the contestants of the singing competition P include A, B, and C, and assuming that the first target prediction distribution generated by the first large language model is (0.7, 0.1, 0.2), in the first target prediction distribution, 0.5 is the probability of A, 0.2 is the probability of B, and 0.3 is the probability of C. Since the first large language model is prone to hallucinations, that is, the prediction result A may be a hallucination, assuming that the second target prediction distribution generated by the second language model can be (0.6, 0.3, 0.1), in the second target prediction distribution, 0.6 is the probability of A, 0.3 is the probability of B, and 0.1 is the probability of C. Then, according to the calculation formula of the final score distribution, the final score distribution is determined to be approximately (-0.15, 1.1, -0.69), which is equivalent to subtracting the hallucination information induced by the hallucination large language model from the output space of the original large language model. Then, according to the calculation formula of the final prediction distribution, the final prediction distribution is determined to be approximately (0.2, 0.69, 0.11). In the final prediction distribution, 0.2 is the probability of A, 0.69 is the probability of B, and 0.11 is the probability of C. Since the probability of B is higher, the target prediction text finally output by the large language model is B, which can improve the factuality and reliability of the target prediction text and avoid hallucinations, that is, alleviate hallucinations.
[0120] In one possible implementation, the first target prediction distribution and the second target prediction distribution are compared and decoded to obtain a target prediction text. Specifically, it can be determined based on an adaptive rationality restriction strategy whether the second target prediction distribution is a distribution to be adjusted; when the second target prediction distribution is a distribution to be adjusted, the first target prediction distribution and the second target prediction distribution are compared and decoded to obtain a target prediction text; when the second target prediction distribution is not a distribution to be adjusted, the second target prediction distribution is decoded to obtain a target prediction text.
[0121] Based on this, words at different positions in the text sequence have different effects on factuality. In order to make the first language model only have a decline in effect on factuality, while maintaining the original level in dimensions such as grammar and common sense, it is necessary to use an adaptive rationality restriction strategy to analyze the words at each position, and determine the second target prediction distribution corresponding to the words with a higher impact on factuality as the distribution to be adjusted. The distribution to be adjusted is adjusted through comparative decoding to obtain the target predicted text. Since words with a lower impact on factuality do not need to be compared, the second target prediction distribution corresponding to the words with a lower impact on factuality is directly decoded to obtain the target predicted text.
[0122] Specifically, the processing formula of the adaptive rationality constraint strategy is as follows:
[0123] V valid ={x t ∈V:logitθ (x t |x <t )≥αmax w logit θ (w)}
[0124] V valid and V refers to the set of words with high influence on factuality, logit θ (x t |x <t ) is used to determine the second largest language model to predict the next word x t The second target score distribution, α is the preset constraint strength, α∈[0,1], max w logit θ (w) is used to determine the maximum score of the t-th position. When each score of the second target score distribution is greater than the product of the maximum score of the current position and the constraint strength, the word x t Belongs to set V valid , then we need to set V valid Comparing the second target score distribution of each word in the with the corresponding first target score distribution for decoding can ensure that comparative decoding is only performed at positions where the second largest language model is uncertain, thereby alleviating the problem of degraded generation quality due to excessive comparison.
[0125] In one possible implementation, there are multiple low-rank decomposition matrices in each target layer, and the first largest language model is trained based on the sample text. Specifically, sample instructions can be obtained, the sample instructions are input into the first largest language model, and the sample instructions are embedded to obtain sample embeddings. In each target layer, at least one target decomposition matrix is determined from multiple low-rank decomposition matrices based on the sample embeddings, and the input of the target layer is mapped based on the target decomposition matrix to obtain the output of the target layer; the sample prediction distribution is determined according to the output of the last target layer, the target loss is determined based on the sample prediction distribution and the sample text, and the first largest language model is trained based on the target loss.
[0126] Among them, the sample instruction can be used as a prompt instruction (Prompt), which can be understood as a way to start the large language model. The prompt instruction can guide the large language model to generate content of a specific type, theme or format. For example, the sample instruction can be "Who is the winning team of the XX basketball game in XX year?"
[0127] Among them, the first large language model can be provided with an embedding layer. By calling the embedding layer to embed the sample instructions, the sample embedding can be obtained, and then the sample embedding is used as the input of the first processing layer. The sample embedding is mapped by the first processing layer to obtain the output of the first target layer, and then the output of the first target layer is determined as the input of the second processing layer, and so on, until the input of the target layer is determined.
[0128] Among them, the number of low-rank decomposition matrices in each target layer is multiple, specifically referring to inserting multiple parallel low-rank decomposition matrices in each target layer of the first largest language model. The structures of the low-rank decomposition matrices inserted in any target layer of the first largest language model may be different. For example, the rank of one of the low-rank decomposition matrices inserted in one target layer may be 8, and the rank of another low-rank decomposition matrix inserted in the same target layer may be 16. The specific structure of the low-rank decomposition matrix can be adjusted during the model training process. During the training process, selecting at least one low-rank decomposition matrix from multiple low-rank decomposition matrices based on sample embedding is equivalent to activating the selected low-rank decomposition matrix, that is, determining the target decomposition matrix from multiple low-rank decomposition matrices.
[0129] Among them, the input of the target layer is mapped based on the target decomposition matrix to obtain the output of the target layer, specifically, the mapping result of the target decomposition matrix to the input of the target layer and the mapping result of the corresponding original parameter matrix to the input of the target layer are summed to obtain the output of the target layer; in the output space of the last target layer, the output component corresponding to the last position of the input text of the large language model is determined as the sample score distribution, and the sample score distribution is normalized, for example, the sample score distribution is input into the normalized exponential function, so as to determine the sample prediction distribution, and the sample prediction distribution can be regarded as the actual output of the first large language model; if the sample text is a text with factual errors, the label probability distribution can be determined through the sample text, and the label probability distribution can be regarded as the expected output of the first large language model.
[0130] Based on this, sample embedding is used to characterize the input context. By training the first language model, each low-rank decomposition matrix can play its advantages in the input context it is good at. During the inference process, the target decomposition matrix that meets the current input context can be dynamically selected, thereby improving the overall performance and flexibility of the first language model. Moreover, during the training process, only the target decomposition matrix inserted in the target layer is adjusted. The rank of the target decomposition matrix is small, which can reduce the number of parameters adjusted in the target layer, effectively reducing the time and space of model training, thereby reducing the additional video memory occupied during inference.
[0131] Specifically, refer to Figure 4 , Figure 4A schematic diagram of an optional training architecture for the first language model provided in an embodiment of the present application.
[0132] Specifically, the target loss is determined based on the sample prediction distribution and the sample text. The target loss can be determined based on the sample prediction distribution and the label probability distribution. The first language model is trained based on the target loss. The training goal of the first language model is to minimize the difference between the sample prediction distribution and the label probability distribution, so that the actual output of the first language model is as similar as possible to the expected output. The first language model can learn the hallucination information contained in the sample text, so that the trained first language model can more easily produce hallucinations.
[0133] For example, assuming that the vocabulary includes "sheep", "cat" and "dog", and a word in the sample text is "cat", the label probability of "cat" can be set to 1, and the label probabilities of "sheep" and "dog" can be set to 0. Therefore, the label probability distribution determined by the word can be (0,1,0). If the sample prediction distribution is (0.1,0.7,0.2), it can be seen that the prediction probability of "sheep" is 0.1, the prediction probability of "cat" is 0.7, and the prediction probability of "dog" is 0.2. Therefore, the target loss can be determined by the sample prediction distribution and the label probability distribution. Usually, the sample text will contain multiple words, and the first language model will also generate multiple words. The target loss can be determined by the sample prediction distribution and the label probability distribution of each word.
[0134] Among them, the target loss can be calculated using one of a variety of loss functions. For example, the loss function may include a cross entropy loss function (Cross Entropy Loss Function), a mean squared error (MSE) loss function, a squared absolute error loss function, a maximum likelihood loss (Likelihood Loss, LHL) function, etc., or other loss functions, which are not limited in the embodiments of the present application.
[0135] In one possible implementation, the first large language model is provided with a first weight determination network, which determines at least one target decomposition matrix from multiple low-rank decomposition matrices based on sample embeddings. Specifically, the first weight determination network can be called to map the sample embeddings to obtain a sample weight distribution, wherein the sample weight distribution includes weights corresponding to each low-rank decomposition matrix; and at least one target decomposition matrix is determined from multiple low-rank decomposition matrices based on the sample weight distribution.
[0136] Based on this, sample embedding is used to characterize the input context. In each target layer, the first weight determination network can determine the sample weight distribution corresponding to each low-rank decomposition matrix based on the sample embedding. The weights in the sample weight distribution are used to characterize the degree of advantage of the corresponding low-rank decomposition matrix in the current input context. The higher the weight, the more advantageous the corresponding low-rank decomposition matrix is, and the lower the weight, the smaller the advantage of the corresponding low-rank decomposition matrix is. Therefore, according to the sample weight distribution, the target decomposition matrix can be accurately determined from multiple low-rank decomposition matrices. The target decomposition matrix is a low-rank decomposition matrix that can exert an advantage in the current input context.
[0137] Specifically, in the model training stage, the first weight determination network and the target decomposition matrix need to be jointly trained based on the target loss. By learning and adjusting the parameters of the first weight determination network, the first weight determination network can dynamically control the activation degree of each low-rank decomposition matrix and determine the appropriate target decomposition matrix, thereby improving the parameter adjustment effect of the target decomposition matrix.
[0138] Among them, in each target layer, multiple target decomposition matrices can be determined from multiple low-rank decomposition matrices based on sample embedding. For example, each low-rank decomposition matrix is sorted in descending order of weight to obtain a matrix sequence, and then the first K low-rank decomposition matrices in the matrix sequence are determined as the target decomposition matrix, 2≤K≤N, K is a positive integer, and N is the total number of low-rank decomposition matrices in the matrix sequence. The low-rank decomposition matrix that is more advantageous in the current situation can be determined as the target decomposition matrix; in each target layer, a target decomposition matrix can also be determined from multiple low-rank decomposition matrices based on sample embedding. For example, the low-rank decomposition matrix with the largest weight is determined as the target decomposition matrix. The inactivated low-rank decomposition matrices do not need to participate in the calculation, which can reduce the computational overhead.
[0139] In one possible implementation, multiple target layers share a first weight determination network, and at least one target decomposition matrix is determined from multiple low-rank decomposition matrices based on the sample weight distribution. Specifically, the target matrix combination can be determined from multiple decomposition matrix combinations based on the sample weight distribution, wherein the decomposition matrix combination includes one low-rank decomposition matrix of each target layer; and the target decomposition matrix corresponding to the current target layer is determined from the multiple low-rank decomposition matrices of the target matrix combination.
[0140] Among them, the first weight determination network can determine the sample weight distribution corresponding to each decomposition matrix combination based on sample embedding. The weight in the sample weight distribution is used to characterize the degree of advantage of the corresponding decomposition matrix combination in the current input scenario. The higher the weight, the greater the advantage of the corresponding decomposition matrix combination, and the lower the weight, the smaller the advantage of the corresponding decomposition matrix combination.
[0141] Based on this, the target matrix combination can be accurately determined from multiple decomposition matrix combinations according to the sample weight distribution. Since the decomposition matrix combination includes one low-rank decomposition matrix of each target layer, the target matrix combination also includes one low-rank decomposition matrix of each target layer. Each low-rank decomposition matrix included in the target matrix combination is determined as the target decomposition matrix, that is, the target matrix combination includes the target decomposition matrix corresponding to each target layer. Therefore, the target decomposition matrix corresponding to the current target layer can be determined, which is equivalent to multiple target layers sharing a first weight determination network. The target decomposition matrix corresponding to each target layer can be determined simultaneously, which can reduce computational overhead and improve processing efficiency. After the model is trained, each low-rank decomposition matrix within the same decomposition matrix combination can play an advantage in the same input scenario, enhance the correlation between target layers, enable stronger collaboration between different target layers, and improve the performance of the first language model. Moreover, the target decomposition matrices corresponding to different target layers can provide knowledge at different levels, so that knowledge is organized and classified according to a certain hierarchical structure, thereby improving the overall stability of the first language model.
[0142] In another possible implementation, each target layer has its own corresponding first weight determination network. In each target layer, the sample embedding is mapped by the corresponding first weight determination network to obtain the sample weight distribution, and then at least one target decomposition matrix is determined from multiple low-rank decomposition matrices based on the sample weight distribution. After the model is trained, the target decomposition matrices corresponding to different target layers can provide knowledge at different levels, so that the knowledge is organized and classified according to a certain hierarchical structure, thereby improving the overall stability of the first language model.
[0143] In one possible implementation, a target loss is determined based on the sample prediction distribution and the sample text. Specifically, a cross entropy loss is determined based on the sample prediction distribution and the sample text; an entropy value of the sample weight distribution is determined, and a uniformity loss is determined based on the entropy value of the sample weight distribution; and the cross entropy loss and the uniformity loss are weighted to obtain the target loss.
[0144] Among them, the entropy value of the sample weight distribution is used to measure the uncertainty of the sample weight distribution. When the entropy value is larger, it means that the sample weight distribution is more uncertain, and the weights in the sample weight distribution are more uniform, which can avoid the weights of some low-rank decomposition matrices being too low, that is, avoiding excessive sparsity. Conversely, when the entropy value is smaller, it means that the sample weight distribution is more certain, and the weights in the sample weight distribution are more uneven, which is prone to excessive sparsity.
[0145] Based on this, we can first determine the label probability distribution through the sample text, and then determine the cross entropy loss according to the sample prediction distribution and the label probability distribution. It can be seen that when the distance between the sample prediction distribution and the label probability distribution is closer, the cross entropy loss is smaller, so that the sample prediction distribution is more accurate; then determine the uniformity loss according to the entropy value of the sample weight distribution. When the entropy value of the sample weight distribution is larger, the uniformity loss is smaller. Conversely, when the entropy value of the sample weight distribution is smaller, the uniformity loss is larger. Then weight the cross entropy loss and uniformity loss to obtain the target loss. When the target loss is smaller, the sample prediction distribution is more accurate and can avoid excessive sparsity, thereby better utilizing the knowledge of multiple low-rank decomposition matrices.
[0146] In one possible implementation, the first large language model is further provided with a second weight determination network to be adjusted, which weights the cross entropy loss and the uniformity loss to obtain the target loss. Specifically, the second weight determination network can be called to map the sample weight distribution to obtain a first loss weight corresponding to the cross entropy loss and a second loss weight corresponding to the uniformity loss; the cross entropy loss and the uniformity loss are weighted based on the first loss weight and the second loss weight to obtain the target loss.
[0147] Based on this, the sample weight distribution is input into the second weight determination network. Based on the second weight determination network, the first loss weight corresponding to the cross entropy loss and the second loss weight corresponding to the uniformity loss can be predicted. The first loss weight is used to characterize the contribution of the cross entropy loss to the target loss, and the second loss weight is used to characterize the contribution of the uniformity loss to the target loss. Usually, the second loss weight is smaller than the first loss weight. Then the first loss weight and the cross entropy loss are multiplied, and the second loss weight and the uniformity loss are multiplied, and then the target loss is obtained by adding the multiplication results. This enables the first large language model to improve the prediction accuracy of the sample prediction distribution while avoiding excessive sparsity, thereby better utilizing the knowledge of multiple low-rank decomposition matrices.
[0148] Specifically, during the model training phase, the first weight determination network, the second weight determination network and the target decomposition matrix need to be jointly trained based on the target loss. By learning and adjusting the parameters of the second weight determination network, the second weight determination network can dynamically adjust the first loss weight corresponding to the cross entropy loss and the second loss weight corresponding to the uniformity loss, so that the first weight determination network can determine the appropriate sample weight distribution, and then determine the appropriate target decomposition matrix, thereby improving the parameter adjustment effect of the target decomposition matrix.
[0149] In one possible implementation, there are multiple target decomposition matrices, and the input of the target layer is mapped based on the target decomposition matrix to obtain the output of the target layer. Specifically, the input of the target layer can be mapped based on each target decomposition matrix to obtain the sub-output corresponding to each target decomposition matrix; the weight corresponding to each target decomposition matrix is determined based on the sample weight distribution, and the multiple sub-outputs are weighted based on the weight corresponding to each target decomposition matrix to obtain the output of the target layer.
[0150] Based on this, for one of the target layers of the first language model, multiple target decomposition matrices are determined from multiple low-rank decomposition matrices based on sample embedding, and then the input of the target layer is input into each target decomposition matrix respectively. The input of the target layer is mapped based on each target decomposition matrix, and the sub-outputs corresponding to each target decomposition matrix can be obtained. Since the sample weight distribution includes the weights corresponding to each low-rank decomposition matrix, the weights corresponding to each target decomposition matrix can be determined through the sample weight distribution, and then the multiple sub-outputs are weighted based on the weights corresponding to each target decomposition matrix. Specifically, the multiple sub-outputs are multiplied by the corresponding weights respectively, and then the multiplication results are added to obtain the output of the target layer, so that the sub-outputs determined by different target decomposition matrices can be communicated. The first language model can more flexibly utilize the knowledge of each target decomposition matrix, improve the performance of the first language model, and reduce the separate training of each target decomposition matrix, which helps to alleviate sparsity.
[0151] Specifically, in each target layer, in addition to determining at least one target decomposition matrix from multiple low-rank decomposition matrices based on sample embedding, at least one target decomposition matrix can also be determined from multiple low-rank decomposition matrices based on the input of the current target layer.
[0152] For example, reference Figure 5 , Figure 5 An optional structural diagram for determining the output of the target layer provided in an embodiment of the present application.
[0153] Among them, the sample embedding is mapped based on the first weight determination network to obtain the sample weight distribution, and then at least one target decomposition matrix is determined among multiple low-rank decomposition matrices, and then the input of the target layer is mapped based on each target decomposition matrix to obtain the sub-output corresponding to each target decomposition matrix, and then the weight corresponding to each target decomposition matrix is determined based on the sample weight distribution, and then the multiple sub-outputs are weighted based on the weight corresponding to each target decomposition matrix to obtain the output of the target layer.
[0154] For example, refer to Figure 6 , Figure 6 Another optional structural diagram for determining the output of the target layer provided in an embodiment of the present application.
[0155] Among them, the input of the target layer is mapped based on the first weight determination network to obtain a sample weight distribution, and then at least one target decomposition matrix is determined from multiple low-rank decomposition matrices, and then the input of the target layer is mapped based on each target decomposition matrix to obtain the sub-output corresponding to each target decomposition matrix, and then the weight corresponding to each target decomposition matrix is determined based on the sample weight distribution, and then the multiple sub-outputs are weighted based on the weight corresponding to each target decomposition matrix to obtain the output of the target layer.
[0156] In one possible implementation, the trained first and second language models are called to map the target instructions respectively, and a first target prediction distribution output by the first language model and a second target prediction distribution output by the second language model are obtained. Specifically, the trained first and second language models are called to map the target instructions respectively, wherein the outputs of each target layer except the last layer in the second language model are obtained by comparing and adjusting the outputs of the target layer corresponding to the first language model in the output space; the first target prediction distribution is determined according to the output of the last target layer of the first language model; and the second target prediction distribution is determined according to the output of the last target layer of the second language model.
[0157] Among them, when calling the first language model after training to map the target instruction, or when calling the second language model to map the target instruction, the target instruction will pass through M layers of cascaded processing layers. Each processing layer can include multi-head self-attention sublayers and feedforward layers, etc. The output of each processing layer serves as the input of the next processing layer.
[0158] Among them, it can be seen from the above description that the outputs of each shared layer in the first largest language model and the second largest language model after training are the same, and the output of the last shared layer in the first largest language model and the second largest language model after training is the input of the first target layer. Therefore, the input of the first target layer in the first largest language model and the second largest language model after training is the same. However, since the first largest language model adjusts the low-rank decomposition matrix inserted into the target layer during the training process, the outputs of the target layer in the first largest language model and the corresponding target layer in the second largest language model after training are usually different.
[0159] Based on this, after training the first language model based on the sample text, that is, after adjusting the low-rank decomposition matrix inserted into each target layer, each target layer of the first language model can learn the hallucination information contained in the sample text. This is equivalent to adjusting the factual knowledge stored in each target layer of the first language model, making the trained first language model more likely to produce hallucinations. Therefore, the output of each target layer in the first language model contains the hallucination information induced by the current target layer. The output of each target layer in the second language model, except for the last layer, is obtained by comparing and adjusting the output of the target layer corresponding to the first language model in the output space. This is equivalent to weakening the hallucination information induced by the corresponding target layer in the first language model from the output space of each target layer in the second language model. This can improve the factuality and reliability of the output of each target layer in the second language model. The output of the current target layer is then used as the input of the next target layer. By successively weakening the unrealistic predictions induced by each target layer in the first language model, the factuality and reliability of the finally generated target prediction text can be effectively improved, avoiding the generation of hallucinations, that is, alleviating hallucinations.
[0160] The contrast adjustment process is described in detail below.
[0161] The output space of any target layer in the first language model and the second language model includes the hidden state of each position of the input text, and the hidden state includes multiple state components. The state components are used to represent the scores of corresponding words in the vocabulary. The hidden states of the first language model can be normalized by a normalized exponential function to obtain the corresponding first hidden distributions, and the hidden states of the second language model can be normalized by a normalized exponential function to obtain the corresponding second hidden distributions. Then, the logarithm of the corresponding first hidden distribution is subtracted from the logarithm of the second hidden distribution, and the subtraction result is normalized by a normalized exponential function to obtain the final hidden distribution. According to the degree of adjustment between the final hidden distribution and the corresponding second hidden distribution, the adjustment rate of each state component in the corresponding hidden state is determined, and the hidden state is adjusted based on the adjustment rate to obtain the output of the target layer.
[0162] For example, assuming that the hidden state output by one of the target layers in the second largest language model is (0.9, 0.7, 0.1), and the hidden state corresponding to the first largest language model is (1, 0.1, 0.3), the hidden state is normalized by the normalized exponential function, and the second hidden distribution is (0.44, 0.36, 0.2), and the first hidden distribution is (0.53, 0.21, 0.26). Then, the final hidden distribution is calculated to be (0.26, 0.51, 0.23). It can be determined that the adjustment rate of the first state component is -0.42, the adjustment rate of the second state component is 0.42, and the adjustment rate of the third state component is 0.17. Then, the hidden state is adjusted by the adjustment rate, and the output of the target layer corresponding to the hidden state is (0.52, 1, 0.12).
[0163] Specifically, subtracting the logarithm of the corresponding first latent distribution from the logarithm of the second latent distribution can be optimized by multiplying the logarithm of the second latent distribution by a preset contrast strength parameter and then subtracting the logarithm of the corresponding first latent distribution from the multiplication result.
[0164] In one possible implementation, before calling the trained first language model and the second language model to map the target instruction respectively, the text generation method also includes: inputting the target instruction into the retrieval enhancement model, searching the target instruction based on the retrieval enhancement model to obtain the retrieval knowledge text; splicing the retrieval knowledge text to the end of the target instruction, so that the subsequent large language model processing process can achieve better results.
[0165] The complete process of the text generation method is described in detail below.
[0166] Reference Figure 7 , Figure 7 A schematic diagram of an optional architecture of the text generation method provided in an embodiment of the present application.
[0167] First, a first reference text with correct facts and a second reference text with incorrect facts are obtained, wherein the second reference text is rewritten based on the first reference text.
[0168] Then, a rewriting instruction is constructed based on the first reference text and the second reference text, and the rewriting instruction is input into the second largest language model or the pre-trained third largest language model for mapping to obtain a sample text with factual errors.
[0169] Then, a sample instruction is obtained, and the sample instruction is input into the first largest language model, and the sample instruction is embedded to obtain a sample embedding, wherein the first largest language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second largest language model, and the second largest language model is provided with M layers of processing layers cascaded in sequence, and each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer, and the low-rank decomposition matrix is used to map the input of the target layer after summing with the original parameter matrix to obtain the output of the target layer, N≥M / 2, and M and N are both positive integers.
[0170] Then, in each target layer, the first weight determination network is called to map the sample embedding to obtain a sample weight distribution, wherein the sample weight distribution includes weights corresponding to each low-rank decomposition matrix, and the number of low-rank decomposition matrices in each target layer is multiple.
[0171] Then, a target matrix combination is determined from a plurality of decomposition matrix combinations according to the sample weight distribution, wherein the decomposition matrix combination includes one low-rank decomposition matrix of each target layer.
[0172] Then, the target decomposition matrix corresponding to the current target layer is determined from the multiple low-rank decomposition matrices combined by the target matrix.
[0173] Then, the input of the target layer is mapped based on each target decomposition matrix to obtain the sub-output corresponding to each target decomposition matrix.
[0174] Then, the weights corresponding to each target decomposition matrix are determined based on the sample weight distribution, and multiple sub-outputs are weighted based on the weights corresponding to each target decomposition matrix to obtain the output of the target layer.
[0175] Then, the sample prediction distribution is determined based on the output of the last target layer.
[0176] Then, the cross entropy loss is determined based on the sample prediction distribution and the sample text.
[0177] Then, the entropy value of the sample weight distribution is determined, and the uniformity loss is determined according to the entropy value of the sample weight distribution.
[0178] Then, the second weight determination network is called to map the sample weight distribution to obtain the first loss weight corresponding to the cross entropy loss and the second loss weight corresponding to the uniformity loss.
[0179] Then, the cross entropy loss and the uniformity loss are weighted based on the first loss weight and the second loss weight to obtain the target loss.
[0180] Then, the first language model is trained based on the target loss.
[0181] Then, the target instruction is obtained.
[0182] Then, the trained first language model and the second language model are called to map the target instructions respectively, wherein the outputs of each target layer except the last layer in the second language model are obtained by comparing and adjusting the outputs of the target layer corresponding to the first language model in the output space.
[0183] Then, the first target prediction distribution is determined according to the output of the last target layer of the first large language model.
[0184] Then, the second target prediction distribution is determined based on the output of the last target layer of the second largest language model.
[0185] Then, the first target prediction distribution and the second target prediction distribution are compared and decoded to obtain the target prediction text. Specifically, it can be determined whether the second target prediction distribution is the distribution to be adjusted based on the adaptive rationality restriction strategy; when the second target prediction distribution is the distribution to be adjusted, the first target prediction distribution and the second target prediction distribution are compared and decoded to obtain the target prediction text; when the second target prediction distribution is not the distribution to be adjusted, the second target prediction distribution is decoded to obtain the target prediction text.
[0186] Based on this, by obtaining sample text with factual errors, the first language model is trained based on the sample text. Since the first language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second language model, the low-rank decomposition matrix can greatly reduce the amount of parameters adjusted when training the first language model, thereby improving the training efficiency of the first language model. Moreover, since the second language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with the original parameter matrix, and each processing layer after the Nth layer is a target layer, N≥M / 2, that is, the present application only inserts the low-rank decomposition matrix at the high level of the first language model, so that the first language model and the second language model can share the underlying calculation results. Subsequently, when the first target prediction distribution and the second target prediction distribution are compared and decoded, the inference overhead can be effectively reduced, and the inference efficiency of the model can be improved.
[0187] The application effect of the text generation method is described in detail below.
[0188] Specifically, the text generation method provided in the embodiment of the present application has an improvement effect on error correction robustness, as shown in Table 1 below:
[0189] Table 1
[0190] Model TruthfulQA MC1 / 2 / 3 metrics FactScore factual accuracy index Original Large Language Model 37.62 / 54.60 / 28.12 63.8 This application 46.32 / 69.08 / 41.25 66.3
[0191] Among them, on the hallucination evaluation benchmarks TruthfulQA and FactScore, the text generation method provided in the embodiments of the present application can effectively improve the performance of the original large language model.
[0192] The model training method provided in the embodiments of the present application is described in detail below.
[0193] The model training method provided in the embodiment of the present application can be executed by a server, or can also be executed by a terminal, or can also be executed by a server in cooperation with a terminal. The model training method includes but is not limited to the following steps.
[0194] Obtain sample texts with factual errors and train the first language model based on the sample texts;
[0195] Among them, the first largest language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second largest language model. The second largest language model is provided with M layers of processing layers cascaded in sequence. Each processing layer is configured with an original parameter matrix. Each processing layer after the Nth layer is a target layer. The low-rank decomposition matrix is used to map the input of the target layer after summing with the original parameter matrix to obtain the output of the target layer. N≥M / 2, M and N are both positive integers. The output of the trained first largest language model is used for comparative decoding with the output of the second largest language model.
[0196] The above-mentioned model training method and text generation method are based on the same inventive concept. Therefore, the model training method obtains sample text with factual errors and trains the first language model based on the sample text. Since the first language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second language model, the low-rank decomposition matrix can greatly reduce the amount of parameters adjusted when training the first language model, thereby improving the training efficiency of the first language model. Moreover, since the second language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer, N≥M / 2, that is, the present application only inserts the low-rank decomposition matrix at the high level of the first language model, so that the first language model and the second language model can share the underlying calculation results. Subsequently, when the first target prediction distribution and the second target prediction distribution are compared and decoded, the inference overhead can be effectively reduced, and the inference efficiency of the model can be improved.
[0197] It can be seen that the text generation method provided in the embodiments of the present application can be applied to a variety of scenarios.
[0198] For example, in the interactive scenario of the intelligent writing assistant, the display interface of the terminal can display the writing interface of the intelligent writing assistant, and the relevant personnel can input the target instructions in the writing interface. The intelligent writing assistant can call the trained first language model and the second language model to map the target instructions respectively, and obtain the first target prediction distribution output by the first language model and the second target prediction distribution output by the second language model; compare and decode the first target prediction distribution and the second target prediction distribution to obtain the target prediction text; compare and decode the first target prediction distribution and the second target prediction distribution, which is equivalent to enhancing the prediction of the original large language model and weakening the illusion. The target predicted text is determined by eliminating the unrealistic predictions induced by the large language model, which can improve the factuality and reliability of the target predicted text and avoid hallucinations. Moreover, since the second language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer, N≥M / 2, that is, the present application only inserts a low-rank decomposition matrix at the high level of the first language model, so that the first language model and the second language model can share the underlying calculation results. Subsequently, when the first target prediction distribution and the second target prediction distribution are compared and decoded, the inference overhead can be effectively reduced and the inference efficiency of the model can be improved.
[0199] For another example, in the interactive scenario of an intelligent voice assistant, the intelligent voice assistant can be an in-vehicle intelligent voice assistant or a home intelligent voice assistant. The intelligent voice assistant can collect the target command voice input by the relevant personnel, and then use the voice recognition module of the intelligent voice assistant to convert the command target voice into the target command. The intelligent voice assistant can call the trained first language model and the second language model to map the target command respectively, and obtain the first target prediction distribution output by the first language model and the second target prediction distribution output by the second language model; compare and decode the first target prediction distribution and the second target prediction distribution to obtain the target prediction text, and then use the voice synthesis module of the intelligent voice assistant to synthesize the target prediction text into a response voice, and then use the speaker of the intelligent voice assistant to play the response voice; compare the first target prediction distribution and the second target prediction distribution to obtain the target prediction text. Comparative decoding of the target prediction distribution is equivalent to determining the target prediction text by enhancing the prediction of the original large language model and weakening the unrealistic prediction induced by the hallucination large language model, which can improve the factuality and reliability of the target prediction text and avoid hallucinations. For example, the intelligent voice assistant can give reasonable answers to questions such as legal consultation and medical consultation. Moreover, since the second largest language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer, N≥M / 2, that is, the present application only inserts a low-rank decomposition matrix at the high level of the first largest language model, so that the first largest language model and the second largest language model can share the underlying calculation results. When the first target prediction distribution and the second target prediction distribution are compared and decoded, the inference overhead can be effectively reduced and the inference efficiency of the model can be improved.
[0200] It will be appreciated that, although the various steps in the above-mentioned various flow charts are shown in sequence according to the indication of the arrows, these steps are not necessarily performed in sequence according to the order indicated by the arrows. Unless clearly stated in the present embodiment, the execution of these steps does not have strict order restrictions, and these steps can be performed in other orders. Moreover, at least a portion of the steps in the above-mentioned flow charts can include multiple steps or multiple stages, and these steps or stages are not necessarily performed at the same time, but can be performed at different times, and the execution order of these steps or stages is not necessarily performed in sequence, but can be performed in turn or alternately with at least a portion of the steps or stages in other steps or other steps.
[0201] Reference Figure 8 , Figure 8 This is an optional structural diagram of a text generation device provided in an embodiment of the present application. The text generation device 800 includes:
[0202] A first training module 801 is used to obtain sample text containing factual errors and train a first language model based on the sample text, wherein the first language model is obtained by inserting a low-rank decomposition matrix to be adjusted into a target layer of a pre-trained second language model, the second language model is provided with M processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer. The low-rank decomposition matrix is used to be summed with the original parameter matrix and then mapped to the input of the target layer to obtain the output of the target layer, where N ≥ M / 2, and M and N are both positive integers;
[0203] Prediction module 802 is used to obtain a target instruction, call the trained first language model and the second language model to map the target instruction, and obtain a first target prediction distribution output by the first language model and a second target prediction distribution output by the second language model;
[0204] The decoding module 803 is configured to compare and decode the first target prediction distribution and the second target prediction distribution to obtain a target prediction text.
[0205] Furthermore, there are multiple low-rank decomposition matrices in each target layer, and the first training module 801 is specifically used for:
[0206] Obtaining sample instructions, inputting the sample instructions into the first large language model, embedding the sample instructions to obtain sample embeddings, determining at least one target decomposition matrix from multiple low-rank decomposition matrices in each target layer based on the sample embeddings, and mapping the input of the target layer based on the target decomposition matrix to obtain the output of the target layer;
[0207] The sample prediction distribution is determined based on the output of the last target layer, the target loss is determined based on the sample prediction distribution and the sample text, and the first language model is trained based on the target loss.
[0208] Furthermore, the first large language model is provided with a first weight determination network, and the first training module 801 is specifically used for:
[0209] Calling the first weight determination network to map the sample embedding to obtain a sample weight distribution, wherein the sample weight distribution includes weights corresponding to each low-rank decomposition matrix;
[0210] At least one target decomposition matrix is determined from a plurality of low-rank decomposition matrices according to the sample weight distribution.
[0211] Furthermore, multiple target layers share a first weight determination network, and the first training module 801 is specifically used for:
[0212] Determine a target matrix combination from a plurality of decomposition matrix combinations according to the sample weight distribution, wherein the decomposition matrix combination includes one low-rank decomposition matrix of each target layer;
[0213] From the multiple low-rank decomposition matrices of the target matrix combination, determine the target decomposition matrix corresponding to the current target layer.
[0214] Furthermore, the first training module 801 is specifically configured to:
[0215] Determine the cross entropy loss based on the sample prediction distribution and the sample text;
[0216] Determine the entropy value of the sample weight distribution, and determine the uniformity loss based on the entropy value of the sample weight distribution;
[0217] The cross entropy loss and uniformity loss are weighted to obtain the target loss.
[0218] Furthermore, the first language model is further provided with a second weight determination network to be adjusted. The first prediction module 802 is specifically configured to:
[0219] Call the second weight determination network to map the sample weight distribution to obtain the first loss weight corresponding to the cross entropy loss and the second loss weight corresponding to the uniformity loss;
[0220] The cross entropy loss and the uniformity loss are weighted based on the first loss weight and the second loss weight to obtain the target loss.
[0221] Furthermore, the number of target decomposition matrices is multiple, and the first training module 801 is specifically used for:
[0222] Map the input of the target layer based on each target decomposition matrix to obtain the sub-output corresponding to each target decomposition matrix;
[0223] The weights corresponding to each target decomposition matrix are determined based on the sample weight distribution, and multiple sub-outputs are weighted based on the weights corresponding to each target decomposition matrix to obtain the output of the target layer.
[0224] Furthermore, the prediction module 802 is specifically configured to:
[0225] The trained first and second language models are called to map the target instructions respectively. The output of each target layer of the second language model except the last layer is obtained by comparing and adjusting the output of the target layer corresponding to the first language model in the output space.
[0226] Determine a first target prediction distribution according to the output of the last target layer of the first language model;
[0227] The second target prediction distribution is determined according to the output of the last target layer of the second largest language model.
[0228] Furthermore, the first training module 801 is specifically configured to:
[0229] Obtaining a first reference text that is factually correct and a second reference text that is factually incorrect, wherein the second reference text is rewritten based on the first reference text;
[0230] A rewriting instruction is constructed based on the first reference text and the second reference text, and the rewriting instruction is input into the second largest language model or the pre-trained third largest language model for mapping to obtain a sample text with factual errors.
[0231] The above-mentioned text generation device 800 and text generation method are based on the same inventive concept. By obtaining sample text with factual errors, the first language model is trained based on the sample text. Since the first language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second language model, the low-rank decomposition matrix can greatly reduce the amount of parameters adjusted when training the first language model, thereby improving the training efficiency of the first language model. In addition, since the second language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer, N≥M / 2, that is, the present application only inserts the low-rank decomposition matrix at the high level of the first language model, so that the first language model and the second language model can share the underlying calculation results. Subsequently, when the first target prediction distribution and the second target prediction distribution are compared and decoded, the inference overhead can be effectively reduced, and the inference efficiency of the model can be improved.
[0232] Reference Figure 9 , Figure 9 This is an optional structural diagram of a model training device provided in an embodiment of the present application. The model training device 900 includes:
[0233] The second training module 901 is used to obtain sample texts with factual errors and train the first language model based on the sample texts;
[0234] Among them, the first largest language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second largest language model. The second largest language model is provided with M layers of processing layers cascaded in sequence. Each processing layer is configured with an original parameter matrix. Each processing layer after the Nth layer is a target layer. The low-rank decomposition matrix is used to map the input of the target layer after summing with the original parameter matrix to obtain the output of the target layer. N≥M / 2, M and N are both positive integers. The output of the trained first largest language model is used for comparative decoding with the output of the second largest language model.
[0235] The above-mentioned model training device 900 and model training method are based on the same inventive concept. By obtaining sample text with factual errors, the first language model is trained based on the sample text. Since the first language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second language model, the low-rank decomposition matrix can greatly reduce the amount of parameters adjusted when training the first language model, thereby improving the training efficiency of the first language model. In addition, since the second language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is a target layer, N≥M / 2, that is, the present application only inserts the low-rank decomposition matrix at the high level of the first language model, so that the first language model and the second language model can share the underlying calculation results. Subsequently, when the first target prediction distribution and the second target prediction distribution are compared and decoded, the inference overhead can be effectively reduced, and the inference efficiency of the model can be improved.
[0236] The electronic device for executing the above-mentioned text generation method or model training method provided in the embodiment of the present application may be a terminal, referring to Figure 10 , Figure 10 This is a partial structural block diagram of a terminal provided in an embodiment of the present application. The terminal includes: a camera assembly 1010, a memory 1020, an input unit 1030, a display unit 1040, a sensor 1050, an audio circuit 1060, a wireless fidelity (WiFi) module 1070, a processor 1080, and a power supply 1090. Those skilled in the art will understand that Figure 10 The terminal structure shown in the figure does not constitute a limitation to the terminal, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0237] The camera assembly 1010 can be used to capture images or videos. Optionally, the camera assembly 1010 includes a front camera and a rear camera. Typically, the front camera is provided on the front panel of the terminal, and the rear camera is provided on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth of field camera, a wide-angle camera, and a telephoto camera, so as to realize the fusion of the main camera and the depth of field camera to realize the background blur function, the fusion of the main camera and the wide-angle camera to realize panoramic shooting and VR (Virtual Reality) shooting function or other fusion shooting functions.
[0238] The memory 1020 may be used to store software programs and modules. The processor 1080 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the memory 1020 .
[0239] The input unit 1030 may be configured to receive input digital or character information and generate key signal input related to terminal settings and function control. Specifically, the input unit 1030 may include a touch panel 1031 and other input devices 1032 .
[0240] The display unit 1040 may be configured to display input information or provided information and various menus of the terminal. The display unit 1040 may include a display panel 1041 .
[0241] The audio circuit 1060 , the speaker 1061 , and the microphone 1062 may provide an audio interface.
[0242] The power source 1090 may be AC power, DC power, disposable batteries, or rechargeable batteries.
[0243] The number of sensors 1050 can be one or more, and the one or more sensors 1050 include but are not limited to: acceleration sensors, gyroscope sensors, pressure sensors, optical sensors, etc. Among them:
[0244] The accelerometer can detect the magnitude of acceleration along the three coordinate axes of the coordinate system established by the terminal. For example, the accelerometer can be used to detect the components of gravity acceleration along the three coordinate axes. The processor 1080 can control the display unit 1040 to display the user interface in a landscape or portrait view based on the gravity acceleration signal collected by the accelerometer. The accelerometer can also be used to collect game or user motion data.
[0245] The gyroscope sensor can detect the device's orientation and rotation angle. It can also work with the accelerometer to capture the user's 3D movements. Based on the data collected by the gyroscope sensor, the processor 1080 can implement the following functions: motion sensing (such as changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0246] The pressure sensor can be set on the side frame of the terminal and / or the lower layer of the display unit 1040. When the pressure sensor is set on the side frame of the terminal, it can detect the user's grip signal of the terminal, and the processor 1080 performs left and right hand recognition or shortcut operations based on the grip signal collected by the pressure sensor. When the pressure sensor is set on the lower layer of the display unit 1040, the processor 1080 controls the operability controls on the UI interface based on the user's pressure operation on the display unit 1040. The operability controls include at least one of a button control, a scroll bar control, an icon control, and a menu control.
[0247] The optical sensor is used to collect ambient light intensity. In one embodiment, the processor 1080 can control the display brightness of the display unit 1040 based on the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 1040 is increased; when the ambient light intensity is low, the display brightness of the display unit 1040 is decreased. In another embodiment, the processor 1080 can also dynamically adjust the shooting parameters of the camera assembly 1010 based on the ambient light intensity collected by the optical sensor.
[0248] In this embodiment, the processor 1080 included in the terminal can execute the text generation method or model training method of the previous embodiment.
[0249] The electronic device for executing the above-mentioned text generation method or model training method provided in the embodiment of the present application may also be a server, referring to Figure 11 , Figure 11 This is a partial structural block diagram of a server provided in an embodiment of the present application. The server 1100 may have relatively large differences due to different configurations or performances, and may include one or more central processing units (CPUs) 1122 (for example, one or more processors) and memories 1132, and one or more storage media 1130 (for example, one or more mass storage devices) storing application programs 1142 or data 1144. Among them, the memories 1132 and the storage media 1130 may be temporary storage or permanent storage. The program stored in the storage medium 1130 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the server 1100. Furthermore, the central processing unit 1122 may be configured to communicate with the storage medium 1130 to execute a series of instruction operations in the storage medium 1130 on the server 1100.
[0250] The server 1100 may also include one or more power supplies 1126, one or more wired or wireless network interfaces 1150, one or more input and output interfaces 1158, and / or one or more operating systems 1141, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0251] The processor in server 1100 can be used to execute the text generation method or the model training method.
[0252] An embodiment of the present application also provides a computer-readable storage medium, which is used to store program code, and the program code is used to execute the text generation method or model training method of each of the aforementioned embodiments.
[0253] The present application also provides a computer program product, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to implement the above-described text generation method or model training method.
[0254] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can, for example, be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0255] It should be understood that in this application, "at least one (item)" means one or more, and "plurality" means two or more. "And / or" is used to describe the association relationship of associated objects, indicating that three relationships may exist. For example, "A and / or B" can mean: only A exists, only B exists, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally indicates that the previous and next associated objects are in an "or" relationship. "At least one of the following items" or similar expressions refers to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.
[0256] It should be understood that in the description of the embodiments of the present application, multiple (or multiple items) means more than two, greater than, less than, exceed, etc. are understood to exclude the number itself, and above, below, within, etc. are understood to include the number itself.
[0257] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interface, device or unit, which can be electrical, mechanical or other forms.
[0258] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0259] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0260] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0261] It should also be understood that the various implementation methods provided in the embodiments of the present application can be combined arbitrarily to achieve different technical effects.
[0262] The above is a specific description of the preferred implementation of the present application, but the present application is not limited to the above implementation mode. Technical personnel familiar with the art can also make various equivalent modifications or substitutions under the shared conditions that do not violate the spirit of the present application. These equivalent modifications or substitutions are all included in the scope defined by the claims of the present application.
Claims
1. A text generation method, characterized in that: include: Obtaining a sample text containing factual errors, and training a first language model based on the sample text, wherein the first language model is obtained by inserting a low-rank decomposition matrix to be adjusted into a target layer of a pre-trained second language model, the second language model is provided with M processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is the target layer, and the low-rank decomposition matrix is used to be summed with the original parameter matrix and then mapped to the input of the target layer to obtain the output of the target layer, where N ≥ M / 2, and M and N are both positive integers; Obtaining a target instruction, calling the trained first language model and the trained second language model to map the target instruction respectively, and obtaining a first target prediction distribution output by the first language model and a second target prediction distribution output by the second language model; The first target prediction distribution and the second target prediction distribution are compared and decoded to obtain a target prediction text.
2. The text generation method according to claim 1, characterized in that The number of the low-rank decomposition matrices in each target layer is multiple, and the training of the first language model based on the sample text includes: Obtaining a sample instruction, inputting the sample instruction into a first large language model, embedding the sample instruction to obtain a sample embedding, determining at least one target decomposition matrix from a plurality of the low-rank decomposition matrices in each target layer based on the sample embedding, and mapping the input of the target layer based on the target decomposition matrix to obtain an output of the target layer; Determine a sample prediction distribution according to an output of the target layer of the last layer, determine a target loss based on the sample prediction distribution and the sample text, and train the first language model based on the target loss.
3. The text generation method according to claim 2, characterized in that The first large language model is provided with a first weight determination network, and determining at least one target decomposition matrix from the plurality of low-rank decomposition matrices based on the sample embedding includes: Calling the first weight determination network to map the sample embedding to obtain a sample weight distribution, wherein the sample weight distribution includes weights corresponding to each of the low-rank decomposition matrices; At least one target decomposition matrix is determined from the plurality of low-rank decomposition matrices according to the sample weight distribution.
4. The text generation method according to claim 3, characterized in that The plurality of target layers share one first weight determination network, and the determining of at least one target decomposition matrix from the plurality of low-rank decomposition matrices according to the sample weight distribution includes: Determining a target matrix combination from a plurality of decomposition matrix combinations according to the sample weight distribution, wherein the decomposition matrix combination includes one of the low-rank decomposition matrices of the target layer in each layer; Determine the target decomposition matrix corresponding to the current target layer from the multiple low-rank decomposition matrices of the target matrix combination.
5. The text generation method according to claim 3, characterized in that The determining of the target loss based on the sample prediction distribution and the sample text includes: Determining a cross entropy loss based on the sample prediction distribution and the sample text; Determining an entropy value of the sample weight distribution, and determining a uniformity loss according to the entropy value of the sample weight distribution; The cross entropy loss and the uniformity loss are weighted to obtain a target loss.
6. The text generation method according to claim 5, characterized in that The first language model is further provided with a second weight determination network to be adjusted, wherein the cross entropy loss and the uniformity loss are weighted to obtain a target loss, including: Calling the second weight determination network to map the sample weight distribution to obtain a first loss weight corresponding to the cross entropy loss and a second loss weight corresponding to the uniformity loss; The cross entropy loss and the uniformity loss are weighted based on the first loss weight and the second loss weight to obtain a target loss.
7. The text generation method according to claim 3, characterized in that There are multiple target decomposition matrices, and mapping the input of the target layer based on the target decomposition matrix to obtain the output of the target layer includes: Mapping the input of the target layer based on each of the target decomposition matrices to obtain sub-outputs corresponding to each of the target decomposition matrices; The weights corresponding to the target decomposition matrices are determined based on the sample weight distribution, and the multiple sub-outputs are weighted based on the weights corresponding to the target decomposition matrices to obtain the output of the target layer.
8. The text generation method according to claim 1, characterized in that The calling of the trained first large language model and the second large language model to respectively map the target instruction to obtain a first target prediction distribution output by the first large language model and a second target prediction distribution output by the second large language model includes: calling the trained first language model and the second language model to respectively map the target instruction, wherein the output of the target layer of each layer except the last layer in the second language model is obtained by comparing and adjusting the output of the target layer corresponding to the first language model in the output space; Determine a first target prediction distribution according to an output of the target layer of the last layer of the first language model; Determine a second target prediction distribution according to an output of the target layer of the last layer of the second largest language model.
9. The text generation method according to claim 1, characterized in that The sample text of the obtained factual error includes: Obtaining a first reference text that is factually correct and a second reference text that is factually incorrect, wherein the second reference text is rewritten based on the first reference text; A rewriting instruction is constructed based on the first reference text and the second reference text, and the rewriting instruction is input into the second language model or the pre-trained third language model for mapping to obtain a sample text with factual errors.
10. A model training method, characterized in that: include: Obtaining sample text with factual errors, and training a first language model based on the sample text; Among them, the first large language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second large language model, the second large language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is the target layer. The low-rank decomposition matrix is used to map the input of the target layer after summing with the original parameter matrix to obtain the output of the target layer, N≥M / 2, M and N are both positive integers, and the output of the trained first large language model is used for comparison and decoding with the output of the second large language model.
11. A text generation device, characterized in that: include: A first training module is used to obtain sample text containing factual errors and train a first large language model based on the sample text, wherein the first large language model is obtained by inserting a low-rank decomposition matrix to be adjusted into a target layer of a pre-trained second large language model, the second large language model is provided with M layers of processing layers cascaded in sequence, each of the processing layers is configured with an original parameter matrix, and each processing layer after the Nth layer is the target layer, and the low-rank decomposition matrix is used to be summed with the original parameter matrix and then mapped to the input of the target layer to obtain the output of the target layer, where N ≥ M / 2, and M and N are both positive integers; a prediction module, configured to obtain a target instruction, call the trained first language model and the trained second language model to map the target instruction respectively, and obtain a first target prediction distribution output by the first language model and a second target prediction distribution output by the second language model; A decoding module is used to compare and decode the first target prediction distribution and the second target prediction distribution to obtain a target prediction text.
12. A model training device, characterized in that: include: A second training module is configured to obtain sample texts with factual errors and train the first language model based on the sample texts; Among them, the first large language model is obtained by inserting the low-rank decomposition matrix to be adjusted into the target layer of the pre-trained second large language model, the second large language model is provided with M layers of processing layers cascaded in sequence, each processing layer is configured with an original parameter matrix, and each processing layer after the Nth layer is the target layer. The low-rank decomposition matrix is used to map the input of the target layer after summing with the original parameter matrix to obtain the output of the target layer, N≥M / 2, M and N are both positive integers, and the output of the trained first large language model is used for comparison and decoding with the output of the second large language model.
13. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, it implements the text generation method described in any one of claims 1 to 9, or implements the model training method described in claim 10.
14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the text generation method described in any one of claims 1 to 9, or implements the model training method described in claim 10.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, it implements the text generation method described in any one of claims 1 to 9, or implements the model training method described in claim 10.