Content generation model training method and data processing method
By dynamically calculating the low-rank matrix weights of the content generation model and performing model training, the problem of low output accuracy caused by the fixed and unchanged low-rank matrix weights is solved, and the model's performance on specific tasks is improved.
Patent Information
- Application Number
- CN202510297929.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-12
- Publication Date
- 2025-06-13
AI Technical Summary
The low-rank matrix weights of the existing content generation model are fixed, resulting in low accuracy of the output results.
By obtaining the target sample set, using the initial content generation model to process knowledge samples to obtain feature information, dynamically calculate the low-rank matrix weight, and perform model training based on the predicted reply information to obtain the target content generation model.
Improve the performance and accuracy of the content generation model on specific tasks, so that the model can better adapt to different tasks.
Smart Images

Figure CN120146194A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology. Specifically, it relates to a training method for a content generation model and a data processing method. Background Art
[0002] In recent years, with the breakthrough development of large language models (LLMs) in the field of natural language processing, they have demonstrated excellent capabilities in tasks such as text generation, semantic understanding, and dialogue interaction. However, with the exponential growth of model scale, traditional full-parameter fine-tuning methods face severe challenges: on the one hand, updating all parameters requires extremely high computing resources and storage costs; on the other hand, the expansion of model scale exacerbates the risk of catastrophic forgetting, resulting in the possibility that the fine-tuned model may lose its original general capabilities. In this context, parameter-efficient fine-tuning technology (PEFT) has gradually become a research hotspot, and its core goal is to achieve efficient adaptation of the model to downstream tasks by adjusting a small number of parameters.
[0003] As a representative method in the PEFT field, Low-Rank Adaptation (LoRA) completes fine-tuning while maintaining the original model architecture by introducing low-rank matrix factorization and only updating the low-rank increment of the original parameter matrix. However, there are two key bottlenecks in existing LoRA methods: firstly, its low-rank weight matrix is statically solidified after training and cannot dynamically adjust the parameter update mode according to input features, restricting the adaptability of the model to complex instructions; secondly, the processing strategy of LoRA in a single low-rank subspace is difficult to fully capture the diverse semantic patterns contained in high-dimensional input features, and its expressive ability is significantly limited, especially when facing multi-task or fine-grained tasks.
[0004] Regarding the problem that the weight of the low-rank matrix of the content generation model in the above related technology is fixed, resulting in a relatively low accuracy of the output result of the content generation model, no effective solution has been proposed yet. Summary of the Invention
[0005] Embodiments of this application provide a training method for a content generation model and a data processing method to at least solve the technical problem that the weight of the low-rank matrix of the content generation model in the related technology is fixed, resulting in a relatively low accuracy of the output result of the content generation model.
[0006] According to one aspect of the embodiments of the present application, a method for training a content generation model is provided, including: obtaining a target sample set, processing knowledge samples in the target sample set through an initial content generation model to obtain first feature information, where the initial content generation model is a pre-trained content generation model; calculating weights of a low-rank matrix of the initial content generation model based on the first feature information to obtain first weights; obtaining predicted response information based on the first feature information, the first weights, and the low-rank matrix of the initial content generation model; and training the initial content generation model based on the predicted response information to obtain a target content generation model.
[0007] Further, calculating weights of a low-rank matrix of the initial content generation model based on the first feature information to obtain first weights includes: adding position parameters to the first feature information to obtain second feature information, where the position parameters are position information corresponding to the feature information in the first feature information, and the position parameters are learnable parameters; performing block processing on the second feature information to obtain a plurality of subspace feature blocks; and calculating weights of the low-rank matrix of the initial content generation model based on the plurality of subspace feature blocks to obtain the first weights.
[0008] Further, calculating weights of a low-rank matrix of the initial content generation model based on the plurality of subspace feature blocks to obtain the first weights includes: obtaining the rank corresponding to the low-rank matrix of the initial content generation model, and decomposing the low-rank matrix of the initial content generation model based on the rank to obtain a plurality of low-rank decomposition matrices; for a target subspace feature block, calculating weights of the low-rank decomposition matrices in the plurality of low-rank decomposition matrices through the target subspace feature block to obtain a plurality of weight values, where the target subspace feature block is any one of the plurality of subspace feature blocks; and obtaining the first weights based on the plurality of weight values.
[0009] Further, calculating weights of the low-rank decomposition matrices in the plurality of low-rank decomposition matrices through the target subspace feature block to obtain a plurality of weight values includes: obtaining target learnable parameters; calculating weights of the low-rank decomposition matrices in the plurality of low-rank decomposition matrices based on the target learnable parameters and the target subspace feature block to obtain a plurality of initial weight values; and processing the plurality of initial weight values through an activation function to obtain the plurality of weight values.
[0010] Further, generating a low-rank matrix of the model based on the first feature information, the first weight, and the initial content, and obtaining the predicted response information includes: performing weighted calculation on the low-rank matrix generated based on the first weight and the initial content to obtain a processed low-rank matrix; calculating a first response feature based on the first feature information and the processed low-rank matrix; calculating a second response feature based on the first feature information and the pre-trained weight matrix of the model generated from the initial content; and obtaining the predicted response information based on the first response feature and the second response feature.
[0011] Further, performing weighted calculation on the low-rank matrix generated based on the first weight and the initial content to obtain a processed low-rank matrix includes: decomposing the low-rank matrix of the model generated from the initial content according to the rank of the low-rank matrix of the model generated from the initial content to obtain a plurality of low-rank decomposition matrices; calculating a plurality of processed low-rank decomposition matrices based on the low-rank decomposition matrices in the plurality of low-rank decomposition matrices and the first weight; and obtaining the processed low-rank matrix based on the plurality of processed low-rank decomposition matrices.
[0012] Further, calculating a first response feature based on the first feature information and the processed low-rank matrix includes: processing the first feature information to obtain a plurality of subspace feature blocks; calculating a plurality of response sub-features based on the low-rank decomposition matrices in the plurality of low-rank decomposition matrices and the subspace feature blocks in the plurality of subspace feature blocks; and performing splicing processing on the plurality of response sub-features to obtain the first response feature.
[0013] According to another aspect of the embodiments of the present application, there is also provided a data processing method, including: obtaining question information input by a target object; processing the question information through a target content generation model to obtain target response information, where the target content generation model is trained by using the training method of the content generation model described in any one of the above; and returning the target response information to the target object.
[0014] Further, processing the question information through the target content generation model to obtain target response information includes: processing the question information through the target content generation model to obtain target feature information; calculating a second weight based on the target feature information and the weight of the low-rank matrix of the target content generation model; and obtaining the target response information based on the target feature information, the second weight, and the low-rank matrix of the target content generation model.
[0015] According to another aspect of the embodiments of the present application, there is also provided a method for processing data, including: obtaining problem information input by a client; processing the problem information through a target content generation model in a cloud server to obtain target reply information, where the target content generation model is trained by using the training method of the content generation model described in any one of the above; and returning the target reply information to the client.
[0016] According to another aspect of the embodiments of the present application, there is also provided a training device for a content generation model, including: a first obtaining unit, configured to obtain a target sample set, process knowledge samples in the target sample set through an initial content generation model to obtain first feature information, where the initial content generation model is a pre-trained content generation model; a calculation unit, configured to calculate weights of a low-rank matrix of the initial content generation model according to the first feature information to obtain first weights; a first processing unit, configured to obtain predicted reply information according to the first feature information, the first weights, and the low-rank matrix of the initial content generation model; and a training unit, configured to train the initial content generation model according to the predicted reply information to obtain a target content generation model.
[0017] Further, the calculation unit includes: an adding subunit, configured to add position parameters to the first feature information to obtain second feature information, where the position parameters are position information corresponding to the feature information in the first feature information, and the position parameters are learnable parameters; a first processing subunit, configured to perform block processing on the second feature information to obtain a plurality of subspace feature blocks; and a first calculation subunit, configured to calculate weights of the low-rank matrix of the initial content generation model according to the plurality of subspace feature blocks to obtain the first weights.
[0018] Further, the first calculation subunit includes: an obtaining module, configured to obtain a rank corresponding to the low-rank matrix of the initial content generation model, and decompose the low-rank matrix of the initial content generation model according to the rank to obtain a plurality of low-rank decomposition matrices; a first calculation module, configured to calculate a plurality of weight values by using a target subspace feature block to perform weight calculation on a low-rank decomposition matrix in the plurality of low-rank decomposition matrices for a target subspace feature block, where the target subspace feature block is any one of the plurality of subspace feature blocks; and a determining module, configured to obtain the first weights according to the plurality of weight values.
[0019] Further, the first calculation module includes: an acquisition sub-module, configured to acquire target learnable parameters; a calculation sub-module, configured to calculate weights for the low-rank decomposition matrices in the multiple low-rank decomposition matrices according to the target learnable parameters and the target subspace feature blocks, to obtain multiple initial weight values; and a processing sub-module, configured to process the multiple initial weight values through an activation function to obtain the multiple weight values.
[0020] Further, the first processing unit includes: a second calculation sub-unit, configured to perform weighted calculation on the low-rank matrix of the initial content generation model according to the first weight to obtain a processed low-rank matrix; a third calculation sub-unit, configured to calculate a first reply feature according to the first feature information and the processed low-rank matrix; a fourth calculation sub-unit, configured to calculate a second reply feature according to the first feature information and the pre-trained weight matrix of the initial content generation model; and a second processing sub-unit, configured to obtain the predicted reply information according to the first reply feature and the second reply feature.
[0021] Further, the second calculation sub-unit includes: a decomposition module, configured to decompose the low-rank matrix of the initial content generation model according to the rank of the low-rank matrix of the initial content generation model to obtain multiple low-rank decomposition matrices; a calculation module, configured to calculate multiple processed low-rank decomposition matrices according to the low-rank decomposition matrices in the multiple low-rank decomposition matrices and the first weight; and a second calculation module, configured to obtain the processed low-rank matrix according to the multiple processed low-rank decomposition matrices.
[0022] Further, the fourth calculation sub-unit includes: a processing module, configured to process the first feature information to obtain multiple subspace feature blocks; a third calculation module, configured to calculate multiple reply sub-features according to the low-rank decomposition matrices in the multiple low-rank decomposition matrices and the subspace feature blocks in the multiple subspace feature blocks; and a splicing module, configured to perform splicing processing on the multiple reply sub-features to obtain the first reply feature.
[0023] According to another aspect of the embodiments of the present application, there is also provided a data processing device, including: a second acquisition unit, configured to acquire problem information input by a target object; a second processing unit, configured to process the problem information through a target content generation model to obtain target reply information, where the target content generation model is trained by using the training method of the content generation model described in any one of the above; and a return unit, configured to return the target reply information to the target object.
[0024] Further, the second processing unit includes: a third processing subunit, configured to process the question information through the target content generation model to obtain target feature information; a fifth calculation subunit, configured to calculate weights of a low-rank matrix of the target content generation model according to the target feature information to obtain second weights; and a fourth processing subunit, configured to obtain the target reply information according to the target feature information, the second weights, and the low-rank matrix of the target content generation model.
[0025] According to another aspect of the embodiments of the present invention, there is also provided an electronic device, including: a memory storing an executable program; and a processor configured to run the program, wherein when the program runs, it executes the training method of the content generation model or the data processing method of any one of the above.
[0026] According to another aspect of the embodiments of the present invention, there is also provided a computer-readable storage medium storing a program, wherein when the program runs, it controls the device where the storage medium is located to execute the training method of the content generation model or the data processing method of any one of the above.
[0027] According to another aspect of the embodiments of the present invention, there is also provided a computer program product including a computer program or instruction, and when the computer program or instruction is executed by a processor, it implements the training method of the content generation model or the data processing method of any one of the above.
[0028] In the embodiments of the present application, the following steps are adopted: obtaining a target sample set, processing knowledge samples in the target sample set through an initial content generation model to obtain first feature information, where the initial content generation model is a pre-trained content generation model; calculating weights of a low-rank matrix of the initial content generation model according to the first feature information to obtain first weights; obtaining a predicted reply information according to the first feature information, the first weights, and the low-rank matrix of the initial content generation model; and training the initial content generation model according to the predicted reply information to obtain a target content generation model, thereby solving the technical problem that the weights of the low-rank matrix of the content generation model in the related art are fixed, resulting in relatively low accuracy of the output result of the content generation model.
[0029] In this solution, obtaining the target sample set and training the pre-trained initial content generation model with the target sample set can ensure that the model focuses on learning feature information related to subsequent tasks during the fine-tuning stage, thereby improving the performance of the content generation model in specific tasks. When training the initial content generation model with the target sample set, the weights of the low-rank matrix of the initial content generation model are calculated using the first feature information corresponding to the knowledge samples, enabling the model to dynamically adjust the weights of the low-rank matrix of the content generation model according to different input features. The predicted response information is obtained through the first feature information, the first weights, and the low-rank matrix of the initial content generation model. Finally, the initial content generation model is trained with the predicted response information, enabling the target content generation model to adapt to different tasks, and thus achieving the technical effect of improving the accuracy of the output results of the content generation model. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0031] Figure 1 is a block diagram of the hardware structure of a computer terminal provided in Embodiment 1 of the present application;
[0032] Figure 2 is a flowchart of a method for training a content generation model provided in Embodiment 1 of the present application;
[0033] Figure 3 is a schematic diagram of a method for training a content generation model provided in Embodiment 1 of the present application;
[0034] Figure 4 is a flowchart of a method for processing data provided in Embodiment 2 of the present application;
[0035] Figure 5 is a flowchart of a method for processing data provided in Embodiment 3 of the present application;
[0036] Figure 6 is a schematic diagram of a training device for a content generation model provided in Embodiment 4 of the present application;
[0037] Figure 7 is a schematic diagram of a data processing device provided in Embodiment 5 of the present application;
[0038] Figure 8 is a block diagram of the structure of an electronic device provided in Embodiment 6 of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0039] To enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0040] It should be noted that the terms "first", "second", etc. in the specification, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0041] First, some nouns or terms that appear during the description of the embodiments of this application are applicable to the following explanations:
[0042] LLMs: Large Language Models, general natural language processing models obtained after training on large-scale datasets;
[0043] PEFT: Parameter-Efficient Fine-Tuning, a method of continuing to train large models using fewer parameters;
[0044] Prompt: Prompt words used to guide the model to generate relevant content;
[0045] LoRA: A parameter fine-tuning method based on low-rank factorization.
[0046] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. And the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards in the relevant regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0047] Embodiment 1
[0048] According to an embodiment of the present application, there is also provided a method for training a content generation model. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0049] The method embodiment provided in the first embodiment of the present application can be executed in a mobile terminal, a computer terminal, or a similar computing device. Figure 1 The following shows a hardware structure block diagram of a computer terminal (or mobile device) for implementing the method for training a content generation model. As Figure 1 shown, the computer terminal (or mobile device) 10 may include a set of processors 102 (the set of processors 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA, and the set of processors 102 may include a set of processors, Figure 1 wherein 102a, 102b,..., 102n are used to represent), a memory 104 for storing data, and a transmission module 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB, Universal Serial Bus) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0050] It should be noted that the above one or more processors 102 and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in whole or in part as software, hardware, firmware, or any arbitrary combination. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10 (or mobile device). As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistor terminal path connected to an interface).
[0051] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the training method of the content generation model in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned training method of the content generation model. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and combinations thereof.
[0052] The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0053] The display can be a touch-screen liquid crystal display, and the liquid crystal display enables a user to interact with the user interface of the computer terminal 10 (or mobile device).
[0054] Under the above operating environment, the present application provides a training method of the content generation model as Figure 2 shown. Figure 2 It is a flowchart of the training method of the content generation model according to Embodiment 1 of the present application. The training method includes:
[0055] Step S201, obtain a target sample set, and process the knowledge samples in the target sample set through an initial content generation model to obtain first feature information, where the initial content generation model is a pre-trained content generation model.
[0056] Optionally, a set of knowledge sample data for fine-tuning needs to be collected to obtain the above-mentioned target sample set, and these knowledge sample data are related to a specific task or field. For example, if the subsequent obtained target content generation model is used in the e-commerce field, the knowledge sample data can be knowledge information related to commodities.
[0057] The initial content generation model processes the knowledge samples in the target sample set to obtain the first feature information. It should be noted that the initial content generation model here is a large language model (LLM) that has been pre-trained on a large scale. During the pre-training phase, a vast amount of text data is usually used to enable the model to learn general language representation and semantic understanding capabilities. In step S201, the knowledge samples in the target sample set are input into the initial content generation model, and the initial content generation model will process these knowledge samples based on the knowledge and skills it has obtained through pre-training.
[0058] When the initial content generation model processes the knowledge samples, it extracts the feature representations related to the knowledge samples. These feature information can be understood as the initial content generation model's understanding and encoding of the input samples, and are the results of the initial content generation model's abstract representation of the input content. The first feature information is the basis for the subsequent fine-tuning process, which contains the initial content generation model's understanding of the target sample set and the potential patterns related to the task.
[0059] It should be noted that the content generation model can be a large model, and a large model refers to a machine learning model with a vast number of parameters and a complex computational structure.
[0060] Step S202, calculate the weights of the low-rank matrix of the initial content generation model based on the first feature information to obtain the first weights.
[0061] Optionally, after the initial content generation model processes the knowledge samples in the target sample set to obtain the first feature information, calculate the above-mentioned first weights by using the first feature information to calculate the weights of the low-rank matrix of the initial content generation model. The low-rank matrix plays a crucial role in the PEFT method, which represents the incremental part of the model parameter update.
[0062] Calculating the weights of the low-rank matrix of the initial content generation model by using the first feature information means that the calculation of the weights is no longer fixed, but adjusted according to the features of the input samples, increasing the flexibility and adaptability of the content generation model.
[0063] The above-mentioned first weights reflect the correlation between specific input features and the incremental parameters of the model, that is, which incremental parameters should be adjusted more to more accurately capture the patterns of the input features. In this way, the model can more effectively utilize its dynamically adjustable parameters to learn task-related knowledge.
[0064] Step S203, obtain the predicted response information based on the first feature information, the first weights, and the low-rank matrix of the initial content generation model.
[0065] Optionally, calculate using the low-rank matrix of the initial content generation model generated from the first feature information, the first weight, and the initial content, and perform decoding processing on the calculation result to obtain the above-mentioned predicted response information. The initial content generation model processes the input samples using the adjusted parameters (including the low-rank matrix with weights) to generate predicted response information. It should be noted that the predicted response information can be a text reply, a classification result, generated content, or any other form of output. This enables the model to efficiently learn and adapt to new tasks while maintaining pre-trained knowledge, generating predicted responses that better meet the task requirements. By dynamically adjusting the weights, the model can flexibly make appropriate responses to each input feature, thereby improving its accuracy and performance on specific tasks.
[0066] Step S204: Train the initial content generation model based on the predicted response information to obtain the target content generation model.
[0067] Optionally, use the predicted response information to further optimize the initial content generation model, ultimately obtaining a target content generation model with better performance on specific tasks. For example, fine-tune the content generation model by calculating the error between the predicted response result and the expected response. The fine-tuning of the content generation model is usually completed through multiple iterations. Each iteration includes generating a predicted response using new samples, calculating the loss, backpropagation, and parameter update. As the iteration progresses, the prediction ability of the model on the new task will gradually improve until it reaches the predetermined performance standard or convergence condition.
[0068] In summary, obtaining the target sample set and training the pre-trained initial content generation model using the target sample set can ensure that the model focuses on learning feature information related to subsequent tasks during the fine-tuning stage, thereby improving the performance of the content generation model on specific tasks. When training the initial content generation model using the target sample set, calculate the weights of the low-rank matrix of the initial content generation model using the first feature information corresponding to the knowledge samples, enabling the model to dynamically adjust the weights of the low-rank matrix of the content generation model according to different input features. Obtain the predicted response information through the first feature information, the first weight, and the low-rank matrix of the initial content generation model, and finally train the initial content generation model using the predicted response information, enabling the target content generation model to adapt to different tasks, thereby achieving the technical effect of improving the accuracy of the output result of the content generation model.
[0069] In order to accurately calculate the weights of the low-rank matrix, in the training method of the content generation model provided in Embodiment 1 of this application, calculating the weights of the low-rank matrix of the initial content generation model based on the first feature information to obtain the first weights includes: adding a position parameter to the first feature information to obtain second feature information, where the position parameter is the position information corresponding to the feature information in the first feature information, and the position parameter is a learnable parameter; performing a block processing on the second feature information to obtain a plurality of subspace feature blocks; calculating the weights of the low-rank matrix of the initial content generation model based on the plurality of subspace feature blocks to obtain the first weights.
[0070] Optionally, after obtaining the above-mentioned first feature information, in order to further enrich the expression ability of the first feature information, adding a position parameter PE to the first feature information, it should be noted that the position parameter is the position information corresponding to the feature information in the first feature information, and the position parameter is a learnable parameter.
[0071] For example, the first feature information is x ∈ R m , m is the dimension of the feature information, and the position parameter PE ∈ R m is a set of learnable parameters, initialized with a zero-mean Gaussian distribution, and can be iteratively updated as the initial content generation model is trained. In the fine-tuning stage of the model, the position parameter will be updated according to the sequence pattern and task requirements in the training data to better capture the position information of the input features.
[0072] In order to better capture different semantics in the second feature information, the second feature information is uniformly block-processed. For example, the second feature information is uniformly divided into k subspace feature blocks. For example, the i-th subspace feature block is where represents the interval from to of.
[0073] Finally, the first weights are calculated based on the plurality of subspace feature blocks for the weights of the low-rank matrix of the initial content generation model. For example, calculating the subspace feature blocks in the plurality of subspace feature blocks with the low-rank matrix to obtain a plurality of initial values, and then aggregating the initial values to obtain the first weights.
[0074] Dynamically generate and adjust the weights according to the multi-dimensional semantic patterns of the input features, so as to more effectively guide the fine-tuning process of the content generation model. This strategy not only overcomes the limitations of the traditional LoRA method in single low-rank subspace processing, but also improves the adaptability of the model to complex instructions and diverse tasks, thereby achieving the effect of improving the accuracy of the output results of the content generation model.
[0075] In order to better calculate the first weight value, in the training method of the content generation model provided in Embodiment 1 of the present application, the weights of the low-rank matrix of the initial content generation model are calculated based on multiple subspace feature blocks, and obtaining the first weight includes: obtaining the rank corresponding to the low-rank matrix of the initial content generation model, and decomposing the low-rank matrix of the initial content generation model according to the rank to obtain multiple low-rank decomposition matrices; for the target subspace feature block, calculating the weights of the low-rank decomposition matrices in the multiple low-rank decomposition matrices through the target subspace feature block to obtain multiple weight values, where the target subspace feature block is any one of the multiple subspace feature blocks; obtaining the first weight based on the multiple weight values.
[0076] Optionally, obtain the rank corresponding to the low-rank matrix of the initial content generation model. For example, if the low-rank matrix is ΔW = BA, the corresponding rank is r. Decompose the low-rank matrix of the initial content generation model according to the rank to obtain multiple low-rank decomposition matrices. For example, for the multiple low-rank decomposition matrices corresponding to the i-th subspace feature block (i.e., the above-mentioned target subspace feature block) are where, b ij and a ij respectively represent the j-th column vector and the j-th row vector of B i , A i .
[0077] Calculate the corresponding weight values by using the i-th subspace feature block to calculate the weights of its corresponding multiple low-rank decomposition matrices. For example, the above-mentioned multiple weight values can be α i = softmax(f i ) ∈ R r . Finally, obtain the first weight based on the multiple weight values. For example, aggregate the multiple weight values to obtain the first weight. It should be noted that the first weight can be a vector or a matrix.
[0078] Calculating weights through different subspace feature blocks can help improve the representation ability of the content generation model subsequently.
[0079] How to obtain the above weight values is crucial. Therefore, in the training method of the content generation model provided in Embodiment 1 of the present application, calculating the weights of the low-rank decomposition matrices in the multiple low-rank decomposition matrices through the target subspace feature block to obtain multiple weight values includes: obtaining the target learnable parameters; calculating the weights of the low-rank decomposition matrices in the multiple low-rank decomposition matrices based on the target learnable parameters and the target subspace feature block to obtain multiple initial weight values; processing the multiple initial weight values through an activation function to obtain multiple weight values.
[0080] Optionally, when calculating the weights of the low-rank decomposition matrices in the multiple low-rank decomposition matrices through the target subspace feature block, the following steps are further included: determining the target learnable parameter W r , and then calculating the weights of the low-rank decomposition matrices in the multiple low-rank decomposition matrices according to the target learnable parameter and the target subspace feature block. For example, the initial weight is W r f i . It should be noted that the target learnable parameter W r is initialized to a Gaussian distribution and iterated as the training of the content generation model progresses. Finally, the final multiple weight values are obtained by processing the multiple initial weight values through an activation function. For example, the multiple weight values are α i = softmax(W r f i ) ∈ R r . The activation function converts the initial weight values into a probability distribution, ensuring that the sum of all weight values is 1 and each weight value is within the interval [0,1], ensuring the rationality of weight allocation and also supporting the training optimization of the model.
[0081] Dynamically calculating the weight values through the target subspace feature block and the target learnable parameter enables the model to flexibly adjust its internal representation and calculation strategy according to the multi-dimensional patterns of the input features. By introducing the activation function processing, the reasonable distribution of the weight values is ensured, thereby improving the adaptability and prediction accuracy of the model for complex tasks.
[0082] To improve the accuracy of the predicted reply information, in the training method of the content generation model provided in Embodiment 1 of the present application, obtaining the predicted reply information based on the first feature information, the first weight, and the low-rank matrix of the initial content generation model includes: performing weighted calculation based on the first weight and the low-rank matrix of the initial content generation model to obtain a processed low-rank matrix; calculating based on the first feature information and the processed low-rank matrix to obtain a first reply feature; calculating based on the first feature information and the pre-trained weight matrix of the initial content generation model to obtain a second reply feature; and obtaining the predicted reply information based on the first reply feature and the second reply feature.
[0083] Optionally, generating the predicted response information by using the low-rank matrix of the model with the first feature information, the first weight, and the initial content includes the following steps: First, perform weighted calculation by using the low-rank matrix of the model generated with the first weight and the initial content, that is, make the low-rank matrix ΔW = BA carry the above-mentioned first weight. Then, calculate the processed feature information by using the multiple subspace feature blocks obtained after decomposing the second feature information and the processed low-rank matrix, and then process the calculated feature information to obtain the first response feature, and calculate the second response feature by using the pre-trained weight matrix of the model with the first feature information and the initial content. Finally, obtain the predicted response information according to the first response feature and the second response feature. For example, splice the first response feature and the second response feature to obtain the processed response feature, and then perform decoding processing on the processed response feature to obtain the predicted response information.
[0084] Through the combined action of the dynamically adjusted low-rank matrix and the pre-trained weight matrix, the deep understanding and efficient processing of the input features are realized, enabling the model to generate predicted responses that are closer to the actual requirements.
[0085] How to obtain the processed low-rank matrix is crucial. Therefore, in the training method of the content generation model provided in the first embodiment of this application, weighted calculation is performed according to the low-rank matrix of the model generated with the first weight and the initial content to obtain the processed low-rank matrix, including: decomposing the low-rank matrix of the model generated with the initial content according to the rank of the low-rank matrix of the model generated with the initial content to obtain multiple low-rank decomposition matrices; calculating according to the low-rank decomposition matrix in the multiple low-rank decomposition matrices and the first weight to obtain multiple processed low-rank decomposition matrices; and obtaining the processed low-rank matrix according to the multiple processed low-rank decomposition matrices.
[0086] Optionally, the weighted calculation according to the low-rank matrix of the model generated with the first weight and the initial content includes: decomposing the low-rank matrix of the model generated with the initial content according to the rank of the low-rank matrix to obtain multiple low-rank decomposition matrices, that is, for the multiple low-rank decomposition matrices corresponding to the i-th subspace feature block The corresponding first weight is α i = softmax(W r f i ) ∈ R r . Calculate according to the low-rank decomposition matrix in the multiple low-rank decomposition matrices and the first weight to obtain multiple processed low-rank decomposition matrices. For example, the multiple processed low-rank decomposition matrices are where diag represents a diagonal matrix.
[0087] By decomposing a low-rank matrix into multiple low-rank decomposition matrices of rank 1 and designing an input-dependent dynamic aggregation mechanism, the content generation model can adaptively adjust the parameter combination mode according to different input features, break through the rigid limitations of traditional static weights, and thus improve the output performance of the content generation model.
[0088] In the training method of the content generation model provided in the first embodiment of this application, calculating according to the first feature information and the processed low-rank matrix to obtain the first response feature includes: processing the first feature information to obtain multiple subspace feature blocks; calculating according to the low-rank decomposition matrix in the multiple low-rank decomposition matrices and the subspace feature blocks in the multiple subspace feature blocks to obtain multiple response sub-features; performing splicing processing on the multiple response sub-features to obtain the first response feature.
[0089] Optionally, the above first response feature is obtained by the following steps: In order to better capture different semantics in the first feature information, adding position parameters to the first feature information to obtain second feature information, and performing uniform block processing on the second feature information. For example, the second feature information is evenly divided into k subspace feature blocks. For example, the i-th subspace feature block is where represents the interval from to For the i-th subspace feature block (that is, the subspace feature block in the above multiple subspace feature blocks), calculating the subspace feature block and the multiple low-rank decomposition matrices to obtain the corresponding response sub-feature. For example, the response sub-feature corresponding to the i-th subspace feature block is
[0090] After obtaining multiple response sub-features, perform splicing processing on the multiple response sub-features to obtain the first response feature. For example, the first response feature is Δy = Concat(o , …, o 1 , …, o k ).
[0091] By decomposing the input features into multiple subspaces and using low-rank matrices for parameter fine-tuning in each subspace, the model can specifically capture and process diverse semantic patterns in the input.
[0092] In an alternative embodiment, as Figure 3 shown in the schematic diagram of the feature processing layer of the content generation model, it should be noted that there are multiple feature processing layers in the content generation model, and the processing processes of the feature processing layers in the multiple feature processing layers are the same. As Figure 3 shown, the input feature is the output result of the previous feature processing layer. If the current feature processing layer is the last layer, then Figure 3The output feature in it is the combination of the above-mentioned first reply feature and second reply feature. The specific processing process includes: Step 1: Input feature decomposition. For the input feature (the above-mentioned first feature) \(x\in R\) m , it is evenly divided into \(k\) subspace feature blocks, and learnable position parameters are added to the subspace feature blocks. For example, the \(i\)-th subspace feature block is
[0093] Step 2: Calculate dynamic weights. Decompose the low-rank matrix of the content generation model. For the \(i\)-th subspace feature block, the low-rank matrix is decomposed into For \(\Delta W\) i calculate its weight \(\alpha\) i =\(\text{softmax}(W\) r f i )\(\in R\) r .
[0094] Step 3: Dynamically aggregate weights. Re-weight the low-rank matrix with weights, that is to obtain the dynamic weight related to the input.
[0095] Step 4: Subspace output splicing. Calculate the output results of the subspaces using the dynamic weights Perform splicing processing on multiple output results to obtain \(\Delta y=\text{Concat}(o\) 1 ,\(\cdots\),\(o\) k ).
[0096] Step 5: Calculate the original output result and obtain the fine-tuned result by addition. The pre-trained weight matrix of the content generation model is \(W\) 0 , and the corresponding output result is \(y\) 0 =\(W\) 0 x. Then the final output feature is \(y = y\) 0 +\(\Delta y\).
[0097] It should be noted that from the perspective of function approximation, it can be proved that the method provided in this application can produce a smaller approximation error compared with LoRA in the prior art.
[0098] The output of LoRA in the prior art is \(Y\) 1 =\((W + BA)X\), where \(X\) is the input, \(W\) is the pre-trained weight, and \(BA\) is the low-rank matrix; the output of the method provided in this application is \(Y\) 2 =\(W + g(X)BAX\), where \(g(x)\) is the dynamic weight.
[0099] Define the residual term: For the expected target, the mean squared error (MSE) minimized by the two methods is respectively expressed as: To obtain a better dynamic weight function, take the derivative of E2 with respect to g(X) and set it to zero: A better dynamic weight function can be obtained:
[0100] When g = g * the dynamic weight error of this application satisfies:
[0101] Therefore, the method provided in this application can produce a smaller approximation error compared with LoRA in the prior art.
[0102] In the training method of the content generation model provided in the first embodiment of this application, by obtaining a target sample set and processing the knowledge samples in the target sample set through an initial content generation model, first feature information is obtained, where the initial content generation model is a pre-trained content generation model; calculating the weights of the low-rank matrix of the initial content generation model according to the first feature information to obtain first weights; obtaining predicted response information according to the first feature information, the first weights and the low-rank matrix of the initial content generation model; and training the initial content generation model according to the predicted response information to obtain a target content generation model, which solves the technical problem that the weights of the low-rank matrix of the content generation model in the related art are fixed, resulting in relatively low accuracy of the output results of the content generation model.
[0103] In this solution, obtaining a target sample set and training the pre-trained initial content generation model through the target sample set can ensure that the model focuses on learning feature information related to subsequent tasks during the fine-tuning stage, thereby improving the performance of the content generation model in specific tasks. When training the initial content generation model through the target sample set, calculating the weights of the low-rank matrix of the initial content generation model using the first feature information corresponding to the knowledge samples enables the model to dynamically adjust the weights of the low-rank matrix of the content generation model according to different input features. Obtaining predicted response information according to the first feature information, the first weights and the low-rank matrix of the initial content generation model, and finally training the initial content generation model with the predicted response information enables the target content generation model to adapt to different tasks, thereby achieving the technical effect of improving the accuracy of the output results of the content generation model.
[0104] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0105] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation method. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to enable a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of various embodiments of the present application.
[0106] Embodiment 2
[0107] The present application provides a method for processing data as follows Figure 4 shown. Figure 4 is a flowchart of the data processing method according to the second embodiment of the present application. The processing method includes:
[0108] Step S401, obtaining problem information input by a target object;
[0109] Step S402, processing the problem information through a target content generation model to obtain target reply information, where the target content generation model is trained by using the content generation model training method provided in Embodiment 1;
[0110] Step S403, returning the target reply information to the target object.
[0111] Optionally, after training the target content generation model by using the content generation model training method provided in Embodiment 1, the target content generation model can be deployed to a platform that needs to be applied. Then, through the platform, obtain problem information input by the target object. The problem information can be in text form, such as a sentence asking about specific product information, or other forms of data, such as images, audio, etc.
[0112] Then, the problem information is input into the target content generation model for processing. It should be noted that the target content generation model is trained by using the content generation model training method provided in Embodiment 1. The target content generation model will generate a target reply information according to the input problem information by using its internal parameters (including the low-rank matrix obtained by fine-tuning). Finally, the target reply information is returned to the target object.
[0113] In the data processing method provided in the second embodiment of the present application, processing the question information through the target content generation model to obtain the target reply information includes: processing the question information through the target content generation model to obtain target feature information; calculating the weight of the low-rank matrix of the target content generation model according to the target feature information to obtain the second weight; and obtaining the target reply information according to the target feature information, the second weight, and the low-rank matrix of the target content generation model.
[0114] Optionally, processing the question information through the target content generation model includes: vectorizing the question information to obtain target feature information, evenly dividing the target feature information into multiple target subspace features, and decomposing the low-rank matrix to obtain multiple low-rank decomposition matrices, calculating the weights of the multiple low-rank decomposition matrices through the multiple target subspace features to obtain the above-mentioned second weight. Weighting the multiple low-rank decomposition matrices through the second weight to obtain an adjusted low-rank matrix, and finally, processing according to the target feature information, the adjusted low-rank matrix, and the pre-trained weight matrix to obtain the target reply information.
[0115] In summary, when training the initial content generation model through the target sample set, calculating the weight of the low-rank matrix of the initial content generation model using the first feature information corresponding to the knowledge sample enables the model to dynamically adjust the weight of the low-rank matrix of the content generation model according to different input features, obtaining the predicted reply information through the first feature information, the first weight, and the low-rank matrix of the initial content generation model, and finally training the initial content generation model with the predicted reply information, enabling the target content generation model to adapt to different tasks, thereby achieving the technical effect of improving the accuracy of the output result of the content generation model.
[0116] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0117] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal device (which can be a mobile phone, computer, server, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0118] Embodiment 3
[0119] The present application provides a method for processing data as follows Figure 5 shown. Figure 5 is a flowchart of the method for processing data according to Embodiment 3 of the present application. The processing method includes:
[0120] Step S501, obtaining the problem information input by the client;
[0121] Step S502, processing the problem information through a target content generation model in the cloud server to obtain target reply information, where the target content generation model is trained by using the training method of the content generation model provided in Embodiment 1;
[0122] Step S503, returning the target reply information to the client.
[0123] It should be noted that the specific method for processing data in the cloud server is the same as that in Embodiment 2 and will not be elaborated here.
[0124] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.
[0125] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases, the former is a better implementation manner. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions for causing a terminal device (which can be a mobile phone, a computer, a server, or a network device, etc.) to execute the methods of the various embodiments of the present application.
[0126] Embodiment 4
[0127] According to an embodiment of the present application, there is also provided a training device for a content generation model for implementing the training method of the above content generation model, as Figure 6 shown. The device includes: a first acquisition unit 601, a calculation unit 602, a first processing unit 603, and a training unit 604.
[0128] The first acquisition unit 601 is configured to acquire a target sample set, and process the knowledge samples in the target sample set through an initial content generation model to obtain first feature information, where the initial content generation model is a pre-trained content generation model;
[0129] The calculation unit 602 is configured to calculate the weights of the low-rank matrix of the initial content generation model according to the first feature information to obtain first weights;
[0130] The first processing unit 603 is configured to obtain predicted reply information according to the first feature information, the first weights, and the low-rank matrix of the initial content generation model;
[0131] The training unit 604 is configured to train the initial content generation model according to the predicted reply information to obtain a target content generation model.
[0132] In the training device of the content generation model provided in the fourth embodiment of the present application, a target sample set is obtained through the first acquisition unit 601, and the knowledge samples in the target sample set are processed by the initial content generation model to obtain first feature information, where the initial content generation model is a pre-trained content generation model; the calculation unit 602 calculates the weights of the low-rank matrix of the initial content generation model based on the first feature information to obtain first weights; the first processing unit 603 obtains predicted response information based on the first feature information, the first weights, and the low-rank matrix of the initial content generation model; the training unit 604 trains the initial content generation model based on the predicted response information to obtain a target content generation model, solving the technical problem that the weights of the low-rank matrix of the content generation model in the related art are fixed, resulting in a relatively low accuracy of the output result of the content generation model.
[0133] In this solution, obtaining a target sample set and training the pre-trained initial content generation model with the target sample set can ensure that the model focuses on learning feature information related to subsequent tasks during the fine-tuning stage, thereby improving the performance of the content generation model in specific tasks. When training the initial content generation model with the target sample set, the weights of the low-rank matrix of the initial content generation model are calculated using the first feature information corresponding to the knowledge samples, enabling the model to dynamically adjust the weights of the low-rank matrix of the content generation model according to different input features. Predicted response information is obtained based on the first feature information, the first weights, and the low-rank matrix of the initial content generation model. Finally, the initial content generation model is trained with the predicted response information, enabling the target content generation model to adapt to different tasks, and thus achieving the technical effect of improving the accuracy of the output result of the content generation model.
[0134] Optionally, in the training device of the content generation model provided in the fourth embodiment of the present application, the calculation unit includes: an addition subunit, configured to add position parameters to the first feature information to obtain second feature information, where the position parameters are the position information corresponding to the feature information in the first feature information, and the position parameters are learnable parameters; a first processing subunit, configured to perform block processing on the second feature information to obtain a plurality of subspace feature blocks; a first calculation subunit, configured to calculate the weights of the low-rank matrix of the initial content generation model based on the plurality of subspace feature blocks to obtain first weights.
[0135] Optionally, in the training device of the content generation model provided in Embodiment 4 of the present application, the first calculation subunit includes: an acquisition module, configured to acquire the rank corresponding to the low-rank matrix of the initial content generation model, and decompose the low-rank matrix of the initial content generation model according to the rank to obtain a plurality of low-rank decomposition matrices; a first calculation module, configured to calculate weights of the low-rank decomposition matrices in the plurality of low-rank decomposition matrices through a target subspace feature block for the target subspace feature block, where the target subspace feature block is any one of the plurality of subspace feature blocks; a determination module, configured to obtain a first weight according to the plurality of weight values.
[0136] Optionally, in the training device of the content generation model provided in Embodiment 4 of the present application, the first calculation module includes: an acquisition sub-module, configured to acquire a target learnable parameter; a calculation sub-module, configured to calculate weights of the low-rank decomposition matrices in the plurality of low-rank decomposition matrices according to the target learnable parameter and the target subspace feature block to obtain a plurality of initial weight values; a processing sub-module, configured to process the plurality of initial weight values through an activation function to obtain a plurality of weight values.
[0137] Optionally, in the training device of the content generation model provided in Embodiment 4 of the present application, the first processing unit includes: a second calculation subunit, configured to perform weighted calculation according to the first weight and the low-rank matrix of the initial content generation model to obtain a processed low-rank matrix; a third calculation subunit, configured to calculate a first reply feature according to the first feature information and the processed low-rank matrix; a fourth calculation subunit, configured to calculate a second reply feature according to the first feature information and the pre-trained weight matrix of the initial content generation model; a second processing subunit, configured to obtain predicted reply information according to the first reply feature and the second reply feature.
[0138] Optionally, in the training device of the content generation model provided in Embodiment 4 of the present application, the second calculation subunit includes: a decomposition module, configured to decompose the low-rank matrix of the initial content generation model according to the rank of the low-rank matrix of the initial content generation model to obtain a plurality of low-rank decomposition matrices; a calculation module, configured to calculate according to the low-rank decomposition matrices in the plurality of low-rank decomposition matrices and the first weight to obtain a plurality of processed low-rank decomposition matrices; a second calculation module, configured to obtain a processed low-rank matrix according to the plurality of processed low-rank decomposition matrices.
[0139] Optionally, in the training device of the content generation model provided in the fourth embodiment of the present application, the fourth calculation subunit includes: a processing module, configured to process the first feature information to obtain a plurality of subspace feature blocks; a third calculation module, configured to calculate according to the low-rank decomposition matrix in the plurality of low-rank decomposition matrices and the subspace feature blocks in the plurality of subspace feature blocks to obtain a plurality of reply sub-features; a splicing module, configured to splice and process the plurality of reply sub-features to obtain a first reply feature.
[0140] It should be noted here that the above-mentioned first acquisition unit 601, calculation unit 602, first processing unit 603, and training unit 604 correspond to steps S201 to S204 in Embodiment 1. The functions of the four units are the same as those of the corresponding steps in terms of implementation examples and application scenarios, but are not limited to the content disclosed in the above-mentioned Embodiment 1. It should be noted that the above-mentioned modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0141] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as the schemes, application scenarios, and implementation processes provided in Embodiment 1, but are not limited to the schemes provided in Embodiment 1.
[0142] Embodiment 5
[0143] According to an embodiment of the present application, there is also provided a data processing device for implementing the above data processing method, as Figure 7 shown, the device includes: a second acquisition unit 701, a second processing unit 702, and a return unit 703.
[0144] The second acquisition unit 701 is configured to acquire question information input by a target object;
[0145] The second processing unit 702 is configured to process the question information through a target content generation model to obtain target reply information, where the target content generation model is trained by using the training method of the content generation model provided in Embodiment 1;
[0146] The return unit 703 is configured to return the target reply information to the target object.
[0147] Optionally, in the data processing device provided in the fourth embodiment of the present application, the second processing unit includes: a third processing subunit, configured to process the question information through a target content generation model to obtain target feature information; a fifth calculation subunit calculates the weight of the low-rank matrix of the target content generation model according to the target feature information to obtain a second weight; a fourth processing subunit is configured to obtain target reply information according to the target feature information, the second weight, and the low-rank matrix of the target content generation model.
[0148] It should be noted that the above-mentioned second acquisition unit 701, second processing unit 702, and return unit 703 correspond to steps S401 to S403 in Embodiment 2. The functions of the three units are the same as those of the corresponding steps in terms of implementation examples and application scenarios, but are not limited to the content disclosed in the above-mentioned Embodiment 2. It should be noted that the above modules, as part of the device, can run in the computer terminal 10 provided in Embodiment 1.
[0149] It should be noted that the preferred implementation schemes involved in the above embodiments of the present application are the same as those provided in Embodiment 2 in terms of application scenarios and implementation processes, but are not limited to the solutions provided in Embodiment 2.
[0150] Embodiment 6
[0151] An embodiment of the present application can provide an electronic device, which can be any one of the electronic device terminals in the electronic device terminal group. Optionally, in this embodiment, the above electronic device can also be replaced with a terminal device such as a mobile terminal.
[0152] Optionally, in this embodiment, the above electronic device can be at least one of the multiple network devices in the computer network.
[0153] In this embodiment, the above electronic device can execute the program code of the following steps in the above method: obtaining a target sample set, processing the knowledge samples in the target sample set through an initial content generation model to obtain first feature information, where the initial content generation model is a pre-trained content generation model; calculating the weights of the low-rank matrix of the initial content generation model according to the first feature information to obtain first weights; obtaining predicted response information according to the first feature information, the first weights, and the low-rank matrix of the initial content generation model; and training the initial content generation model according to the predicted response information to obtain a target content generation model.
[0154] The above electronic device can execute the program code of the following steps in the above method: calculating the weights of the low-rank matrix of the initial content generation model according to the first feature information to obtain first weights, including: adding position parameters to the first feature information to obtain second feature information, where the position parameters are the position information corresponding to the feature information in the first feature information, and the position parameters are learnable parameters; performing block processing on the second feature information to obtain multiple subspace feature blocks; and calculating the weights of the low-rank matrix of the initial content generation model according to the multiple subspace feature blocks to obtain first weights.
[0155] The above electronic device can execute the program code of the following steps in the above method: calculating the weights of the low-rank matrix of the initial content generation model based on multiple subspace feature blocks, and obtaining the first weights, including: obtaining the rank corresponding to the low-rank matrix of the initial content generation model, and decomposing the low-rank matrix of the initial content generation model according to the rank to obtain multiple low-rank decomposition matrices; for the target subspace feature block, calculating the weights of the low-rank decomposition matrices in the multiple low-rank decomposition matrices through the target subspace feature block to obtain multiple weight values, where the target subspace feature block is any one of the multiple subspace feature blocks; obtaining the first weights based on the multiple weight values.
[0156] The above electronic device can execute the program code of the following steps in the above method: calculating the weights of the low-rank decomposition matrices in the multiple low-rank decomposition matrices through the target subspace feature block to obtain multiple weight values, including: obtaining the target learnable parameters; calculating the weights of the low-rank decomposition matrices in the multiple low-rank decomposition matrices according to the target learnable parameters and the target subspace feature block to obtain multiple initial weight values; processing the multiple initial weight values through an activation function to obtain multiple weight values.
[0157] The above electronic device can execute the program code of the following steps in the above method: obtaining the predicted response information based on the first feature information, the first weights, and the low-rank matrix of the initial content generation model, including: performing weighted calculation based on the first weights and the low-rank matrix of the initial content generation model to obtain the processed low-rank matrix; calculating based on the first feature information and the processed low-rank matrix to obtain the first response feature; calculating based on the first feature information and the pre-trained weight matrix of the initial content generation model to obtain the second response feature; obtaining the predicted response information based on the first response feature and the second response feature.
[0158] The above electronic device can execute the program code of the following steps in the above method: performing weighted calculation based on the first weights and the low-rank matrix of the initial content generation model to obtain the processed low-rank matrix, including: decomposing the low-rank matrix of the initial content generation model according to the rank of the low-rank matrix of the initial content generation model to obtain multiple low-rank decomposition matrices; calculating according to the low-rank decomposition matrices in the multiple low-rank decomposition matrices and the first weights to obtain multiple processed low-rank decomposition matrices; obtaining the processed low-rank matrix based on the multiple processed low-rank decomposition matrices.
[0159] The above electronic device may execute the program code of the following steps in the above method: calculate based on the first feature information and the processed low-rank matrix, and obtain a first response feature, including: process the first feature information to obtain a plurality of subspace feature blocks; calculate based on the low-rank decomposition matrix in the plurality of low-rank decomposition matrices and the subspace feature blocks in the plurality of subspace feature blocks to obtain a plurality of response sub-features; perform a splicing process on the plurality of response sub-features to obtain the first response feature.
[0160] The above electronic device may execute the program code of the following steps in the above method: obtain the question information input by the target object; process the question information through the target content generation model to obtain the target response information, where the target content generation model is trained by using the training method of the content generation model in any one of the above; return the target response information to the target object.
[0161] The above electronic device may execute the program code of the following steps in the above method: process the question information through the target content generation model to obtain the target response information, including: process the question information through the target content generation model to obtain the target feature information; calculate the weight of the low-rank matrix of the target content generation model based on the target feature information to obtain a second weight; obtain the target response information based on the target feature information, the second weight, and the low-rank matrix of the target content generation model.
[0162] The above electronic device may execute the program code of the following steps in the above method: obtain the question information input by the client; process the question information through the target content generation model in the cloud server to obtain the target response information, where the target content generation model is trained by using the training method of the content generation model in any one of the above; return the target response information to the client.
[0163] Optionally, Figure 8 is a structural block diagram of an electronic device according to an embodiment of the present application. As Figure 8 shown, the electronic device 80 may include: one or more ( Figure 8 only one is shown in the figure) processors 802, a memory 804. The electronic device 80 may further include a storage controller to control and manage the memory 804 through the storage controller; the electronic device 80 may further include a peripheral interface to connect a radio frequency module, an audio module, a display screen, etc. through the peripheral interface.
[0164] Among them, the memory can be used to store software programs and modules, such as the program instructions / modules corresponding to the training method of the content generation model and the data processing method and device in the embodiments of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory, that is, implements the above-mentioned training method of the content generation model and the data processing method. The memory may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory may further include a memory remotely provided with respect to the processor, and these remote memories can be connected to the electronic device 80 through a network. Examples of the above network include but are not limited to the Internet, enterprise intranets, local area networks, mobile communication networks, and combinations thereof.
[0165] The processor can call the information and application programs stored in the memory through the transmission device to execute the following steps: obtaining a target sample set, processing the knowledge samples in the target sample set through an initial content generation model to obtain first feature information, where the initial content generation model is a pre-trained content generation model; calculating the weights of the low-rank matrix of the initial content generation model according to the first feature information to obtain first weights; obtaining predicted response information according to the first feature information, the first weights, and the low-rank matrix of the initial content generation model; training the initial content generation model according to the predicted response information to obtain a target content generation model.
[0166] Optionally, the above processor can also execute the program code of the following steps: calculating the weights of the low-rank matrix of the initial content generation model according to the first feature information to obtain first weights includes: adding position parameters to the first feature information to obtain second feature information, where the position parameters are the position information corresponding to the feature information in the first feature information, and the position parameters are learnable parameters; performing block processing on the second feature information to obtain a plurality of subspace feature blocks; calculating the weights of the low-rank matrix of the initial content generation model according to the plurality of subspace feature blocks to obtain first weights.
[0167] Optionally, the above processor can also execute the program code of the following steps: calculating the weights of the low-rank matrix of the initial content generation model according to the plurality of subspace feature blocks to obtain first weights includes: obtaining the rank corresponding to the low-rank matrix of the initial content generation model, and decomposing the low-rank matrix of the initial content generation model according to the rank to obtain a plurality of low-rank decomposition matrices; for a target subspace feature block, calculating the weights of the low-rank decomposition matrices in the plurality of low-rank decomposition matrices through the target subspace feature block to obtain a plurality of weight values, where the target subspace feature block is any one of the plurality of subspace feature blocks; obtaining first weights according to the plurality of weight values.
[0168] Optionally, the above-mentioned processor can also execute the program code of the following steps: calculate the weights of the low-rank decomposition matrices in multiple low-rank decomposition matrices through the target subspace feature block, and obtain multiple weight values, including: obtaining the target learnable parameters; calculating the weights of the low-rank decomposition matrices in multiple low-rank decomposition matrices according to the target learnable parameters and the target subspace feature block to obtain multiple initial weight values; processing the multiple initial weight values through an activation function to obtain multiple weight values.
[0169] Optionally, the above-mentioned processor can also execute the program code of the following steps: generate a low-rank matrix of the model according to the first feature information, the first weight, and the initial content, and obtain the predicted reply information, including: performing weighted calculation on the low-rank matrix generated by the model according to the first weight and the initial content to obtain the processed low-rank matrix; calculating according to the first feature information and the processed low-rank matrix to obtain the first reply feature; calculating according to the pre-trained weight matrix of the model generated by the first feature information and the initial content to obtain the second reply feature; obtaining the predicted reply information according to the first reply feature and the second reply feature.
[0170] Optionally, the above-mentioned processor can also execute the program code of the following steps: perform weighted calculation on the low-rank matrix generated by the model according to the first weight and the initial content to obtain the processed low-rank matrix, including: decomposing the low-rank matrix generated by the model according to the initial content according to the rank of the low-rank matrix generated by the model according to the initial content to obtain multiple low-rank decomposition matrices; calculating according to the low-rank decomposition matrices in the multiple low-rank decomposition matrices and the first weight to obtain multiple processed low-rank decomposition matrices; obtaining the processed low-rank matrix according to the multiple processed low-rank decomposition matrices.
[0171] Optionally, the above-mentioned processor can also execute the program code of the following steps: calculate according to the first feature information and the processed low-rank matrix to obtain the first reply feature, including: processing the first feature information to obtain multiple subspace feature blocks; calculating according to the low-rank decomposition matrices in the multiple low-rank decomposition matrices and the subspace feature blocks in the multiple subspace feature blocks to obtain multiple reply sub-features; performing splicing processing on the multiple reply sub-features to obtain the first reply feature.
[0172] Optionally, the above-mentioned processor can also execute the program code of the following steps: obtain the question information input by the target object; process the question information through the target content generation model to obtain the target reply information, where the target content generation model is trained by using the training method of the content generation model in any one of the above; return the target reply information to the target object.
[0173] Optionally, the above processor may also execute the program code of the following steps: processing the question information through the target content generation model to obtain the target response information, including: processing the question information through the target content generation model to obtain the target feature information; calculating the weights of the low-rank matrix of the target content generation model according to the target feature information to obtain the second weights; and obtaining the target response information according to the target feature information, the second weights and the low-rank matrix of the target content generation model.
[0174] Optionally, the above processor may also execute the program code of the following steps: obtaining the question information input by the client; processing the question information through the target content generation model in the cloud server to obtain the target response information, where the target content generation model is trained by using the training method of the content generation model of any one of the above; and returning the target response information to the client.
[0175] Those of ordinary skill in the art can understand that Figure 8 The structure shown is only schematic, and the electronic device 80 may also be a terminal device such as a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a handheld computer, and a Mobile Internet Device (MID), a PAD, etc. Figure 8 It does not limit the structure of the above electronic device. For example, the electronic device 80 may further include more or fewer components (such as a network interface, a display device, etc.) than those shown in Figure 8 the figure, or have a different configuration from that shown in Figure 8 the figure.
[0176] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the relevant hardware of the terminal device through a program, and the program can be stored in a computer-readable storage medium, and the storage medium may include: a flash drive, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disc, etc.
[0177] Embodiment 7
[0178] The embodiment of the present application also provides a computer program product. Optionally, in this embodiment, the above computer program product may be used to store the program code executed by the training method of the content generation model and the data processing method provided in the first embodiment above.
[0179] Optionally, in this embodiment, the above computer program product may be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0180] The serial numbers of the embodiments of the present application above are only for description and do not represent the superiority or inferiority of the embodiments.
[0181] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0182] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0183] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0184] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0185] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, read-only memories (ROMs), random access memories (RAMs), mobile hard disks, magnetic disks or optical discs that can store program codes.
[0186] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A training method for a content generation model, characterized in that: include: Acquire a target sample set, and process the knowledge samples in the target sample set through an initial content generation model to obtain first feature information, wherein the initial content generation model is a pre-trained content generation model; Calculating the weight of the low-rank matrix of the initial content generation model according to the first feature information to obtain a first weight; Obtaining predicted reply information according to the first feature information, the first weight, and the low-rank matrix of the initial content generation model; The initial content generation model is trained according to the predicted reply information to obtain a target content generation model.
2. The method according to claim 1, characterized in that Calculating the weight of the low-rank matrix of the initial content generation model according to the first feature information to obtain the first weight includes: Adding a position parameter to the first feature information to obtain second feature information, wherein the position parameter is the position information corresponding to the feature information in the first feature information, and the position parameter is a learnable parameter; Performing block processing on the second feature information to obtain a plurality of subspace feature blocks; The weight of the low-rank matrix of the initial content generation model is calculated according to the multiple subspace feature blocks to obtain the first weight.
3. The method according to claim 2, characterized in that Calculating the weight of the low-rank matrix of the initial content generation model according to the multiple subspace feature blocks to obtain the first weight includes: Obtaining a rank corresponding to the low-rank matrix of the initial content generation model, and decomposing the low-rank matrix of the initial content generation model according to the rank to obtain a plurality of low-rank decomposition matrices; For a target subspace feature block, weight calculation is performed on a low-rank decomposition matrix in the multiple low-rank decomposition matrices through the target subspace feature block to obtain multiple weight values, wherein the target subspace feature block is any one of the multiple subspace feature blocks; The first weight is obtained according to the multiple weight values.
4. The method according to claim 3, characterized in that The weight calculation is performed on the low-rank decomposition matrices in the multiple low-rank decomposition matrices by using the target subspace feature block, and the multiple weight values obtained include: Get the target learnable parameters; According to the target learnable parameter and the target subspace feature block, weight calculation is performed on the low-rank decomposition matrices in the multiple low-rank decomposition matrices to obtain multiple initial weight values; The multiple initial weight values are processed by an activation function to obtain the multiple weight values.
5. The method according to claim 1, characterized in that The predicted response information is obtained according to the first feature information, the first weight, and the low-rank matrix of the initial content generation model, including: Performing weighted calculation according to the first weight and the low-rank matrix of the initial content generation model to obtain a processed low-rank matrix; Performing calculation based on the first feature information and the processed low-rank matrix to obtain a first response feature; Calculating according to the first feature information and the pre-trained weight matrix of the initial content generation model to obtain a second answer feature; The predicted reply information is obtained according to the first reply feature and the second reply feature.
6. The method according to claim 5, characterized in that The low-rank matrix obtained by performing weighted calculation according to the first weight and the low-rank matrix of the initial content generation model includes: Decomposing the low-rank matrix of the initial content generation model according to the rank of the low-rank matrix of the initial content generation model to obtain multiple low-rank decomposition matrices; Calculating according to a low-rank decomposition matrix among the multiple low-rank decomposition matrices and the first weight to obtain a plurality of processed low-rank decomposition matrices; The processed low-rank matrix is obtained according to the multiple processed low-rank decomposition matrices.
7. The method according to claim 6, characterized in that The first response feature obtained by calculating according to the first feature information and the processed low-rank matrix includes: Processing the first feature information to obtain a plurality of subspace feature blocks; Calculating according to the low-rank decomposition matrices in the multiple low-rank decomposition matrices and the subspace feature blocks in the multiple subspace feature blocks to obtain multiple response sub-features; The multiple reply sub-features are concatenated to obtain the first reply feature.
8. A data processing method, characterized in that: include: Get the question information input by the target object; Processing the question information through a target content generation model to obtain target answer information, wherein the target content generation model is trained using the training method for a content generation model described in any one of claims 1 to 7; The target reply information is returned to the target object.
9. The method according to claim 8, characterized in that The question information is processed by the target content generation model to obtain target answer information including: Processing the question information through the target content generation model to obtain target feature information; Calculating the weight of the low-rank matrix of the target content generation model according to the target feature information to obtain a second weight; The target response information is obtained based on the target feature information, the second weight and the low-rank matrix of the target content generation model.
10. A data processing method, characterized in that: include: Get the question information entered by the client; The question information is processed in the cloud server by a target content generation model to obtain target answer information, wherein the target content generation model is trained by the content generation model training method described in any one of claims 1 to 7; The target reply information is returned to the client.
11. A training device for a content generation model, characterized in that: include: A first acquisition unit is used to acquire a target sample set, and process the knowledge samples in the target sample set through an initial content generation model to obtain first feature information, wherein the initial content generation model is a pre-trained content generation model; a calculation unit, configured to calculate a weight of a low-rank matrix of the initial content generation model according to the first feature information to obtain a first weight; A first processing unit, configured to obtain predicted answer information according to the first feature information, the first weight and the low-rank matrix of the initial content generation model; A training unit is used to train the initial content generation model according to the predicted reply information to obtain a target content generation model.
12. A data processing device, characterized in that: include: A second acquisition unit is used to acquire question information input by a target object; A second processing unit, configured to process the question information through a target content generation model to obtain target answer information, wherein the target content generation model is trained using the training method for the content generation model described in any one of claims 1 to 7; A returning unit is used to return the target reply information to the target object.
13. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is run, the device where the storage medium is located is controlled to execute the content generation model training method described in any one of claims 1 to 7, or the data processing method described in any one of claims 8 to 10.
14. An electronic device, characterized in that: include: A memory storing an executable program; A processor, used to run the program, wherein when the program is run, the training method of the content generation model described in any one of claims 1 to 7, or the data processing method described in any one of claims 8 to 10 is executed.
15. A computer program product, characterized in that It comprises a computer program or instructions, which, when executed by a processor, implements the training method of the content generation model according to any one of claims 1 to 7, or the data processing method according to any one of claims 8 to 10.