Task processing method, device and equipment based on dynamic LoRA network and medium
By calculating the correlation between the feature vectors of the task processing text and the LoRA network, dynamically selecting and fusing the target LoRA network to the basic model, solving the problem of high task adaptability and resource utilization of large language models, and achieving efficient and accurate task processing.
Patent Information
- Application Number
- CN202510569119.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-08-12
AI Technical Summary
The existing large language models have limitations in model fine-tuning and task adaptability, which are difficult to adjust dynamically, and are inefficient, have high resource usage and slow response speed in multimodal tasks and edge computing scenarios.
By receiving the feature vectors of the task processing text, the correlation with several LoRA networks is calculated, the activation probability is determined based on the pre-configured gated network, and the target LoRA network is dynamically fused into the basic model, and the low-rank decomposition technology is used to reduce redundant computing and storage.
It improves the efficiency and accuracy of task processing, reduces computing resource consumption, and significantly reduces memory usage and computing time when processing large-scale language model, providing a flexible and efficient solution.
Smart Images

Figure CN120469778A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of large model applications, and in particular to a task processing method, device, computer equipment and computer storage medium based on a dynamic LoRA network. Background Art
[0002] In recent years, with the rapid development of natural language processing (NLP) and artificial intelligence technologies, large language models (LLMs) have demonstrated strong performance across a wide range of tasks. However, traditional technologies have limitations in model fine-tuning and task adaptability. For example, while LoRA (Low-Rank Adaptation) technology can reduce the number of parameters, its adaptability is limited by its static loading method, making dynamic adjustment difficult. The MoE (Mixture of Experts) architecture has high memory overhead and is difficult to flexibly adjust to specific task requirements. Furthermore, existing technologies also face challenges such as low efficiency, high resource utilization, and slow response times when handling multimodal tasks and edge computing scenarios. Improving the efficiency and accuracy of large language models in task processing has become a pressing issue. Summary of the Invention
[0003] The purpose of this application is to provide a task processing method, device, computer equipment and computer storage medium based on a dynamic LoRA network, so as to at least solve the problem of low efficiency and accuracy of task processing by the current large language model.
[0004] To solve the above technical problems, the present application provides a task processing method based on a dynamic LoRA network, comprising: receiving a task processing text, and identifying a feature vector of the task processing text, wherein the feature vector at least indicates semantic and contextual features of the task processing text; Calculate the correlation with several candidate LoRA networks based on the feature vector of the task-processed text; Determine the activation probability of each candidate LoRA network according to the correlation, and determine the target LoRA network to be activated according to the activation probability based on the preconfigured gating network; The target LoRA network is integrated into the basic model to obtain a fusion model; The fusion model is used to perform task processing based on the semantic and contextual features of the task processing text, so as to load the target LoRA network and output the processing result of the task processing text.
[0005] Optionally, after receiving the task processing text and identifying the feature vector of the task processing text, the method further includes: Processing text matching domain or label information for the task; The feature vector is updated after weighting the key information in the special effect vector according to the domain or label information based on a preconfigured attention mechanism.
[0006] Optionally, the preconfigured gated network-based determination of a target LoRA network to be activated according to the activation probability includes: Get the scene of the current task processing; Configuring the gating network according to the scenario to adjust the probabilistic activation strategy of the gating network; The target LoRA network to be activated is determined based on the adjusted door network probability activation strategy.
[0007] Optionally, the selected LoRA network is stored in a distributed parameter storage system, and the target LoRA network is integrated into the basic model to obtain a fusion model, including: Retrieve the target LoRA network from the distributed parameter storage system based on the product quantization indexing mechanism; The target LoRA network is fused into the basic model through low-rank decomposition to obtain a fusion model.
[0008] Optionally, the target LoRA network includes at least two of a first LoRA network and a second LoRA network, and fusing the target LoRA network into a basic model by low-rank decomposition to obtain a fusion model includes: Obtaining a first correlation of the first LoRA network and a second correlation of the second LoRA network; configuring a first weight of a first LoRA network according to the first correlation and configuring a second weight of a second LoRA network according to the second correlation; By low-rank decomposition, the first LoRA network and the second LoRA network are fused into a basic model according to the first weight and the second weight to obtain a fusion model.
[0009] Optionally, when the feature vector of the task processing text represents that the current task is a multimodal scenario, the gating network is configured as a cross-modal gating network, and the method further includes: The target LoRA network of the corresponding modality is activated according to the cross-modal gating network to realize joint routing decision through the cross-modal gating network and output the multimodal collaborative processing result of the task processing text.
[0010] Optionally, when the request for the task processing text comes from an edge device, after the pre-configured gated network determines the target LoRA network to be activated according to the activation probability, the method further includes: In response to the request of the task processing text, the target LoRA network is sent to the edge device, so that the edge device downloads and caches the target LoRA network.
[0011] To solve the above technical problems, the present application also provides a task processing device based on a dynamic LoRA network, comprising: A task receiving module, configured to receive a task processing text and identify a feature vector of the task processing text, wherein the feature vector at least indicates semantic and contextual features of the task processing text; A correlation calculation module is used to calculate the correlation with several candidate LoRA networks based on the feature vector of the task processing text; A matching activation module is used to determine the activation probability of each candidate LoRA network according to the correlation, and determine the target LoRA network to be activated according to the activation probability based on the preconfigured gating network; A network fusion module is used to fuse the target LoRA network into the basic model to obtain a fusion model; A processing output module is used to perform task processing based on the semantic and contextual features of the task processing text using a fusion model to output the processing result of the task processing text after loading the target LoRA network.
[0012] In order to solve the above technical problems, the present application also provides a computer device, including a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor performs the steps of the above-mentioned task processing method based on the dynamic LoRA network.
[0013] In order to solve the above technical problems, the present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the above-mentioned task processing method based on the dynamic LoRA network.
[0014] The beneficial effects of the embodiments created by the present application are as follows: receiving a task processing text, identifying a feature vector of the task processing text, wherein the feature vector at least indicates the semantic and contextual features of the task processing text; calculating the correlation with a number of candidate LoRA networks based on the feature vector of the task processing text; determining the activation probability of each candidate LoRA network based on the correlation, and determining the target LoRA network to be activated based on the activation probability based on the preconfigured gating network; fusing the target LoRA network into the basic model to obtain a fusion model; using the fusion model to perform task processing based on the semantic and contextual features of the task processing text, so as to load the target LoRA network and then output the result. The processing result of the task processing text is output; by dynamically loading the target LoRA network and adopting low-rank decomposition technology, low-rank parameters related to the task are loaded only when needed, avoiding redundant calculation and storage of the entire model. This dynamic adjustment method significantly reduces the consumption of computing resources, especially when processing large-scale language models, it can effectively reduce video memory usage and computing time, and the fusion model can efficiently utilize the knowledge of the target LoRA network to perform accurate task processing on the task processing text and output high-quality results. This method of dynamically selecting and fusing LoRA networks not only improves the efficiency and performance of task processing, but also provides a flexible and efficient solution for natural language processing tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The above and / or additional aspects and advantages of the present application will become apparent and easily understood from the following description of the embodiments in conjunction with the accompanying drawings, in which: Figure 1 This is a basic flow chart of a task processing method based on a dynamic LoRA network according to a specific embodiment of the present application; Figure 2 This is a schematic diagram of the basic structure of a task processing device based on a dynamic LoRA network according to a specific embodiment of the present application; Figure 3 This is a basic structural block diagram of a computer device according to a specific embodiment of the present application. DETAILED DESCRIPTION
[0016] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.
[0017] Those skilled in the art will understand that, unless otherwise stated, the singular forms "a," "an," "said," and "the" used herein may also include plural forms. It should be further understood that the term "comprising" used in the specification of this application refers to the presence of the stated features, integers, steps, operations, elements, and / or components, but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0018] Those skilled in the art will understand that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and, unless specifically defined as such, will not be interpreted in an idealized or overly formal sense.
[0019] Those skilled in the art will appreciate that the term "terminal" as used herein includes both devices that are wireless signal receivers, i.e., devices that only have wireless signal receivers without transmission capabilities, and devices that have receiving and transmitting hardware capable of performing two-way communication over a two-way communication link. Such devices may include: cellular or other communication devices with single-line displays, multi-line displays, or cellular or other communication devices without multi-line displays; PCS (Personal Communications Service) devices that may combine voice, data processing, fax, and / or data communication capabilities; PDAs (Personal Digital Assistants) that may include a radio frequency receiver, a pager, Internet / Intranet access, a web browser, a notepad, a calendar, and / or a GPS (Global Positioning System) receiver; and conventional laptop and / or palmtop computers or other devices that have and / or include a radio frequency receiver. As used herein, a "terminal" can be portable, transportable, installed in a vehicle (air, sea, and / or land), or adapted and / or configured to operate locally and / or in a distributed manner at any other location on Earth and / or in space. A "terminal" as used herein can also refer to a communication terminal, an Internet access terminal, or a music / video playback terminal, such as a PDA, a mobile internet device (MID), and / or a mobile phone with music / video playback capabilities, as well as devices such as smart televisions and set-top boxes.
[0020] The hardware referred to by names such as "server", "client", and "service node" in this application is essentially an electronic device with capabilities equivalent to those of a personal computer. It is a hardware device that has the necessary components revealed by the von Neumann principle, such as a central processing unit (including an arithmetic unit and a controller), a memory, an input device, and an output device. Computer programs are stored in its memory, and the central processing unit loads the program stored in the external memory into the internal memory for execution, executes the instructions in the program, and interacts with the input and output devices to complete specific functions.
[0021] It should be noted that the concept of "server" referred to in this application can also be extended to server clusters. Based on the network deployment principles understood by those skilled in the art, the servers described should be logically divided. In physical space, these servers can be independent of each other but callable through interfaces, or integrated into a single physical computer or a computer cluster. Those skilled in the art should understand this flexibility and should not use it to constrain the implementation of the network deployment method of this application.
[0022] Unless expressly specified, one or more technical features of the present application can be deployed on a server for implementation and accessed by a client through a remote call to obtain an online service interface provided by the server, or can be directly deployed and run on a client for implementation.
[0023] Unless explicitly specified, the AI models referenced or may be referenced in this application can be deployed on a remote server and remotely called on the client, or can be deployed and directly called on a client with sufficient device capabilities. In some embodiments, when it runs on the client, its corresponding intelligence can be obtained through transfer learning to reduce the requirements for the client's hardware operating resources and avoid excessive occupation of the client's hardware operating resources.
[0024] Unless explicitly specified, the various data involved in this application can be stored remotely on a server or on a local terminal device, as long as they are suitable for being called by the technical solution of this application.
[0025] Those skilled in the art should be aware that although the various methods of this application are described based on the same concept and thus exhibit commonality, unless otherwise specified, these methods can be independently executed. Similarly, the various embodiments disclosed in this application are all based on the same inventive concept. Therefore, concepts with the same expression, as well as concepts that are appropriately transformed for convenience despite different expression, should be understood as equivalent.
[0026] Unless expressly stated to be mutually exclusive, the various embodiments disclosed in this application may be cross-combined with the relevant technical features of the various embodiments to flexibly construct new embodiments, as long as such combination does not deviate from the creative spirit of this application and can meet the needs of the prior art or resolve certain deficiencies in the prior art. Those skilled in the art should be aware of such flexibility.
[0027] See also Figure 1 , Figure 1 Schematic diagram of the basic flow of the task processing method based on the dynamic LoRA network in this embodiment.
[0028] like Figure 1 Shown, including: S1100, receiving a task processing text, and identifying a feature vector of the task processing text, wherein the feature vector at least indicates semantic and contextual features of the task processing text; This embodiment is applied to task processing scenarios in various fields such as finance, e-commerce, insurance, education, medical care, and law. In this embodiment, a task processing system equipped with a basic model (large language model) is configured to provide users with various services. The task processing system can be deployed on a server and accessed by the client through a remote call to obtain an online service interface provided by the server, or it can be directly deployed and run on the client to access it. The task processing system can also be deployed as a functional module on a web page or client. When a user needs to process a task, he or she can access the task processing system in a predetermined manner. In response to the user accessing the task processing system and sending a task, the task sent by the user is defined as a "task processing text" and the feature vector of the task processing text is identified. The reception of task processing text and the identification of feature vectors are the starting points of the entire dynamic LoRA network task processing method. Specifically, the task processing text is first received through an input interface, which can be an API, a file reader or a network transmission module, for receiving text data input by the user. The received task processing text is then preprocessed. The preprocessing includes text cleaning (removing useless symbols, stop words, etc.), word segmentation (splitting continuous text into words or subword units) and normalization (unifying text formats, such as lowercase). Preprocessing the task processing text can effectively reduce noise and improve the efficiency and accuracy of subsequent processing. Furthermore, the preprocessed task processing text undergoes feature extraction and encoding to obtain the feature vector of the task processing text. In one embodiment, in order to better capture the semantic and contextual features of the task processing text, a pre-trained language model based on the Transformer architecture (such as BERT or RoBERTa) is used to extract feature vectors, and the input text is passed through the encoder layer to generate a high-dimensional context feature vector, which can effectively represent the semantic and contextual information of the text. For example, the task processing text "I want to book a flight from Beijing to Shanghai on Friday" is encoded to output a 768-dimensional feature vector; further, the feature vector can be preliminarily classified to determine whether the task processing text belongs to the probability distribution of sentiment analysis, entity recognition, or intent classification.
[0029] It should be noted that the semantics of this embodiment may include detailed explanation of the task processing text, intent classification, key information extraction, semantic labeling, etc., to provide a comprehensive understanding and analysis of the meaning of the task processing text.
[0030] It should be noted that the tasks sent by the user in this embodiment can be sent in the form of text, voice, image, etc., and the task processing system can process the corresponding tasks.
[0031] S1200, calculating the correlation with several LoRA networks to be selected based on the feature vector of the task processing text; After receiving the task processing text and identifying the feature vector of the task processing text, the correlation with several candidate LoRA networks is calculated based on the feature vector of the task processing text. The candidate LoRA network is an expert LoRA (Low-Rank Adaptation) module. In order to more accurately calculate the correlation between the feature vector of the task processing text and the candidate LoRA network, this embodiment configures a learnable key vector (Key) for each candidate LoRA network, and calculates the correlation with the feature vector of the task processing text as a query vector (Query). Specifically, any one or more of the following calculation methods are combined: (1) based on cosine similarity, the correlation between the task processing text feature vector and each LoRA network feature vector is measured by calculating the cosine value between them. The value range of cosine similarity is between [-1, 1]. The closer the value is to 1, the higher the correlation; (2) based on distance measurement (such as Euclidean distance or Manhattan distance), the correlation is evaluated by calculating the distance between the feature vector and each LoRA network feature vector. The smaller the distance, the higher the correlation; (3) combined with a deep learning model, such as a neural network or an attention mechanism, the feature vector of the task processing text is further processed and matched to more accurately evaluate the correlation. For example, a small neural network can be used to jointly encode the feature vector of the task processing text and the feature vector of the LoRA network, and then the correlation between them is predicted through an output layer.
[0032] It should be noted that the LoRA network to be selected in this embodiment refers to a set of neural network modules that have been pre-trained and optimized for specific tasks or domains. These LoRA networks are stored in a dedicated library, and each LoRA network has unique parameters and structure, capable of handling specific types of tasks. For example, some LoRA networks may focus on natural language understanding tasks such as sentiment analysis or question-answering systems, while other LoRA networks may be optimized for text generation or machine translation tasks. Each LoRA network is described by a feature vector that reflects the semantic domain, task type, and performance characteristics of the LoRA network.
[0033] It should be pointed out that in order to improve the efficiency and accuracy of correlation calculation, this embodiment also configures an optimization strategy, including a hierarchical clustering method, to group the selected LoRA networks according to task type or field, and then perform correlation calculation within each group, thereby reducing the amount of calculation and improving the calculation speed.
[0034] It should be noted that this embodiment also configures the introduction of external knowledge bases or domain expert knowledge to enhance the correlation calculation. For example, the domain information of the task processing text can be matched with the domain label of the LoRA network as an important reference factor for the correlation calculation. Through these optimization and expansion strategies, the correlation between the task processing text and the candidate LoRA network can be more accurately evaluated, providing a reliable basis for the subsequent LoRA network selection.
[0035] Through the above implementation, the correlation between the feature vector of the task processing text and the selected LoRA network can be calculated efficiently and accurately, providing important support for the dynamic selection of the most suitable LoRA network.
[0036] S1300, determining the activation probability of each candidate LoRA network according to the correlation, and determining the target LoRA network to be activated according to the activation probability based on the preconfigured gated network; After calculating the correlation with several candidate LoRA networks based on the feature vector of the task processing text, the activation probability of each candidate LoRA network is determined according to the correlation, and the target LoRA network to be activated is determined according to the activation probability based on the preconfigured gating network. In the task processing method of the dynamic LoRA network, determining the activation probability of the candidate LoRA network and selecting the target LoRA network are key links to achieve efficient task processing. Specifically, the activation probability is calculated using a softmax function to convert the correlation score into a probability value. In one embodiment, for each candidate LoRA network, its activation probability can be calculated by the following formula: <center> \( P(e_i) = \frac{\exp(\text{similarity}(f_t, f_{e_i}))}{\sum_{j=1}^{N} \exp(\text{similarity}(f_t, f_{e_j}))} \) < / center> Where P(e_i) represents the activation probability of the i-th candidate LoRa network, similarity(f_t, f_{e_i}) represents the correlation score between the task processing text feature vector f_t and the i-th candidate LoRa network feature vector f_{e_i}, and N represents the total number of candidate LoRa networks. The softmax function can be used to normalize the correlation scores to probability values, so that the sum of the activation probabilities of all candidate LoRa networks is 1.
[0037] In order to more accurately select the target LoRA network, this embodiment introduces a preconfigured gated network. The function of the gated network is to dynamically select the most suitable LoRA network for task processing based on the activation probability. When the gated network selects the target LoRA network based on the activation probability, it can adopt multiple strategies: (1) Select the LoRA network with the highest activation probability as the target LoRA network. The advantage of this method is that it is simple and efficient, but in some cases it may ignore other LoRA networks with higher activation probabilities; (2) Use a weighted selection method to perform a weighted combination of multiple LoRA networks based on the activation probability to generate a comprehensive LoRA network. This method can fully utilize the advantages of multiple LoRA networks and improve the performance of task processing; (3) Introduce randomness and select the LoRA network with a certain probability based on the activation probability, thereby increasing the flexibility and robustness of the system. Using different strategies in different scenarios can effectively improve the accuracy and efficiency of the target LoRA network activation.
[0038] It should be noted that the gating network in this embodiment can adopt a variety of architectures, such as a multi-layer perceptron (MLP), a convolutional neural network (CNN), or a Transformer-based architecture. In practical applications, the gating network learns from training data to optimize its selection strategy. The training data includes the task processing text and its corresponding optimal LoRa network selection result. Through learning, the gating network can automatically adjust its selection strategy based on the feature vector of the task processing text and the activation probability of the candidate LoRa network, thereby improving the accuracy and efficiency of the selection.
[0039] S1400: Fusing the target LoRA network into the basic model to obtain a fusion model; After determining the activation probability of each candidate LoRA network based on the correlation, and determining the target LoRA network to be activated based on the activation probability based on the preconfigured gating network, the target LoRA network is integrated into the basic model to obtain a fusion model. In order to efficiently integrate the target LoRA network into the basic model, this application adopts low-rank decomposition technology. Low-rank decomposition is a method of decomposing a matrix into two or more low-rank matrices. In this way, the number of parameters can be significantly reduced while maintaining the performance of the model. Specifically, the parameter matrix of the target LoRA network can be expressed as the product of two low-rank matrices, that is, W=UV T, where U and V are low-rank matrices. This decomposition method not only reduces the overhead of parameter storage and computation, but also makes the fusion process more efficient. During the fusion process, the low-rank decomposition parameters of the target LoRA network are dynamically loaded into the base model, where the base model can be a pre-trained large language model (LLM) such as BERT, GPT, or its variants. The specific fusion method is to merge the low-rank matrices U and V of the target LoRA network with the corresponding layers of the base model. For example, if the target LoRA network is a fine-tuned module for a specific task, its low-rank decomposition parameters can be inserted into the hidden layer of the base model and fused with the parameters of the base model through linear combination. This fusion method can seamlessly integrate the characteristics of the target LoRA network into the base model without retraining the base model.
[0040] It should be noted that, in this embodiment, when integrating the target LoRA network, only the parameters related to the target LoRA network will be updated, while the other parts of the basic model remain unchanged.
[0041] Through the above implementation, the target LoRA network can be efficiently integrated into the basic model, forming a fusion model that combines the advantages of the basic model and the LoRA network, providing strong support for subsequent task processing.
[0042] S1500: Perform task processing based on the semantic and contextual features of the task processing text using a fusion model, so as to load the target LoRA network and then output a processing result of the task processing text.
[0043] After integrating the target LoRa network into the base model to generate a fused model, the fused model is used to perform task processing based on the semantic and contextual features of the task-processing text. After loading the target LoRa network, the processing result of the task-processing text is output. After the fused model is prepared, the task-processing text enters the fused model through an input interface. The fused model first encodes the task-processing text using the encoder layer of the base model to generate contextual feature vectors. These feature vectors contain the semantic and contextual information of the task-processing text. Subsequently, the low-rank decomposition parameters of the target LoRa network are dynamically loaded into the fused model and interact with the parameters of the base model, enabling the fused model to leverage the target LoRa network's task-specific knowledge to more accurately process the task-processing text. The specific implementation of task processing depends on the task type. For example, for a text classification task, the fused model inputs the feature vector of the task-processing text into a classifier. The classifier uses the knowledge of the target LoRa network to classify the text and output the classification result. For a text generation task, the fused model generates corresponding text content based on the contextual features of the task-processing text. The target LoRa network provides domain-specific generative knowledge in this process, making the generated text more suitable for the task requirements. If it is a question-answering task, the fusion model will extract the answer based on the question text and context information through the knowledge of the target LoRA network and output the answer text.
[0044] In the above embodiment, by receiving the task processing text, identifying the feature vector of the task processing text, wherein the feature vector at least indicates the semantic and contextual features of the task processing text; calculating the correlation with several candidate LoRA networks according to the feature vector of the task processing text; determining the activation probability of each candidate LoRA network according to the correlation, and determining the target LoRA network to be activated according to the activation probability based on the preconfigured gating network; fusing the target LoRA network into the basic model to obtain a fusion model; using the fusion model to perform task processing based on the semantic and contextual features of the task processing text, so as to load the target LoRA network and output the result. The processing results of the task processing text are obtained; by dynamically loading the target LoRA network and adopting low-rank decomposition technology, only low-rank parameters related to the task are loaded when needed, avoiding redundant calculation and storage of the entire model. This dynamic adjustment method significantly reduces the consumption of computing resources, especially when processing large-scale language models, it can effectively reduce video memory usage and computing time, and the fusion model can efficiently utilize the knowledge of the target LoRA network to perform accurate task processing on the task processing text and output high-quality results. This method of dynamically selecting and fusing LoRA networks not only improves the efficiency and performance of task processing, but also provides a flexible and efficient solution for natural language processing tasks.
[0045] In some implementations, after receiving a task processing text and identifying a feature vector of the task processing text, S1100 includes: S1111, processing text matching fields or label information for the task; In this embodiment, after receiving the task processing text and identifying the feature vector of the task processing text, the task processing text is matched with domain or label information. The domain or label information refers to the subject, domain or task type related to the task processing text. This information can be obtained through a predefined domain classification table or label library. For example, in a natural language processing task, the domain information can be "news reports", "medical health", "finance and economics", etc., and the label information can be "sentiment analysis", "question and answer system", "text generation", etc. These domain or label information can be pre-defined by domain experts and stored in the system's knowledge base for matching during task processing. In order to match the task processing text with the corresponding domain or label information, in one embodiment, this embodiment is based on keyword matching, by pre-extracting the key feature words of each domain or label, and searching for these keywords in the task processing text. For example, if the task processing text contains keywords such as "stocks", "market", and "investment", it can be matched to the "finance and economics" field; in addition, machine learning-based classifiers, such as support vector machines (SVM) or deep learning models (such as CNN or Transformer), can be used to perform field or label classification on the task processing text. These classifiers can be trained with a large amount of labeled data to improve the accuracy and robustness of matching.
[0046] It should be noted that in this embodiment, task processing documents may involve multiple fields or labels. To handle this, the system can employ a multi-label classification algorithm, allowing task processing documents to be assigned multiple fields or labels. During the matching process, the system can sort the fields or labels based on their confidence or relevance scores, selecting the most relevant fields or labels as matching results.
[0047] S1112. Based on a preconfigured attention mechanism, weight the key information in the special effect vector according to the domain or label information and then update the feature vector.
[0048] After matching domain or label information to the task text, the key information in the feature vector is weighted using an attention mechanism. The core idea of the attention mechanism is to enable the model to dynamically focus on important parts of the feature vector while ignoring unimportant parts. In this application, a pre-configured attention mechanism can weight key information in the feature vector based on the domain or label information. For example, for task text in the "health" domain, the attention mechanism may assign higher weights to medical terminology or health-related features, while lowering the weights for other irrelevant features. This domain-based attention mechanism can be learned using training data to optimize the weight allocation strategy. During the weighted feature vector update process, the feature vector of the task text is first input into the pre-configured attention module. The attention module calculates weights for each feature dimension based on the matched domain or label information. These weights are typically calculated using a learnable parameter matrix and the domain or label embedding vector to ensure that the weights are highly correlated with the domain or label information. The calculated weights are then multiplied by each dimension of the feature vector to obtain a weighted feature vector. Finally, the weighted feature vector is used as the updated feature vector for subsequent task processing. This weighted update process can highlight the key information in the feature vector, allowing the model to better focus on task-related features, thereby improving task processing performance.
[0049] In order to further improve the quality of feature vectors, this embodiment performs optimization processing on the basis of the extracted feature vectors. By configuring a multi-task learning mechanism, the feature vectors are combined with task-related labels or domain information to further enrich the semantic expression ability of the feature vectors. By configuring an attention mechanism, the key information in the feature vectors is weighted to enhance attention to important semantic units. The feature vector finally generated not only contains the original semantic information of the text, but also integrates the contextual environment and task-related background information, providing a solid foundation for subsequent LoRA network selection.
[0050] In some embodiments, the S1300 determines the target LoRA network to be activated according to the activation probability based on the preconfigured gated network, including: S1311. Obtain the current task processing scenario; In this embodiment, when determining the target LoRA network to activate based on the activation probability using a preconfigured gated network, the system first obtains the current task processing scenario. The task processing scenario refers to the context of the current task, including the task type, domain, user requirements, and input data characteristics. For example, task types can include text classification, sentiment analysis, question-and-answer systems, or text generation; and domains can include news, healthcare, finance, or education. The definition of scenarios needs to be categorized and labeled based on the actual application scenario to ensure accurate system identification. For example, task processing scenarios can be categorized into "real-time question-and-answer," "batch text classification," and "multi-domain text generation," each corresponding to different task requirements and optimization goals. Specifically, in a question-and-answer system, the user-input question text may contain contextual information, such as "disease consultation in the medical field" or "investment advice in the financial field." The system can extract contextual information from the input text using natural language processing techniques (such as keyword extraction and named entity recognition). Contextual information can also be obtained through user contextual information (such as historical user behavior and preferences) or task metadata (such as task source and task type label). For example, if a user initiates a consultation task on a healthcare platform, the system can directly identify the context as "medical field."
[0051] It should be noted that in actual applications, task processing scenarios may change dynamically. For example, users may switch between different task types or fields on the same platform. To adapt to such dynamic changes, the system needs to update scenario information in real time. By introducing a context-aware module, changes in task input can be monitored in real time and scenario classification can be adjusted dynamically. For example, when a user switches from "news reading" to "stock analysis," the system can automatically identify the scenario change and update the task processing scenario to "financial field." In addition, the acquisition and classification strategies of scenario information can be further optimized through user feedback or a posteriori analysis of task processing results.
[0052] S1312. Configuring a gated network according to the scenario to adjust a probabilistic activation strategy of the gated network; After obtaining the current task processing scenario, the gating network is configured according to the scenario to adjust the probability activation strategy of the gating network. The gating network is one of the core components of the dynamic LoRA network task processing method. Its main function is to dynamically select the most suitable LoRA network based on the feature vector of the task processing text and the scene information. Specifically, the scene embedding vector is introduced, and the scene embedding vector is spliced or interacted with the feature vector of the task processing text as the input of the gating network. For example, for the task processing scenario of "medical field", the embedding vector of the medical field can be combined with the feature vector of the task processing text, so that the gating network can dynamically adjust the activation probability according to the scene information. In addition, the scene perception of the gating network can be fine-tuned through training data, so that it can better adapt to the task requirements in different scenarios. For example, by fine-tuning the gating network on medical field data, it can more accurately select the LoRA network suitable for medical tasks.
[0053] S1313: Determine the target LoRA network to be activated according to the adjusted door network probability activation strategy.
[0054] After configuring the gated network based on the scenario and adjusting its probabilistic activation strategy, the target LoRA network to be activated is determined based on the adjusted gated network probabilistic activation strategy. The activation probability of each candidate LoRA network is calculated based on the adjusted gated network activation strategy. The output of the gated network is a probability distribution representing the probability of each LoRA network being activated. Based on these activation probabilities, the LoRA network with the highest activation probability is selected as the target LoRA network. For example, if the activation probabilities output by the gated network are [0.7, 0.2, 0.1], the LoRA network with an activation probability of 0.7 is selected as the target LoRA network. This selection strategy is simple and efficient, quickly determining the most suitable LoRA network for the task at hand. In some cases, task processing may require the synergy of multiple LoRA networks. For example, in complex multi-domain tasks, multiple LoRA networks may need to be activated simultaneously to improve task processing performance. In this case, multiple LoRA networks can be selected based on their activation probabilities and integrated into the base model. For example, a LoRA network with an activation probability greater than a certain threshold (such as 0.3) can be selected as the target LoRA network. In this way, the system can fully utilize the advantages of multiple LoRA networks to improve the accuracy and robustness of task processing.
[0055] It's important to note that during task processing, the target LoRA network selection may need to be dynamically adjusted. For example, if the task processing scenario changes or the task processing results are unsatisfactory, the system can dynamically adjust the target LoRA network based on feedback. By introducing a feedback mechanism, the activation strategy of the gated network can be adjusted based on the accuracy of the task processing results or user satisfaction. For example, if the task processing results of a target LoRA network are unsatisfactory, its activation probability can be reduced while the activation probability of other LoRA networks can be increased. This dynamic adjustment mechanism allows the system to adapt to changes in task processing in real time, improving overall system performance.
[0056] This implementation achieves efficient and accurate target LoRA network selection by obtaining scene information of the current task processing and dynamically adjusting the activation strategy of the gated network according to the scene. This method can significantly improve the adaptability and flexibility of the system, enabling it to better cope with the diverse needs of different task scenarios. Through scene-aware activation strategy adjustment, the system can dynamically select the LoRA network that best suits the current task, thereby improving the accuracy and efficiency of task processing. In addition, this dynamic adjustment mechanism can further optimize the activation strategy based on feedback from task processing results, further improving the performance and robustness of the system.
[0057] In some embodiments, the target LoRA network is stored in a distributed parameter storage system, and S1400 integrates the target LoRA network into the basic model to obtain a fusion model, including: S1411, retrieving the target LoRA network from the distributed parameter storage system based on the product quantization indexing mechanism; In this embodiment, the LoRA network to be selected is stored in a distributed parameter storage system, which is an infrastructure for efficiently storing and managing a large number of LoRA network parameters. The distributed parameter storage system is usually composed of multiple storage nodes, each of which is responsible for storing a portion of the parameters of the LoRA network. The parameters of each LoRA network are divided into multiple parts and stored in different nodes to ensure the high availability and fault tolerance of the system. The target LoRA network is then retrieved from the distributed parameter storage system based on a product quantization indexing mechanism. The product quantization indexing mechanism is used to quickly retrieve the LoRA network most relevant to the task processing text. Specifically, the vector of each LoRA network is product quantized, that is, the vector is decomposed into multiple sub-vectors, and the sub-vectors are quantized using a codebook in each subspace. Then, a product quantization-based index structure is constructed. This structure uses the encoded vector to quickly retrieve the LoRA network most similar to the feature vector.
[0058] S1412. Fusing the target LoRA network into the basic model through low-rank decomposition to obtain a fusion model.
[0059] After retrieving the target LoRA network, the target LoRA network is integrated into the base model to obtain a fusion model. Before integrating the target LoRA network, the system needs to load the parameters of the target LoRA network from the distributed parameter storage system. These parameters include matrices after low-rank decomposition or other optimized parameter representations. In order to ensure the integrity and consistency of the parameters during the loading process and avoid parameter errors caused by network delays or node failures in the distributed storage environment, after loading is completed, the parameters of the target LoRA network are temporarily stored in the memory and prepared for fusion with the base model. In order to improve loading efficiency, this embodiment adopts an asynchronous loading mechanism so that the base model can continue to perform other tasks while waiting for parameter loading. Furthermore, this embodiment uses low-rank decomposition to represent the parameters of the target LoRA network as the product of two low-rank matrices, which can significantly reduce the number of parameters while retaining the key features of the LoRA network. During fusion, the low-rank matrix is linearly combined with the corresponding layer of the base model. For example, if the target LoRA network is a fine-tuning module for a specific task, its low-rank matrix can be inserted into the hidden layer of the base model. The specific method of fusion is to add or multiply the low-rank matrix with the weight matrix of the basic model through matrix operations, so as to achieve dynamic update of parameters. This fusion strategy can not only quickly adapt to new task requirements, but also avoid large-scale retraining of the basic model.
[0060] This embodiment significantly improves the storage efficiency and retrieval speed of the system by storing the candidate LoRA networks in a distributed parameter storage system and using an indexing mechanism based on product quantization to retrieve the target LoRA network. The product quantization indexing mechanism can quickly locate the LoRA network most relevant to the task processing text in large-scale parameter storage, reducing the retrieval time and improving the response speed of the system. At the same time, the target LoRA network is integrated into the basic model through low-rank decomposition technology, further optimizing the computing and storage overhead, enabling the system to efficiently handle diverse tasks. This distributed storage and dynamic fusion method not only improves the scalability and flexibility of the system, but also enhances the robustness and adaptability of the system, enabling it to better cope with complex and changing task requirements.
[0061] In some embodiments, the target LoRA network includes at least two of a first LoRA network and a second LoRA network, and the S1412 fuses the target LoRA network into the basic model through low-rank decomposition to obtain a fusion model, including: S1421: Obtain a first correlation of the first LoRA network and a second correlation of the second LoRA network; In this embodiment, the target LoRA network includes at least two of a first LoRA network and a second LoRA network. The first correlation of the first LoRA network and the second correlation of the second LoRA network can be obtained through the above S1200.
[0062] S1422: configuring a first weight of a first LoRA network according to the first correlation and configuring a second weight of a second LoRA network according to the second correlation; In this embodiment, after obtaining the relevance of the target LoRA network, the first weight of the first LoRA network is configured according to the first relevance, and the second weight of the second LoRA network is configured according to the second relevance. In the fusion of multiple LoRA networks, the relevance of each LoRA network determines its contribution weight in the fusion model. By dynamically configuring weights based on relevance, the system can better utilize the advantages of each LoRA network, improve the performance of the fusion model, and ensure that LoRA networks that are more relevant to the task processing context play a greater role in the fusion model, thereby improving the accuracy and efficiency of task processing.
[0063] It can be pointed out that in order to further optimize the weight configuration, this embodiment also introduces an extension strategy, including adjusting the weight in combination with the context information of the task processing text or the historical selection record. If the task processing text has similar contextual features or there are similar task processing situations in the historical selection records, the weight of the relevant LoRA network can be appropriately increased; in addition, the weight configuration strategy can be dynamically adjusted according to the results of task processing by introducing a feedback mechanism. If the task processing result of a LoRA network is not ideal, its weight can be appropriately reduced to improve the performance of the system.
[0064] S1423: Through low-rank decomposition, the first LoRA network and the second LoRA network are integrated into the basic model according to the first weight and the second weight to obtain a fusion model.
[0065] In this embodiment, after configuring the first weight of the first LoRA network according to the first correlation and configuring the second weight of the second LoRA network according to the second correlation, the first LoRA network and the second LoRA network are fused into the basic model according to the first weight and the second weight through low-rank decomposition to obtain a fusion model. The present application adopts low-rank decomposition technology to represent the parameters of each LoRA network as the product of two low-rank matrices. This representation method significantly reduces the number of parameters while retaining the key features of the LoRA network. During fusion, the low-rank matrix of each LoRA network is weightedly fused with the corresponding layer of the basic model. Specifically, for the first LoRA network and the second LoRA network, the low-rank matrix is weightedly combined with the weight matrix of the basic model according to their first weight and second weight, respectively. In this way, the contribution of each LoRA network is dynamically adjusted according to its weight, thereby achieving efficient fusion.
[0066] In some embodiments, when the feature vector of the task-processing text represents that the current task is a multimodal scenario, the gating network is configured as a cross-modal gating network, and the method further includes: S1521: Activate the target LoRA network of the corresponding modality according to the cross-modal gating network to implement joint routing decision through the cross-modal gating network, and output the multimodal collaborative processing result of the task processing text.
[0067] In this embodiment, when the feature vector of the task processing text represents that the current task is a multimodal scenario, the gating network is configured as a cross-modal gating network. The cross-modal gating network is a gating mechanism specifically used to process multimodal tasks. Its core function is to dynamically select the most suitable modal LoRA network based on the feature vector of the task processing text (containing multimodal information). The cross-modal gating network is usually composed of a multimodal input layer, a feature fusion layer and a gating decision layer. The multimodal input layer is responsible for receiving feature vectors from different modalities (such as text, images, audio, etc.). The feature fusion layer fuses the features of different modalities through cross-modal fusion technology (such as attention mechanism, multimodal Transformer, etc.) to generate a joint feature representation; the gating decision layer calculates the activation probability of each modality LoRA network based on the joint feature representation.
[0068] When the feature vector of the task processing text represents a multimodal scenario, the cross-modal gating network calculates the activation probability of each modal LoRa network based on the joint feature representation. For example, if the task processing text contains both text and image information, the cross-modal gating network calculates the activation probabilities of the text modal LoRa network and the image modal LoRa network. Based on these probabilities, the gating network selects the modal LoRa network with the highest activation probability as the target LoRa network. If the task requires the collaborative processing of multiple modal LoRa networks, the gating network can simultaneously activate multiple modal LoRa networks and fuse their outputs. After activating the target LoRa network for the corresponding modality, the cross-modal gating network makes a joint routing decision and outputs the multimodal collaborative processing result. In multimodal task processing, information from different modalities must be processed collaboratively to achieve optimal performance. Joint routing decision means that the cross-modal gating network dynamically adjusts the contribution weights of the different modal LoRa networks based on the multimodal characteristics of the task processing text and fuses their outputs to generate the final multimodal collaborative processing result. This approach fully leverages the strengths of different modalities and improves the accuracy and efficiency of task processing. In this embodiment, the joint routing decision can be made through weighted fusion, that is, the outputs of the LoRA networks of different modalities are weighted and summed according to their activation probabilities. In this way, the cross-modal gating network can dynamically adjust the contribution weights of different modalities according to the multimodal characteristics of the task processing text, thereby achieving efficient joint routing decision.
[0069] This embodiment realizes dynamic modal selection and joint routing decision in multimodal task processing by configuring a cross-modal gating network. This method can significantly improve the adaptability and flexibility of the system, enabling it to efficiently process multimodal tasks. Through the cross-modal gating network, the system can dynamically select the most suitable modal LoRA network according to the multimodal characteristics of the task, and realize the collaborative processing of multimodal information through joint routing decisions, thereby improving the accuracy and efficiency of task processing. In addition, this dynamic adjustment mechanism can further optimize the joint routing strategy based on the feedback of the task processing results, further improving the performance and robustness of the system.
[0070] In some embodiments, when the request for the task processing text comes from an edge device, after S1300 determines the target LoRA network to be activated according to the activation probability based on the preconfigured gated network, it also includes: S1331. In response to the request of the task processing text, send the target LoRA network to the edge device, so that the edge device downloads and caches the target LoRA network.
[0071] In this embodiment, when the task processing document request comes from an edge device, such as one with poor performance or limited computing resources, it is necessary to ensure that the edge device can successfully download and cache the network to achieve efficient local task processing. Specifically, in response to the task processing document request, the target LoRA network is sent to the edge device, allowing it to download and cache the target LoRA network. After downloading, the edge device needs to cache the target LoRA network in local storage to enable rapid model loading for subsequent task processing. The caching mechanism needs to take into account the dynamic changes in device storage capacity and task requirements. For example, an LRU (least recently used) caching strategy can be adopted to dynamically manage cache space based on model usage frequency and time. If device storage space is insufficient, the system can automatically delete the least recently used model to make room for the new target LoRA network. Furthermore, a hierarchical caching mechanism can be used to cache model parameters in the device's fast storage (such as memory or SSD) while caching the model structure in slower storage (such as a hard drive) to optimize cache performance. The cached target LoRA network needs to be regularly verified for integrity and validity. The integrity of the model file can be verified by a checksum (such as MD5 or SHA-256) to ensure that no data is corrupted during the download process. If the cached model file is found to be corrupted or outdated, the client application can automatically re-download the model file to ensure the accuracy of the cached data. In addition, the system can also ensure that the edge device always caches the latest version of the target LoRA network through a version control mechanism. For example, the server can embed a version number in the model file, and the client application checks the version number when downloading. If it finds that the locally cached model version is lower than the version on the server, it automatically updates the cache. This implementation significantly improves the task processing efficiency and response speed in the edge computing environment by sending the target LoRA network to the edge device and caching it. This method can effectively reduce the edge device's dependence on cloud servers, reduce network latency and bandwidth consumption, and improve the reliability and privacy of the system. Through the dynamic caching mechanism, the edge device can quickly load the target LoRA network according to task requirements and achieve efficient local task processing. In addition, this optimization strategy can also dynamically adjust the caching strategy according to the device's storage capacity and usage scenarios, further improving the flexibility and adaptability of the system.
[0072] Please refer to the following for details: Figure 2 , Figure 2 This is a schematic diagram of the basic structure of the task processing device based on the dynamic LoRA network in this embodiment.
[0073] like Figure 2As shown, a task processing device based on a dynamic LoRA network includes: a task receiving module 1100, which is used to receive a task processing text and identify a feature vector of the task processing text, wherein the feature vector at least indicates the semantic and contextual features of the task processing text; a correlation calculation module 1200, which is used to calculate the correlation with several candidate LoRA networks based on the feature vector of the task processing text; a matching activation module 1300, which is used to determine the activation probability of each candidate LoRA network based on the correlation, and determine the target LoRA network to be activated based on the activation probability based on the preconfigured gating network; a network fusion module 1400, which is used to fuse the target LoRA network into the basic model to obtain a fusion model; a processing output module 1500, which is used to use the fusion model to perform task processing based on the semantic and contextual features of the task processing text, so as to output the processing result of the task processing text after loading the target LoRA network.
[0074] The above-mentioned task processing device based on the dynamic LoRA network receives a task processing text and identifies a feature vector of the task processing text, wherein the feature vector at least indicates the semantic and contextual features of the task processing text; calculates the correlation with several candidate LoRA networks according to the feature vector of the task processing text; determines the activation probability of each candidate LoRA network according to the correlation, and determines the target LoRA network to be activated according to the activation probability based on the preconfigured gating network; integrates the target LoRA network into the basic model to obtain a fusion model; uses the fusion model to perform task processing based on the semantic and contextual features of the task processing text to load the target LoRA The network outputs the processing result of the task processing text; by dynamically loading the target LoRA network and adopting low-rank decomposition technology, only low-rank parameters related to the task are loaded when needed, avoiding redundant calculation and storage of the entire model. This dynamic adjustment method significantly reduces the consumption of computing resources, especially when processing large-scale language models, it can effectively reduce video memory occupancy and computing time, and the fusion model can efficiently utilize the knowledge of the target LoRA network to perform accurate task processing on the task processing text and output high-quality results. This method of dynamically selecting and fusing LoRA networks not only improves the efficiency and performance of task processing, but also provides a flexible and efficient solution for natural language processing tasks.
[0075] Optionally, the task receiving module is further configured to: Processing text matching domain or label information for the task; The feature vector is updated after weighting the key information in the special effect vector according to the domain or label information based on a preconfigured attention mechanism.
[0076] Optionally, the matching activation module is further configured to: Get the scene of the current task processing; Configuring the gating network according to the scenario to adjust the probabilistic activation strategy of the gating network; The target LoRA network to be activated is determined based on the adjusted door network probability activation strategy.
[0077] Optionally, the LoRA network to be selected is stored in a distributed parameter storage system; the network fusion module is further used to: Retrieve the target LoRA network from the distributed parameter storage system based on the product quantization indexing mechanism; The target LoRA network is fused into the basic model through low-rank decomposition to obtain a fusion model.
[0078] Optionally, the target LoRA network includes at least two of a first LoRA network and a second LoRA network; and the network fusion module is further configured to: Obtaining a first correlation of the first LoRA network and a second correlation of the second LoRA network; configuring a first weight of a first LoRA network according to the first correlation and configuring a second weight of a second LoRA network according to the second correlation; By low-rank decomposition, the first LoRA network and the second LoRA network are fused into a basic model according to the first weight and the second weight to obtain a fusion model.
[0079] Optionally, when the feature vector of the task processing text represents that the current task is a multimodal scenario, the gating network is configured as a cross-modal gating network, and the matching activation module is further used to: The target LoRA network of the corresponding modality is activated according to the cross-modal gating network to realize joint routing decision through the cross-modal gating network and output the multimodal collaborative processing result of the task processing text.
[0080] Optionally, when the request for task processing text comes from an edge device, the apparatus is further configured to: In response to the request of the task processing text, the target LoRA network is sent to the edge device, so that the edge device downloads and caches the target LoRA network.
[0081] To solve the above technical problems, the present application also provides a computer device. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.
[0082] like Figure 3As shown, a schematic diagram of the internal structure of a computer device. The computer device includes a processor, a non-volatile storage medium, a memory and a network interface connected via a system bus. Among them, the non-volatile storage medium of the computer device stores an operating system, a database and computer-readable instructions, and the database may store a control information sequence, and the computer-readable instructions are executed by the processor. The processor of the computer device is used to provide computing and control capabilities to support the operation of the entire computer device. The memory of the computer device may store computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor may execute a task processing method based on a dynamic LoRA network. The network interface of the computer device is used to connect and communicate with the terminal. Those skilled in the art will understand that, Figure 3 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0083] In this embodiment, the processor is used to execute Figure 2 The memory stores the program code and various data required to execute the specific functions of the task receiving module 1100, correlation calculation module 1200, matching activation module 1300, network fusion module 1400, and processing output module 1500. The network interface is used to transmit data between user terminals or servers. The memory in this embodiment stores the program code and data required to execute all submodules of the task processing device based on the dynamic LoRA network. The server can call the server's program code and data to execute the functions of all submodules.
[0084] The computer device receives a task processing text and identifies a feature vector of the task processing text, wherein the feature vector at least indicates the semantic and contextual features of the task processing text; calculates the correlation with a plurality of candidate LoRA networks according to the feature vector of the task processing text; determines the activation probability of each candidate LoRA network according to the correlation, and determines the target LoRA network to be activated according to the activation probability based on the preconfigured gating network; integrates the target LoRA network into the basic model to obtain a fusion model; uses the fusion model to perform task processing based on the semantic and contextual features of the task processing text, so as to load the target LoRA network and output the task processing result; The processing results of the task processing text; by dynamically loading the target LoRA network and adopting low-rank decomposition technology, only low-rank parameters related to the task are loaded when needed, avoiding redundant calculation and storage of the entire model. This dynamic adjustment method significantly reduces the consumption of computing resources, especially when processing large-scale language models, it can effectively reduce video memory usage and computing time, and the fusion model can efficiently utilize the knowledge of the target LoRA network to perform accurate task processing on the task processing text and output high-quality results. This method of dynamically selecting and fusing LoRA networks not only improves the efficiency and performance of task processing, but also provides a flexible and efficient solution for natural language processing tasks.
[0085] The present application also provides a storage medium storing computer-readable instructions. When the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the task processing method based on the dynamic LoRA network described in any of the above embodiments.
[0086] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When executed, the program can include the processes in the above-described method embodiments. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM).
[0087] Those skilled in the art will appreciate that the steps, measures, and schemes in the various operations, methods, and processes discussed in this application may be interchanged, modified, combined, or deleted. Furthermore, other steps, measures, and schemes in the various operations, methods, and processes discussed in this application may also be interchanged, modified, rearranged, decomposed, combined, or deleted. Furthermore, steps, measures, and schemes in the prior art that are similar to those disclosed in this application may also be interchanged, modified, rearranged, decomposed, combined, or deleted.
[0088] The above description is only part of the implementation methods of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A task processing method based on a dynamic LoRA network, characterized in that: include: receiving a task processing text, and identifying a feature vector of the task processing text, wherein the feature vector at least indicates semantic and contextual features of the task processing text; Calculate the correlation with several candidate LoRA networks based on the feature vector of the task-processed text; Determine the activation probability of each candidate LoRA network according to the correlation, and determine the target LoRA network to be activated according to the activation probability based on the preconfigured gating network; The target LoRA network is integrated into the basic model to obtain a fusion model; The fusion model is used to perform task processing based on the semantic and contextual features of the task processing text, so as to load the target LoRA network and output the processing result of the task processing text.
2. The task processing method based on the dynamic LoRA network according to claim 1, characterized in that: After receiving the task processing text and identifying the feature vector of the task processing text, the method further includes: Processing text matching domain or label information for the task; The feature vector is updated after weighting the key information in the special effect vector according to the domain or label information based on a preconfigured attention mechanism.
3. The task processing method based on the dynamic LoRA network according to claim 1, characterized in that: The preconfigured gated network determines a target LoRA network to be activated according to the activation probability, including: Get the scene of the current task processing; Configuring the gating network according to the scenario to adjust the probabilistic activation strategy of the gating network; The target LoRA network to be activated is determined based on the adjusted door network probability activation strategy.
4. The task processing method based on the dynamic LoRA network according to claim 1, characterized in that: The selected LoRA network is stored in a distributed parameter storage system, and the target LoRA network is integrated into the basic model to obtain a fusion model, including: Retrieve the target LoRA network from the distributed parameter storage system based on the product quantization indexing mechanism; The target LoRA network is fused into the basic model through low-rank decomposition to obtain a fusion model.
5. The task processing method based on the dynamic LoRA network according to claim 4, characterized in that: The target LoRA network includes at least two of a first LoRA network and a second LoRA network, and the target LoRA network is fused into a basic model by low-rank decomposition to obtain a fusion model, including: Obtaining a first correlation of the first LoRA network and a second correlation of the second LoRA network; configuring a first weight of a first LoRA network according to the first correlation and configuring a second weight of a second LoRA network according to the second correlation; By low-rank decomposition, the first LoRA network and the second LoRA network are fused into a basic model according to the first weight and the second weight to obtain a fusion model.
6. The task processing method based on the dynamic LoRA network according to claim 1, characterized in that: When the feature vector of the task processing text represents that the current task is a multimodal scenario, the gating network is configured as a cross-modal gating network, and the method further includes: The target LoRA network of the corresponding modality is activated according to the cross-modal gating network to realize joint routing decision through the cross-modal gating network and output the multimodal collaborative processing result of the task processing text.
7. The task processing method based on dynamic LoRA network according to claim 1, characterized in that: When the request for the task processing text comes from an edge device, after the pre-configured gated network determines the target LoRA network to be activated according to the activation probability, the method further includes: In response to the request of the task processing text, the target LoRA network is sent to the edge device, so that the edge device downloads and caches the target LoRA network.
8. A task processing device based on a dynamic LoRA network, characterized in that: include: A task receiving module, configured to receive a task processing text and identify a feature vector of the task processing text, wherein the feature vector at least indicates semantic and contextual features of the task processing text; A correlation calculation module is used to calculate the correlation with several candidate LoRA networks based on the feature vector of the task processing text; A matching activation module is used to determine the activation probability of each candidate LoRA network according to the correlation, and determine the target LoRA network to be activated according to the activation probability based on the preconfigured gating network; A network fusion module is used to fuse the target LoRA network into the basic model to obtain a fusion model; A processing output module is used to perform task processing based on the semantic and contextual features of the task processing text using a fusion model to output the processing result of the task processing text after loading the target LoRA network.
9. A computer device comprising a memory and a processor, wherein the memory stores computer-readable instructions, and when the computer-readable instructions are executed by the processor, the processor performs the steps of the task processing method based on the dynamic LoRA network according to any one of claims 1 to 7.
10. A storage medium storing computer-readable instructions, wherein when the computer-readable instructions are executed by one or more processors, the one or more processors execute the steps of the task processing method based on the dynamic LoRA network according to any one of claims 1 to 7.
Citation Information
Cited By
Large model adaptive training method and system based on task semantic perception
CN120688563A
A task semantic perception-based large model self-adaptive training method and system
CN120688563B
Multi-tenant ai inference service platform lora hot switching method and system, medium
CN122661349A