Model processing method and device, data processing method and device, equipment and storage medium
By quantizing the pre-trained large model and transforming the low-rank adaptive structure, and fine-tuning the weight, we generate a lightweight large model on the end side, solving the problem of accuracy loss after quantization, and achieving efficient multimodal task execution on edge devices.
Patent Information
- Application Number
- CN202510115345.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-16
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Large models may lead to accuracy losses during the quantization process, affecting model performance, especially in resource-constrained edge equipment environments.
By obtaining pre-trained multimodal large model for quantization, the model structure is transformed using the pre-constructed low-rank adaptive structure, and the weights of the low-rank matrix in the modified model are trained and fine-tuned to generate a lightweight large model on the end.
The accuracy of the quantized large model is restored, and the problem of inaccurate multimodal task results caused by low accuracy is avoided, and efficient multimodal task execution on edge devices is achieved.
Smart Images

Figure CN120012844A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a model processing method, a data processing method, a device, an electronic device and a storage medium. Background Art
[0002] With the rapid development of artificial intelligence technology, large models have achieved remarkable results in many fields such as natural language processing and computer vision, becoming an important force in promoting innovation in intelligent applications. However, large models usually have billions or even tens of billions of parameters, which leads to a sharp increase in the demand for computing and storage resources, limiting the application of large models in resource-constrained environments such as edge devices.
[0003] In order to solve these problems faced by large models in practical applications, model quantization methods have emerged. Model quantization mainly reduces the numerical accuracy of the model and represents the parameters of the large model from high-precision floating-point types to low-precision integer forms, thereby achieving model compression and acceleration without significantly reducing the accuracy of the model. Although large model quantization technology has many advantages, it still faces many technical difficulties in practical applications. Among them, the most important problem is that the quantization process may cause a certain degree of accuracy loss, affecting the performance of the model. Therefore, how to restore the accuracy of the quantized model has become a technical problem that needs to be solved urgently. Summary of the invention
[0004] The present invention provides a model processing method, a data processing method, a device, an electronic device, a storage medium and a computer program product.
[0005] According to one aspect of the present invention, there is provided a model processing method, comprising:
[0006] Obtain a pre-trained multimodal large model, and quantize the multimodal large model to obtain a first lightweight large model;
[0007] Using a pre-built low-rank adaptive structure, the model structure of the first lightweight large model is modified to obtain a second lightweight large model; wherein the low-rank adaptive structure is composed of a first low-rank matrix, an activation function, and a second low-rank matrix connected in series;
[0008] The weights of the first low-rank matrix and the second low-rank matrix in the second lightweight model are trained and fine-tuned, and the trained and fine-tuned second lightweight model is used as an end-side lightweight model deployed in an end-side device to perform multimodal tasks.
[0009] According to another aspect of the present invention, there is provided a data processing method, comprising:
[0010] Acquire the data to be processed, and input the data to be processed into the lightweight large model on the end side; wherein the lightweight large model on the end side is obtained by processing according to the model processing method in the embodiment of the present invention;
[0011] Through the lightweight large model on the terminal, at least one multimodal task including content generation tasks, understanding and question-answering tasks, analysis and classification tasks, and auxiliary decision-making tasks is performed on the processed data to obtain the data processing results.
[0012] According to another aspect of the present invention, there is provided a model processing device, comprising:
[0013] A quantization module, used to obtain a pre-trained multimodal large model and quantize the multimodal large model to obtain a first lightweight large model;
[0014] A transformation module, used to transform the model structure of the first lightweight large model by using a pre-built low-rank adaptive structure to obtain a second lightweight large model; wherein the low-rank adaptive structure is composed of a first low-rank matrix, an activation function, and a second low-rank matrix connected in series;
[0015] A fine-tuning module is used to train and fine-tune the weights of the first low-rank matrix and the second low-rank matrix in the second lightweight large model, and use the trained and fine-tuned second lightweight large model as an end-side lightweight large model deployed in an end-side device to perform multimodal tasks.
[0016] According to another aspect of the present invention, there is provided a data processing device, comprising:
[0017] A data acquisition module is used to acquire data to be processed and input the data to be processed into a lightweight large model on the end side; wherein the lightweight large model on the end side is obtained by processing according to the model processing method in the embodiment of the present invention;
[0018] The data processing module is used to perform at least one multimodal task including content generation tasks, understanding and question-answering tasks, analysis and classification tasks, and auxiliary decision-making tasks on the data to be processed through a lightweight large model on the terminal side to obtain data processing results.
[0019] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0020] at least one processor; and
[0021] a memory communicatively connected to at least one processor; wherein,
[0022] The memory stores a computer program that can be executed by at least one processor. The computer program is executed by at least one processor so that the at least one processor can execute the model processing method or data processing method of the embodiment of the present invention.
[0023] According to another aspect of the present invention, a computer-readable storage medium is provided, which stores computer instructions for implementing the model processing method or data processing method of an embodiment of the present invention when executed by a processor.
[0024] According to another aspect of the present invention, a computer program product is provided, comprising a computer program, and when the computer program is executed by a processor, the steps in the above method are implemented.
[0025] The technical solution of the embodiment of the present invention can restore the accuracy of the large model after quantization, avoiding the problem of inaccurate results of executing multimodal tasks due to the low accuracy of the model after quantization. That is, the use of the solution of the present invention can more accurately execute various multimodal tasks.
[0026] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0028] Figure 1 It is a flow chart of a model processing method provided by an embodiment of the present invention;
[0029] Figure 2a is a flow chart of another model processing method provided by an embodiment of the present invention;
[0030] Figure 2b It is a schematic diagram of a submodule of a module to be modified that a low-rank adaptive structure provided by an embodiment of the present invention is accessed;
[0031] Figure 3 It is a flowchart of a data processing method provided by an embodiment of the present invention;
[0032] Figure 4 is a structural schematic diagram of a model processing device provided by an embodiment of the present invention;
[0033] Figure 5 is a structural schematic diagram of a data processing device provided by an embodiment of the present invention;
[0034] Figure 6It is a structural schematic diagram of an electronic device for implementing the model processing method or data processing method of an embodiment of the present invention. DETAILED DESCRIPTION
[0035] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0036] The solution of the present invention is to connect the newly designed low-rank adaptive structure to the quantized large model on the basis of the newly designed low-rank adaptive structure, and fine-tune the weight of the connected low-rank adaptive structure to complete the accuracy recovery of the quantized large model. Specifically, a low-rank adaptive structure including a trainable low-rank matrix and an activation function is connected to a specific module or a specific layer (such as an attention layer) of the quantized large model. During training and fine-tuning, only the weight of the connected low-rank matrix is updated, while other parameters of the quantized large model remain unchanged. The specific implementation process can be found in the following embodiments.
[0037] Embodiment 1
[0038] Figure 1 A flowchart of a model processing method provided in an embodiment of the present invention. This embodiment is applicable to the scenario of restoring the accuracy of a large model after quantization. The method can be executed by a model processing device. The model processing device can be implemented in the form of hardware and / or software. The model processing device can be configured in an electronic device.
[0039] like Figure 1 As shown, the model processing method includes:
[0040] S101. Obtain a pre-trained multimodal large model, and quantize the multimodal large model to obtain a first lightweight large model.
[0041] In the embodiment of the present invention, the multimodal large model refers to a large model obtained by training multimodal information such as text, image, video, audio, etc.
[0042] In an embodiment of the present invention, model lightweight methods such as weight pruning, parameter quantization, and distillation training can be used to quantize the trained multimodal large model to obtain a first lightweight large model with far fewer parameters than the multimodal large model.
[0043] Take the distillation training method to quantize the trained multimodal large model as an example. The basic idea of distillation training is to use a pre-trained large teacher model to guide the training of a student model with fewer parameters. When quantizing, the trained multimodal large model can be used as the teacher model, and a new student model with far fewer parameters than the multimodal large model can be constructed; the training data is input into the teacher model, and the teacher model will output the calculation results; the same training data is input into the student model, and the student model will also output the calculation results; the first loss is calculated based on the calculation results output by the teacher model and the calculation results output by the student model; the second loss is calculated based on the calculation results output by the student model and the true label corresponding to the training data; the total loss is determined based on the first loss and the second loss; the gradient is calculated based on the total loss through the back propagation method; the student model is updated based on the calculated gradient; and finally the trained student model is used as the first lightweight large model.
[0044] S102. Using a pre-built low-rank adaptive structure, the model structure of the first lightweight large model is modified to obtain a second lightweight large model.
[0045] In an embodiment of the present invention, the low-rank adaptive structure is composed of a first low-rank matrix, an activation function, and a second low-rank matrix connected in series. It should be noted that, compared to a low-rank adaptive structure that only includes two low-rank matrices connected in series, the present invention creatively adds a nonlinear activation function (such as a ReLU function) between the two low-rank matrices, so that the newly constructed low-rank adaptive structure can better capture the complex features in the data, thereby ensuring that the model connected to the low-rank adaptive structure can learn more complex function mappings, which helps to improve the calculation accuracy of the model.
[0046] In an embodiment of the present invention, the first lightweight model can be pre-divided into multiple modules according to functions, and each module can include at least one submodule. Exemplarily, the first lightweight model is a Transformer model, and the first lightweight model can be divided into an input embedding module, a self-attention module, a feedforward network module, and an output layer module; wherein the input embedding module can include a text embedding submodule, an image embedding submodule, etc.; the self-attention module can include a query (Query) calculation submodule, a key (Key) calculation submodule, and a value (Value) calculation submodule; the feedforward network module can include an up-projection (up_proj) submodule, a down-projection (down_proj) submodule, etc. It should be noted that each submodule can be composed of one or more layers of networks in the model. On this basis, when the model structure of the first lightweight model is modified, each module of the first lightweight model can be modified or some modules of the first lightweight model can be modified. When any module is modified, the submodules included in the module are first determined, and for any submodule, the constructed low-rank adaptive structure is connected in parallel to the submodule. Through such modification, a second lightweight model can be obtained.
[0047] S103. Perform training and fine-tuning on the weights of the first low-rank matrix and the second low-rank matrix in the second lightweight model, and use the trained and fine-tuned second lightweight model as an end-side lightweight model deployed in an end-side device to perform multimodal tasks.
[0048] In an optional implementation, first obtain sample data in the training set, and input the sample data into the multimodal large model and the second lightweight large model. Then, based on the output data of the multimodal large model and the output data of the second lightweight large model, determine the distillation loss between the multimodal large model and the second lightweight large model. Specifically, the output loss can be calculated using a preset loss function based on the output data of the multimodal large model and the output data of the second lightweight large model; then, using the preset loss function, calculate the loss between the output data of the second lightweight large model and the sample label. Further, perform a weighted summation on the two losses to obtain the distillation loss. Finally, according to the distillation loss, adjust the weights of the first low-rank matrix and the second low-rank matrix in the second lightweight large model. For example, based on the distillation loss, use the back-propagation algorithm to calculate the gradient; and use the stochastic gradient descent method to update the weights of the first low-rank matrix and the second low-rank matrix in the second lightweight large model.
[0049] It should be noted that when fine-tuning the second lightweight model, only the weights of the newly connected first / second low-rank matrices are adjusted, and the other weight parameters of the second lightweight model remain unchanged, so that the second lightweight model can learn the knowledge of specific tasks through the low-rank matrix without changing the main weight of the model. In addition, it can be understood that by introducing the trainable first low-rank matrix, activation function and second low-rank matrix in the second lightweight model, additional trainable parameters are introduced to the model. These parameters can learn more detailed patterns and features in the data, thereby enhancing the model's adaptability to complex tasks. For example, when processing multimodal data, the low-rank matrix can capture the complex correlation between different modalities, allowing the model to better integrate information and thus improve accuracy.
[0050] In an embodiment of the present invention, the second lightweight large model after training and fine-tuning can be used as a terminal-side lightweight large model deployed in a terminal-side device to perform multimodal tasks; wherein the multimodal tasks may include at least one of content generation tasks (such as image generation tasks, text generation tasks), understanding and question-answering tasks (such as visual question-answering tasks), analysis and classification tasks (such as image classification and description tasks), and auxiliary decision-making tasks (such as navigation path planning tasks).
[0051] The embodiment of the present invention can restore the accuracy of the large model after quantization, avoiding the problem of inaccurate results of multimodal tasks due to the low accuracy of the model after quantization. That is, the use of the embodiment of the present invention can more accurately execute various multimodal tasks.
[0052] Embodiment 2
[0053] Figure 2a A flow chart of a model processing method provided by an embodiment of the present invention. Figure 2a , the method comprises the following steps:
[0054] S201. Obtain a pre-trained multimodal large model, and quantize the multimodal large model to obtain a first lightweight large model.
[0055] In the embodiment of the present invention, in order to transform the model structure of the first lightweight large model, the present invention creatively constructs a new low-rank adaptive structure, which is composed of a first low-rank matrix, an activation function, and a second low-rank matrix in series. On this basis, the process of transforming the model structure of the first lightweight large model using the pre-constructed low-rank adaptive structure can be seen in steps S202-S203.
[0056] S202. Determine at least one module to be modified from the model structure of the first lightweight large model; wherein any module to be modified includes at least one submodule.
[0057] In an embodiment of the present invention, the first lightweight model can be pre-divided into multiple modules according to functions, and each module can include at least one submodule. Exemplarily, the first lightweight model is a Transformer model, and the first lightweight model can be divided into an input embedding module, a self-attention module, a feedforward network module, and an output layer module; wherein the input embedding module can include a text embedding submodule, an image embedding submodule, etc.; the attention module can include a query calculation submodule, a key calculation submodule, and a value calculation submodule; the feedforward network module can include an up-projection submodule, a down-projection submodule, etc.
[0058] In the embodiment of the present invention, considering that each module of the first lightweight large model is modified, the model structure of the first lightweight large model becomes complicated. Therefore, only part of the modules of the first lightweight large model are modified. Optionally, the modules to be modified in the first lightweight large model can be determined by means of an ablation experiment. In specific implementation, for each module in the first quantized large model, the low-rank adaptive structure is sequentially connected to each sub-module included in each module in parallel to obtain a third quantized large model; the third quantized large model is subjected to an ablation experiment to obtain an ablation experiment result; wherein the ablation experiment result includes the degree of reduction in device computing power consumption and the extent of the decline in model performance after deleting the low-rank adaptive structure of one or more modules in the third quantized large model; according to the computing power limit of the end-side device where the end-side lightweight large model is to be deployed and the pre-set model performance degradation threshold, combined with the ablation experiment results, at least one module to be modified is determined from the model structure of the first lightweight large model.
[0059] S203. Connect the low-rank adaptive structure to each sub-module of the module to be modified in turn in parallel to obtain a second lightweight large model.
[0060] For example, see Figure 2b , which shows a schematic diagram of a submodule of a module to be modified in which a low-rank adaptive structure is connected. Figure 2b It can be seen that after the low-rank adaptive structure is connected to each submodule of the module to be modified in parallel, the final output of any submodule is equal to the original output of the submodule plus the output result of the low-rank adaptive structure connected in parallel to the submodule. That is, when the hidden layer variable enters a submodule, it is simultaneously input into the low-rank adaptive structure connected in parallel to the submodule; the submodule obtains the original output according to the input hidden layer variable; and when the hidden layer variable passes through the low-rank adaptive structure, the hidden layer variable is first multiplied by the first low-rank matrix, the multiplication result is acted upon by the activation function, and the result of the action is multiplied by the second low-rank matrix; the final product result is added to the original output of the submodule as the final output of this submodule.
[0061] S204. Perform training and fine-tuning on the weights of the first low-rank matrix and the second low-rank matrix in the second lightweight model, and use the trained and fine-tuned second lightweight model as an end-side lightweight model deployed in an end-side device to perform multimodal tasks.
[0062] In an embodiment of the present invention, some modules of the first lightweight large model can be modified to avoid the complexity of the model structure due to the modification of all modules; and this scheme can restore the accuracy of the large model after quantization to avoid the problem of inaccurate results of the executed multimodal tasks due to the low accuracy of the model after quantization.
[0063] Embodiment 3
[0064] Figure 3 A flowchart of a data processing method is provided for an embodiment of the present invention. This embodiment is applicable to scenarios where a multimodal task is performed using a lightweight large model on the end side. The method can be performed by a data processing device, which can be implemented in the form of hardware and / or software, and can be configured in an electronic device. Figure 3 , the data processing method comprises the following steps:
[0065] S301. Obtain data to be processed, and input the data to be processed into a lightweight large model on the end side.
[0066] In the embodiment of the present invention, the lightweight large model of the end side is obtained by the model processing method disclosed in the above embodiment; the specific process can be referred to the description of the above embodiment, which will not be repeated here.
[0067] In this embodiment, the data to be processed may be text data, image data, video data, point cloud data, etc., which are not specifically limited here. The data to be processed may be obtained from a specified data source or through a network, which are not specifically limited here.
[0068] S302: Using the terminal-side lightweight large model, perform at least one multimodal task including content generation tasks, understanding and question-answering tasks, analysis and classification tasks, and auxiliary decision-making tasks on the data to be processed to obtain data processing results.
[0069] In the embodiment of the present invention, the content generation task includes a text generation task, an image generation task, a video generation task, etc. Among them, performing a text generation task on the data to be processed means that the lightweight large model on the end generates natural language text according to the input image, audio and other modal data to be processed. For example, the lightweight large model on the end generates a beautiful descriptive text based on the input landscape picture, or generates a news report, a story, etc. based on the input video content. It can be understood that the generation result of the model is the data processing result. Similarly, performing an image generation task on the data to be processed means that the lightweight large model on the end generates a corresponding image based on the input text description data (i.e., the data to be processed). For example, the lightweight large model on the end generates an image that meets the description based on the input text description such as "a cute kitten playing on the grass". At this time, the generated image is the data processing result. Similarly, performing a video generation task on the data to be processed means that the lightweight large model on the end generates video content, including scene layout, character actions, dialogue, etc., based on the input script and story outline (i.e., the data to be processed). At this time, the generated video is the data processing result.
[0070] Understanding and question-answering tasks can include visual question-answering tasks and emotion recognition tasks. On this basis, performing visual question-answering tasks on the data to be processed means that the lightweight large model on the end understands the input image or video (data to be processed) and answers questions related to it. At this time, the answer content is the data processing result. Similarly, performing emotion recognition tasks on the data to be processed means that the lightweight large model on the end recognizes the emotional state of a person based on the input image, text, and voice data to be processed, and uses the recognized emotional state as the data processing result.
[0071] Analysis and classification tasks can include image classification tasks and video analysis tasks. On this basis, performing image classification tasks on the data to be processed means that the lightweight large model on the end classifies the input image (i.e., the data to be processed) into different categories, and can generate corresponding text descriptions based on the image content. At this time, the text description and image category are the data processing results. Similarly, performing video analysis on the data to be processed means that the lightweight large model on the end performs character tracking, action recognition, and scene recognition processing on the input video content (i.e., the data to be processed), and uses the tracking results and recognition results as the data processing results.
[0072] Decision-making support tasks can include navigation path planning tasks. On this basis, performing navigation path planning tasks on the data to be processed means that the lightweight large model on the end side can accurately perceive the vehicle's surrounding environment based on the data collected by the laser radar and camera in the vehicle (i.e., the data to be processed), so as to plan the path. The final data processing result is the planned driving path.
[0073] In the embodiment of the present invention, after the processed terminal-side lightweight large model is deployed on the terminal-side device, various multimodal tasks can be accurately executed.
[0074] Embodiment 4
[0075] Figure 4 Schematic diagram of a model processing device provided by an embodiment of the present invention. This embodiment is applicable to the scenario of restoring the accuracy of a large model after quantization. Figure 4 As shown, the model processing device includes:
[0076] A quantization module 401 is used to obtain a pre-trained multimodal large model and perform quantization processing on the multimodal large model to obtain a first lightweight large model;
[0077] A transformation module 402 is used to transform the model structure of the first lightweight large model by using a pre-constructed low-rank adaptive structure to obtain a second lightweight large model; wherein the low-rank adaptive structure is composed of a first low-rank matrix, an activation function, and a second low-rank matrix connected in series;
[0078] The fine-tuning module 403 is used to train and fine-tune the weights of the first low-rank matrix and the second low-rank matrix in the second lightweight model, and use the trained and fine-tuned second lightweight model as an end-side lightweight model deployed in an end-side device to perform multimodal tasks.
[0079] In some embodiments, the transformation module 402 includes:
[0080] A screening unit, used to determine at least one module to be modified from the model structure of the first lightweight large model; wherein any module to be modified includes at least one submodule;
[0081] The transformation unit is used to connect the low-rank adaptive structure to each sub-module of the module to be modified in parallel, so as to obtain a second lightweight large model.
[0082] In some embodiments, the fine-tuning module 403 is specifically used to:
[0083] Obtain sample data in the training set, and input the sample data into the multimodal large model and the second lightweight large model;
[0084] Determine a distillation loss between the multimodal large model and the second lightweight large model according to the output data of the multimodal large model and the output data of the second lightweight large model;
[0085] According to the distillation loss, the weights of the first low-rank matrix and the second low-rank matrix in the second lightweight large model are adjusted.
[0086] In some embodiments, after the low-rank adaptive structure is connected in parallel to each sub-module of the module to be modified, the final output of any sub-module is equal to the original output of the sub-module plus the output result of the low-rank adaptive structure connected in parallel to the sub-module.
[0087] In some embodiments, the model processing device further comprises:
[0088] The ablation pre-processing module is used for sequentially connecting the low-rank adaptive structure to each sub-module included in each module in the first quantized large model in parallel to obtain a third quantized large model;
[0089] an ablation module, used to perform an ablation experiment on the third large quantized model to obtain an ablation experiment result; wherein the ablation experiment result includes a reduction degree of device computing power consumption and a reduction degree of model performance after deleting the low-rank adaptive structure of one or more modules in the third large quantized model;
[0090] Accordingly, the screening unit is specifically used for:
[0091] According to the computing power limitation of the terminal device and the pre-set model performance degradation threshold, combined with the ablation experiment results, at least one module to be modified is determined from the model structure of the first lightweight large model.
[0092] The model processing device provided in the embodiment of the present invention can execute the model processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0093] Embodiment 5
[0094] Figure 5 Schematic diagram of a data processing device provided by an embodiment of the present invention. This embodiment is applicable to scenarios where a lightweight large model on the end side is used to perform multimodal tasks. Figure 5 As shown, the data processing device includes:
[0095] A data acquisition module 501 is used to acquire data to be processed and input the data to be processed into a lightweight large model on the end side; wherein the lightweight large model on the end side is obtained by processing according to any one of the methods in claims 1-5;
[0096] The data processing module 502 is used to perform at least one multimodal task including content generation tasks, understanding and question-answering tasks, analysis and classification tasks, and auxiliary decision-making tasks on the data to be processed through a lightweight large model on the terminal side to obtain a data processing result.
[0097] The data processing device provided in the embodiment of the present invention can execute the data processing method provided in any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0098] According to an embodiment of the present invention, the present invention also provides an electronic device, a readable storage medium and a computer program product.
[0099] Embodiment 6
[0100] Figure 6 The components, their connections and relationships, and their functions shown herein are merely examples and are not intended to limit implementation of the invention described and / or claimed herein.
[0101] like Figure 6 As shown, the electronic device 10 includes at least one processor 11, and a memory connected to the at least one processor 11, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., wherein the memory stores a computer program that can be executed by at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 to the random access memory (RAM) 13. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0102] A number of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0103] The processor 11 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as executing a model processing method or a data processing method.
[0104] In some embodiments, the model processing method or the data processing method may be implemented as a computer program, which is tangibly contained in a computer-readable storage medium, such as a storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the model processing method or the data processing method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to execute the model processing method or the data processing method in any other appropriate manner (e.g., by means of firmware).
[0105] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0106] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable model processing device, so that when the computer program is executed by the processor, the functions / operations specified in the flow chart and / or block diagram are implemented. The computer program may be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0107] In the context of the present invention, a computer-readable storage medium may be a tangible medium that may contain or store a computer program for use by or in combination with an instruction execution system, device or equipment. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0108] To provide interaction with a user, the systems and techniques described herein may be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user may be received in any form (including acoustic input, voice input, or tactile input).
[0109] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system may be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0110] A computing system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The client and server relationship is generated by computer programs running on the corresponding computers and having a client-server relationship with each other. The server may be a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of difficult management and weak business scalability in traditional physical hosts and VPS services.
[0111] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps described in the present invention can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solution of the present invention can be achieved, and this document does not limit this.
[0112] The above specific implementations do not constitute a limitation on the protection scope of the present invention. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present invention should be included in the protection scope of the present invention.
Claims
1. A model processing method, characterized in that: include: Obtain a pre-trained multimodal large model, and quantize the multimodal large model to obtain a first lightweight large model; Using a pre-constructed low-rank adaptive structure, the model structure of the first lightweight large model is modified to obtain a second lightweight large model; wherein the low-rank adaptive structure is composed of a first low-rank matrix, an activation function, and a second low-rank matrix connected in series; The weights of the first low-rank matrix and the second low-rank matrix in the second lightweight model are trained and fine-tuned, and the second lightweight model after training and fine-tuning is used as an end-side lightweight model deployed in an end-side device to perform multimodal tasks.
2. The method according to claim 1, characterized in that The method of using the pre-built low-rank adaptive structure to transform the model structure of the first lightweight large model to obtain the second lightweight large model includes: Determine at least one module to be modified from the model structure of the first lightweight large model; wherein any of the modules to be modified includes at least one submodule; The low-rank adaptive structure is sequentially connected to each submodule of the module to be modified in parallel to obtain a second lightweight large model.
3. The method according to claim 1, characterized in that The training and fine-tuning of the weights of the first low-rank matrix and the second low-rank matrix in the second lightweight large model includes: Obtaining sample data in a training set, and inputting the sample data into the multimodal large model and the second lightweight large model; Determining a distillation loss between the multimodal large model and the second lightweight large model according to the output data of the multimodal large model and the output data of the second lightweight large model; According to the distillation loss, weights of the first low-rank matrix and the second low-rank matrix in the second lightweight large model are adjusted.
4. The method according to claim 2, characterized in that: After the low-rank adaptive structure is connected to each sub-module of the module to be modified in parallel, the final output of any sub-module is equal to the original output of the sub-module plus the output result of the low-rank adaptive structure connected in parallel to the sub-module.
5. The method according to claim 2, characterized in that: Also includes: For each module in the first quantized large model, the low-rank adaptive structure is sequentially connected to each submodule included in each module in parallel to obtain a third quantized large model; Performing an ablation experiment on the third quantized large model to obtain an ablation experiment result; wherein the ablation experiment result includes a reduction degree of device computing power consumption and a reduction degree of model performance after deleting the low-rank adaptive structure of one or more modules in the third quantized large model; Accordingly, at least one module to be modified is determined from the model structure of the first lightweight large model, including: According to the computing power limitation of the terminal device and a preset model performance degradation threshold, combined with the ablation experiment results, at least one module to be modified is determined from the model structure of the first lightweight large model.
6. A data processing method, characterized in that: include: Acquire the data to be processed, and input the data to be processed into the lightweight large model on the end side; wherein the lightweight large model on the end side is obtained by processing according to any one of the methods in claims 1-5; Through the terminal-side lightweight large model, at least one multimodal task including content generation tasks, understanding and question-answering tasks, analysis and classification tasks, and auxiliary decision-making tasks is performed on the data to be processed to obtain a data processing result.
7. A model processing device, characterized in that: include: A quantization module, used to obtain a pre-trained multimodal large model and quantize the multimodal large model to obtain a first lightweight large model; A transformation module, used to transform the model structure of the first lightweight large model by using a pre-constructed low-rank adaptive structure to obtain a second lightweight large model; wherein the low-rank adaptive structure is composed of a first low-rank matrix, an activation function, and a second low-rank matrix connected in series; A fine-tuning module is used to train and fine-tune the weights of the first low-rank matrix and the second low-rank matrix in the second lightweight large model, and use the trained and fine-tuned second lightweight large model as an end-side lightweight large model deployed in an end-side device to perform multimodal tasks.
8. A data processing device, characterized in that: include: A data acquisition module, used for acquiring data to be processed and inputting the data to be processed into a lightweight large model on the end side; wherein the lightweight large model on the end side is obtained by processing according to any one of the methods in claims 1-5; The data processing module is used to perform at least one multimodal task including content generation tasks, understanding and question-answering tasks, analysis and classification tasks, and auxiliary decision-making tasks on the data to be processed through the terminal-side lightweight large model to obtain data processing results.
9. An electronic device, characterized in that: include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can perform the method according to any one of claims 1 to 6.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the method according to any one of claims 1 to 6 when executed.
Citation Information
Patent Citations
Data processing method, device and equipment, readable storage medium and program product
CN118211140A
End-to-end learning method, system and equipment based on multi-modal large model
CN118211643A
Log anomaly detection method based on efficient fine tuning of adaptive low-rank parameters
CN118260689A
Large language model acceleration method and device
CN118569324A
Model training method and device, equipment and storage medium
CN118643323A