Content processing method and device, electronic equipment and storage medium
By loading the trainingable matrix matching task type to update the mapping layer of the content processing model and retrieving the pre-cache reference state, the problem of different task types in the prior art requiring specialized models is solved, and low-cost and high-efficiency content processing is achieved.
Patent Information
- Application Number
- CN202510262549.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-06
- Publication Date
- 2025-06-24
AI Technical Summary
In the prior art, processing tasks of different task types require the deployment of dedicated processing models, resulting in high development costs and low response efficiency.
By obtaining input content and task type information, the trainable matrix matching the mapping layer of the content processing model, the original weight matrix of the mapping layer is updated, and the pre-cached reference state is retrieved to achieve fast mapping and processing result output.
It implements the use of a content processing model to process multiple task types, reduces development costs and improves the response efficiency of the content processing model.
Smart Images

Figure CN120196792A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technology, and in particular to a content processing method, device, electronic device and storage medium. Background Art
[0002] With the development of artificial intelligence technology, machine learning models are widely used in various tasks. For example, in a generation task, the input content is provided to the model that processes the generation task, and the corresponding prediction content can be generated. For example, in a classification task, the input content is provided to the model that processes the classification task, and the corresponding classification results can be predicted. However, in the related technology, processing tasks of different task types require the deployment of dedicated processing models for processing, which has high development costs, and the response efficiency of traditional processing models is low. Summary of the invention
[0003] The following is a summary of the subject matter of the detailed description of the present disclosure. This summary is not intended to limit the scope of the claims.
[0004] The embodiments of the present disclosure provide a content processing method, device, electronic device, and storage medium, which can reduce development costs and improve the response efficiency of the content processing model.
[0005] On the one hand, an embodiment of the present disclosure provides a content processing method, including:
[0006] Get input content and task type information;
[0007] Loading a first trainable matrix matching a mapping layer of a content processing model according to the task type information, and updating an original weight matrix of the mapping layer based on the first trainable matrix;
[0008] Retrieving a pre-cached reference state according to the input content and the first trainable matrix, wherein the reference state is generated by calling the mapping layer to map the reference content after updating the original weight matrix based on the first trainable matrix, and the reference content matches a prefix of the input content;
[0009] The updated mapping layer is called, the reference state is reused to map the input content to obtain mapping features, and the first output layer corresponding to the task type information in the content processing model is called to output a processing result based on the mapping features.
[0010] On the other hand, the present disclosure also provides a content processing device, including:
[0011] The acquisition module is used to obtain input content and task type information;
[0012] An updating module, configured to load a first trainable matrix matching a mapping layer of a content processing model according to the task type information, and update an original weight matrix of the mapping layer based on the first trainable matrix;
[0013] a retrieval module, configured to retrieve a pre-cached reference state according to the input content and the first trainable matrix, wherein the reference state is generated by calling the mapping layer to map the reference content after updating the original weight matrix based on the first trainable matrix, and the reference content matches a prefix of the input content;
[0014] The output module is used to call the updated mapping layer, reuse the reference state to map the input content to obtain mapping features, and call the first output layer corresponding to the task type information in the content processing model to output the processing result based on the mapping features.
[0015] Furthermore, the content processing model is further provided with at least one second output layer having a function similar to that of the first output layer, and the updating module is specifically used for:
[0016] Loading a second trainable matrix corresponding to the second output layer;
[0017] Based on the first trainable matrix and the second trainable matrix, an operation is performed with the original weight matrix of the mapping layer to obtain the updated original weight matrix.
[0018] Furthermore, the number of the second output layers is multiple, and the above-mentioned update module is specifically used for:
[0019] Performing an operation on the first trainable matrix and the original weight matrix of the mapping layer to obtain a first operation result;
[0020] sorting the plurality of second output layers according to the functional similarity between the second output layers and the first output layers;
[0021] Based on the sorting result, the second trainable matrix corresponding to the second output layer is iteratively operated with the first operation result in turn to obtain the updated original weight matrix.
[0022] Furthermore, the above update module is specifically used for:
[0023] Traversing the sorting results in sequence, for the first second output layer in the sorting results, determining the element-by-element multiplication result between the second trainable matrix and the first operation result, summing the element-by-element multiplication result and the first operation result to obtain the second operation result of the first round;
[0024] For any remaining second output layer in the sorting result, determine the element-by-element multiplication result between the second trainable matrix and the second operation result of the previous round, and sum the element-by-element multiplication result with the second operation result of the previous round to obtain the second operation result of the current round;
[0025] The second operation result obtained in the last round is determined as the updated original weight matrix.
[0026] Furthermore, the above update module is specifically used for:
[0027] Traversing the sorting results in sequence, for the first second output layer in the sorting results, concatenating the second trainable matrix with the first trainable matrix, performing cross-attention on the concatenated result and the first operation result, and then summing the result with the first operation result to obtain the third operation result of the first round;
[0028] For any remaining second output layer in the sorting result, concatenate all the second trainable matrices corresponding to the first second output layer to the current second output layer with the first trainable matrix, cross-attend the concatenation result with the third operation result of the previous round, and sum them with the third operation result of the previous round to obtain the third operation result of the current round;
[0029] The third operation result obtained in the last round is determined as the updated original weight matrix.
[0030] Furthermore, the above update module is specifically used for:
[0031] Using the first operation result as a query matrix and the concatenation result as a key matrix and a value matrix;
[0032] A third operation result of the first round is obtained by performing cross attention on the query matrix, the key matrix and the value matrix and summing the results with the first operation result.
[0033] Furthermore, the prefix of the input content includes prompt text and non-text prompt data, and the above retrieval module is specifically used for:
[0034] Retrieving a candidate state that matches the first trainable matrix and is pre-cached according to the first trainable matrix, wherein the candidate state is generated by calling the mapping layer to map candidate content after the original weight matrix is updated based on the first trainable matrix, and the candidate content includes candidate text and non-text candidate data;
[0035] Extracting a first data feature of the non-text prompt data, and extracting a second data feature of the non-text candidate data;
[0036] Determine the semantic matching degree between the prompt text and the candidate text, and determine the target similarity between the first data feature and the second data feature;
[0037] When the semantic matching degree is greater than or equal to a first preset threshold, and the target similarity is greater than or equal to a second preset threshold, the candidate state is determined as the retrieved reference state.
[0038] Furthermore, the number of the non-text prompt data is multiple, the number of the first data features and the number of the second data features are the same as the number of the non-text prompt data, and the above-mentioned retrieval module is specifically used for:
[0039] Determining a feature weight of the first data feature corresponding to each of the non-text prompt data according to an input order of each of the non-text prompt data in the input content;
[0040] For each of the first data features, determining an initial similarity between the first data feature and the corresponding second data feature;
[0041] The initial similarities are weighted and summed based on the feature weights to obtain the target similarity.
[0042] Furthermore, the output module is specifically used for:
[0043] Retrieving pre-cached quantization parameters according to the task type information;
[0044] The original weight matrix updated based on the first trainable matrix is quantized according to the quantization parameter, the quantized mapping layer is called, and the reference state is reused to map the input content to obtain a mapping feature.
[0045] Further, the quantization parameter includes a scaling factor and a zero point, and the output module is specifically used for:
[0046] Performing periodic mapping on the original weight matrix updated based on the first trainable matrix to obtain a periodic mapping result;
[0047] Compressing the original weight matrix to obtain a compression result, and determining a product result between the compression result and the periodic mapping result;
[0048] Performing nonlinear adjustment on the scaling factor, adjusting the scaling factor after the nonlinear adjustment according to the zero point and a preset first offset parameter to obtain a first offset result, and determining a ratio between the product result and the first offset result;
[0049] The ratio is adjusted according to the zero point and a preset second offset parameter to obtain a second offset result, and the second offset result is rounded to obtain a quantization result of the original weight matrix after updating based on the first trainable matrix.
[0050] Furthermore, the output module is specifically used for:
[0051] For each weight value of the original weight matrix, determining a summation result between an absolute value of the weight value and a preset offset;
[0052] The summation result is input into a logarithmic function for operation to obtain a compression result.
[0053] On the other hand, an embodiment of the present disclosure further provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the above-mentioned content processing method when executing the computer program.
[0054] On the other hand, an embodiment of the present disclosure further provides a computer-readable storage medium, wherein the storage medium stores a computer program, and the computer program is executed by a processor to implement the above-mentioned content processing method.
[0055] On the other hand, the embodiment of the present disclosure further provides a computer program product, which includes a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the above-mentioned content processing method.
[0056] The disclosed embodiment includes at least the following beneficial effects: by obtaining input content and task type information, and then loading a first trainable matrix matching the mapping layer of the content processing model according to the task type information, since the task type information can characterize the current task type, it is equivalent to loading the first trainable matrix dedicated to the current task type, and then updating the original weight matrix of the mapping layer based on the first trainable matrix, so that the mapping layer can better adapt to the current task type, and for different task type information, it is only necessary to load the corresponding first trainable matrix based on the task type information, thereby realizing the use of one content processing model to process processing tasks of multiple task types, without the need to deploy multiple dedicated models, and effectively reducing development costs, and then retrieving the pre-cached reference state according to the input content and the first trainable matrix, since the mapping layer respectively Before mapping the reference content and the input content, the original weight matrix is updated based on the same first trainable matrix, and the prefixes of the reference content and the input content are matched, so the mapping result of the prefix of the input content is the same or similar to the reference state. Therefore, the retrieved reference state can be used as the mapping result of the prefix of the input content, and then the updated mapping layer is called, and the reference state is reused to map the input content to obtain the mapping feature, and the first output layer corresponding to the task type information in the content processing model is called to output the processing result based on the mapping feature, that is, the processing result is output based on the first output layer dedicated to the current task type, which can ensure the accuracy of the processing result, and reuse the reference state in the mapping process of the input content, without re-mapping the prefix of the input content, which can reduce repeated calculations, thereby effectively improving the response efficiency of the content processing model.
[0057] Other features and advantages of the present disclosure will be set forth in the following description, and in part will be apparent from the description, or may be learned by practicing the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The accompanying drawings are used to provide further understanding of the technical solution of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solution of the present disclosure and do not constitute a limitation on the technical solution of the present disclosure.
[0059] Figure 1 A schematic diagram of an optional implementation environment provided for an embodiment of the present disclosure;
[0060] Figure 2 A schematic diagram of an optional flow chart of a content processing method provided in an embodiment of the present disclosure;
[0061] Figure 3 An optional flowchart of inserting a first trainable matrix provided in an embodiment of the present disclosure;
[0062] Figure 4An optional flow chart of obtaining mapping features provided in an embodiment of the present disclosure;
[0063] Figure 5 An optional flowchart for determining a processing result provided in an embodiment of the present disclosure;
[0064] Figure 6 Another optional flowchart for determining a processing result provided in an embodiment of the present disclosure;
[0065] Figure 7 An optional flowchart of inserting a first trainable matrix and a second trainable matrix provided in an embodiment of the present disclosure;
[0066] Figure 8 A schematic diagram of an optional process for updating an original weight matrix provided in an embodiment of the present disclosure;
[0067] Fig. 9 Another optional flow chart of obtaining mapping features provided in the embodiment of the present disclosure;
[0068] Fig.10 A schematic diagram of an optional architecture of a content processing method provided in an embodiment of the present disclosure;
[0069] Fig.11 A schematic diagram of an optional structure of a content processing device provided in an embodiment of the present disclosure;
[0070] Fig.12 A partial structural block diagram of a terminal provided in an embodiment of the present disclosure;
[0071] Fig.13 A partial structural block diagram of a server provided in an embodiment of the present disclosure. DETAILED DESCRIPTION
[0072] In order to make the purpose, technical solution and advantages of the present disclosure more clear, the present disclosure is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present disclosure and are not used to limit the present disclosure.
[0073] It should be noted that in each specific implementation manner of the present disclosure, when it comes to performing relevant processing based on data related to the characteristics of the target object, such as target object attribute information or a set of attribute information, the permission or consent of the target object will be obtained first. Moreover, the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. Among them, the target object can be a user. In addition, when the embodiments of the present disclosure need to obtain the target object attribute information, the separate permission or separate consent of the target object will be obtained by means of a pop-up window or jumping to a confirmation page. After clearly obtaining the separate permission or separate consent of the target object, the necessary data related to the target object for the normal operation of the embodiments of the present disclosure will be obtained.
[0074] In the embodiments of the present disclosure, the term "module" or "unit" refers to a computer program with a predetermined function or a part of a computer program, which works together with other related parts to achieve a predetermined goal, and can be fully or partially implemented by using software, hardware (such as a processing circuit or a memory), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be a part of the overall module or unit that includes the function of the module or unit.
[0075] With the development of artificial intelligence technology, machine learning models are widely used in various tasks. For example, in a generation task, providing input content to a model for processing the generation task can generate corresponding predicted content. Another example is that in a classification task, providing input content to a model for processing the classification task can predict the corresponding classification result. However, in the related art, dedicated processing models need to be deployed for processing tasks of different task types, resulting in a high development cost, and the response efficiency of traditional processing models is relatively low.
[0076] Based on this, the embodiments of the present disclosure provide a content processing method, apparatus, electronic device, and storage medium, which can reduce the development cost and improve the response efficiency of the content processing model.
[0077] Refer to Figure 1 , Figure 1 which is a schematic diagram of an optional implementation environment provided by the embodiments of the present disclosure. This implementation environment includes a terminal 101 and a server 102. Among them, the terminal 101 and the server 102 are connected through a communication network.
[0078] Exemplarily, the server 102 can obtain the input content and task type information sent by the terminal 101; load the first trainable matrix matching the mapping layer of the content processing model according to the task type information, and update the original weight matrix of the mapping layer based on the first trainable matrix; retrieve the pre-cached reference state according to the input content and the first trainable matrix, where the reference state is generated by calling the mapping layer to map the reference content after updating the original weight matrix based on the first trainable matrix, and the reference content matches the prefix of the input content; call the updated mapping layer, reuse the reference state to map the input content to obtain mapping features, and call the first output layer corresponding to the task type information in the content processing model to output a processing result based on the mapping features; the server 102 sends the processing result to the terminal 101.
[0079] By obtaining the input content and task type information, and then loading the first trainable matrix matching the mapping layer of the content processing model according to the task type information. Since the task type information can represent the current task type, it is equivalent to loading the first trainable matrix dedicated to the current task type. Then, update the original weight matrix of the mapping layer based on the first trainable matrix, enabling the mapping layer to better adapt to the current task type. Moreover, for different task type information, only need to load the corresponding first trainable matrix based on the task type information, which realizes the processing task of using one content processing model to handle multiple task types, without the need to deploy multiple dedicated models, and can effectively reduce the development cost. Then, retrieve the pre-cached reference state according to the input content and the first trainable matrix. Since the original weight matrix is updated based on the same first trainable matrix before the mapping layer maps the reference content and the input content respectively, and the reference content matches the prefix of the input content, the mapping result of the prefix of the input content is the same or similar to the reference state. Therefore, the retrieved reference state can be used as the mapping result of the prefix of the input content. Then, call the updated mapping layer, reuse the reference state to map the input content to obtain mapping features, and call the first output layer corresponding to the task type information in the content processing model to output a processing result based on the mapping features, that is, output the processing result based on the first output layer dedicated to the current task type, which can ensure the accuracy of the processing result. Moreover, reuse the reference state during the mapping process of the input content, without having to remap the prefix of the input content, which can reduce repeated calculations and thus effectively improve the response efficiency of the content processing model.
[0080] The server 102 can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. Additionally, the server 102 can also be a node server in a blockchain network.
[0081] The terminal 101 can be a mobile phone, a computer, an intelligent voice interaction device, an intelligent home appliance, a vehicle-mounted terminal, etc., but is not limited thereto. The terminal 101 and the server 102 can be directly or indirectly connected through wired or wireless communication means, and the embodiments of the present disclosure do not limit this here.
[0082] Referring to Figure 2 , Figure 2 FIG. is an optional flowchart of the content processing method provided by the embodiments of the present disclosure. This content processing method can be executed by the server, or can also be executed by the terminal, or can also be executed by the server in cooperation with the terminal. This content processing method includes but is not limited to the following steps 201 to step 204.
[0083] Step 201: Obtain the input content and the task type information.
[0084] Among them, the input content refers to the content that is input into the content processing model and processed by the content processing model. The input content can be understood by the content processing model and generate an appropriate response through reasoning. The input content can be described in detail from the following aspects.
[0085] The source of the input content can be diversified. For example, the input content can be from user input, data collected by sensors, or pre-stored content. Exemplarily, a prompt instruction can be pre-stored, and subsequently, the prompt instruction can be directly used as the input content or a part of the input content. Suppose the question text input by the user is "Please explain the large language model", and the pre-stored prompt instruction is "Think and answer from the perspective of non-professionals", then the concatenation result of the prompt instruction and the question text can be used as the input content.
[0086] The form of the input content can be diversified. For example, the input content can include one or more types of modal data. Exemplarily, the input content can include one or more of text modal data, image modal data, and audio modal data.
[0087] The scale of the input content can be diverse. For example, the input content can be a long text or a short text, the input content can also be an image or multiple images, and the input content can also be a combination of a long text and multiple images.
[0088] Among them, the task type information is used to characterize the task type of the processing task of the input content. Specifically, unique task type information can be configured for different task types, so that different task types can be distinguished through different task type information. There can be multiple implementation manners for configuring the task type information, which are not limited in the embodiments of the present disclosure.
[0089] For example, since the processing task can include a generation task for predicting the next token and a classification task for predicting the class probability distribution, the task type of the processing task can be simply classified into a generation task type or a classification task type. At this time, the task type information can include generation task information and classification task information. For example, the generation task information can be configured as [GEN], and the classification task information can be configured as [CLS]. By [GEN], it is characterized that the task type of the processing task of the input content is the generation task type, and by [CLS], it is characterized that the task type of the processing task of the input content is the classification task type.
[0090] For another example, the task type can have a two-level structure. In the first level of the task type, the processing task can be classified into a generation task type or a classification task type. In the second level of the task type, according to the specific task requirements of the processing task, the processing task can be further classified into a generation subtask type under the generation task type or a classification subtask type under the classification task type.
[0091] Exemplarily, the generated sub-task types may include translation generation sub-task types, summary generation sub-task types, etc., and the classification sub-task types may include sentiment classification sub-task types, topic classification sub-task types, etc. At this time, the task type information may include translation generation sub-task information, summary generation sub-task information, sentiment classification sub-task information, and topic classification sub-task information. The generation task information and the translation generation sub-task information may be jointly configured as [GEN_1], the generation task information and the summary generation sub-task information may be jointly configured as [GEN_2], the classification task information and the sentiment classification sub-task information may be jointly configured as [CLS_1], and the classification task information and the subject classification sub-task information may be jointly configured as [CLS_2]. Through [GEN_1], it is characterized that the task type of the processing task of the input content is the translation generation sub-task type under the generation task type. Through [GEN_2], it is characterized that the task type of the processing task of the input content is the summary generation sub-task type under the generation task type. Through [CLS_1], it is characterized that the task type of the processing task of the input content is the sentiment classification sub-task type under the classification task type. Through [CLS_2], it is characterized that the task type of the processing task of the input content is the topic classification sub-task type under the classification task type.
[0092] In addition, the task type information may also be configured as other content, as long as it can be used to distinguish the specific task type through the task type information.
[0093] In a possible implementation manner, the input content and the task type information are directly input by the user to the local electronic device, or the input content and the task type information are input by the user to other electronic devices, and then the other electronic devices send the input content and the task type information to the local electronic device.
[0094] Herein, the local electronic device refers to the electronic device that processes and deploys the content processing model, and the other electronic devices refer to the electronic devices other than the local electronic device.
[0095] It should be noted that when the other electronic devices send the input content and the task type information to the local electronic device, specifically, the other electronic devices may generate a processing request according to the input content and the task type information, so that the processing request carries the input content and the task type information, and then sends the input content and the task type information to the local electronic device through a general application programming interface (API). This realizes seamless processing of processing requests corresponding to multiple task types by designing a general API, which can simplify the system architecture, thereby reducing the usage cost and integration cost. In addition, in order to improve the concurrent processing ability, different processing requests may also be sent through different APIs, which is not limited in this embodiment of the present disclosure.
[0096] It should be noted that the task type information can be input into the local electronic device or other electronic devices in various ways. For example, the display interface of the electronic device displays check boxes including various task types. When one of the check boxes is checked by the user, the task type information is generated according to the task type corresponding to the checked check box. Another example is that when the input content is input into the electronic device by the user, the electronic device performs task type prediction on the input content, and generates task type information according to the task type prediction result.
[0097] Step 202: Load the first trainable matrix that matches the mapping layer of the content processing model according to the task type information, and update the original weight matrix of the mapping layer based on the first trainable matrix.
[0098] Among them, the content processing model is a model used to process the input content. The content processing model is provided with a mapping layer, and the mapping layer is used to map the input content. Through the mapping, more abstract and informative features can be extracted from the input content.
[0099] Among them, the weight matrix can also be called a parameter matrix. The weight matrix is used to perform a linear transformation on the input data to map the input data into a new feature space. The original weight matrix of the mapping layer can be a weight matrix learned in the pre-training stage. Specifically, it means that in the pre-training stage of the content processing model, by adjusting the original weight matrix, the original weight matrix can learn the general knowledge obtained by the model through training on a large-scale data set. For example, when the content processing model is a language model, in the pre-training stage, it can learn general knowledge such as vocabulary, grammar structure, context relationship, and common world knowledge.
[0100] In addition, the trainable matrix can be a weight matrix learned in the fine-tuning training stage. Specifically, it means that in the fine-tuning training stage of the content processing model, by adjusting the trainable matrix, the trainable matrix can learn the supplementary knowledge related to the task.
[0101] It is understandable that trainable matrices corresponding to various task types can be deployed for the mapping layer. Usually, multiple trainable matrices need to be deployed. The first trainable matrix is one of the multiple trainable matrices. During the fine-tuning training phase, each trainable matrix needs to be adjusted separately. Before adjusting any trainable matrix, the trainable matrix matching the mapping layer can be loaded separately according to the task type information, that is, the trainable matrix dedicated to the task type is determined. Then, the original weight matrix is updated based on the loaded trainable matrix, which is equivalent to inserting the dedicated trainable matrix into the mapping layer. It is also necessary to obtain the training data corresponding to the loaded trainable matrix. This training data is data related to the task type corresponding to the loaded trainable matrix. For example, if the task type corresponding to the loaded trainable matrix is the abstract generation sub-task type, the training samples are samples related to the abstract generation sub-task type. Exemplarily, the training samples include the article text and the abstract text of the article text. Then, the original weight matrix is frozen, and the loaded trainable matrix is trained based on the training samples, realizing that while keeping the general knowledge contained in the original weight matrix unchanged, the trainable matrix can learn supplementary knowledge related to the task type information, ensuring that the trainable matrix can play its advantages in the scenarios of the task types it is good at. Then, before training other trainable matrices, the trainable matrix inserted into the mapping layer needs to be unloaded, and other trainable matrices are inserted into the mapping layer.
[0102] Based on this, in the inference phase, the first trainable matrix matching the mapping layer can also be loaded according to the task type information, which is equivalent to determining the first trainable matrix from the trainable matrices, that is, determining the first trainable matrix dedicated to the current task type. Then, the original weight matrix of the mapping layer is updated based on the first trainable matrix, which is equivalent to combining the first trainable matrix with the original weight matrix, realizing the combination of the supplementary knowledge related to the task type information contained in the first trainable matrix and the general knowledge contained in the original weight matrix, being able to play the advantages of the first trainable matrix, making the mapping layer with both general knowledge and supplementary knowledge better adapt to the current task type. The mapping layer can focus on and retain the key information related to the task type in the input content, thereby improving the processing effect of the content processing model. Moreover, for different task type information, only the corresponding first trainable matrix needs to be loaded based on the task type information, realizing the processing task of using one content processing model to process multiple task types, without the need to deploy multiple dedicated models, being able to effectively reduce the development cost, system complexity, and maintenance cost. Moreover, when adding or reducing task types, only the management component of the first trainable matrix needs to be independently updated, being able to improve the maintainability and scalability of the system.
[0103] Exemplarily, refer to Figure 3 , Figure 3An alternative process diagram for inserting the first trainable matrix provided by the embodiments of the present disclosure.
[0104] Among them, according to the task type information, the first trainable matrix matching the mapping layer is loaded from multiple trainable matrices, and then the original weight matrix of the mapping layer is updated based on the first trainable matrix. For example, the first trainable matrix is added to the original weight matrix, or the first trainable matrix is weighted and added to the original weight matrix, so as to realize inserting the first trainable matrix into the mapping layer. Equivalent to developing an efficient trainable matrix switching mechanism, it can dynamically select and apply an appropriate first trainable matrix at the granularity of a single request, realize the fast loading and unloading of the first trainable matrix, minimize the switching overhead. In addition, data between different tasks can also be isolated to ensure that data between different tasks does not interfere with each other.
[0105] In a possible implementation, under normal circumstances, the model parameters after training of the mapping layer can be regarded as the sum result of the original parameter matrix and the parameter update amount matrix. At this time, the parameter update amount matrix can be approximately expressed by the first low-rank decomposition matrix, and the first low-rank decomposition matrix can be used as the first trainable matrix. Therefore, the first trainable matrix can be expressed as the product of two low-rank matrices. Based on this, in the fine-tuning training stage, by freezing the original weight matrix and adjusting the two low-rank matrices of the first trainable matrix, the number of training parameters can be effectively reduced, thereby effectively reducing the time and space of model training.
[0106] The principle of reducing the number of training parameters is as follows: Assume that the dimension of the original parameter matrix W0 is N×N, where N is a positive integer. The dimension of the first low-rank matrix A can be N×d, and the dimension of the second low-rank matrix B can be d×N. The dimension of the first trainable matrix obtained by multiplying the two low-rank matrices is N×N, which is the same as the dimension of the original parameter matrix W0 and does not change the dimension of the output data. It can be seen that when fully adjusting the parameter update amount matrix of the mapping layer, the number of adjusted parameters is N*N, while when adjusting the first trainable matrix, the number of adjusted parameters is N*d + d*N. Since d is much smaller than N, where d is the rank of the first trainable matrix, the number of adjusted parameters is reduced from N*N to 2*N*d. It can be seen that the number of adjusted parameters can be effectively reduced.
[0107] Step 203: Retrieve the pre-cached reference state according to the input content and the first trainable matrix.
[0108] Among them, the reference state is generated by calling the mapping layer to map the reference content after updating the original weight matrix based on the first trainable matrix, and the reference content matches the prefix of the input content.
[0109] It can be understood that since the reference state is generated by mapping the reference content through the mapping layer after updating the original weight matrix based on the first trainable matrix, the reference content matches the first trainable matrix. Also, since the input content matches the first trainable matrix, the reference content and the input content correspond to the same task type.
[0110] Specifically, the reference content is the content that is processed before the content processing model processes the input content. The task type information corresponding to the reference content is the same as the task type information corresponding to the input content. Similar to the process of processing the input content, when processing the reference content, it is also necessary to load the first trainable matrix according to the task type information, update the original weight matrix of the mapping layer based on the first trainable matrix, and then call the mapping layer to map the reference content, which can generate the reference state of the reference content. Then, the reference state is cached, and the cached reference state is associated with the first trainable matrix loaded when processing the reference content. Therefore, the pre-cached reference state can be retrieved according to the input content and the first trainable matrix.
[0111] It should be noted that before the content processing model processes the input content, whenever the original weight matrix is updated based on the trainable matrix and the mapping layer is called to map the input historical content, the intermediate state generated by mapping the historical prefix based on the historical content can be pre-cached. At the same time, the intermediate state is associated with the corresponding historical prefix and the corresponding trainable matrix; among them, the target length of the historical prefix can be configured according to actual needs. For example, a fixed target length or a dynamically adjustable target length can be configured. Specifically, the target length can be configured according to the system performance. When the length of the historical content is less than or equal to the target length, the historical content is used as the historical prefix. When the length of the historical content is greater than the target length, the part of the historical content within the target window is used as the historical prefix, where the length of the target window is the target length, and the starting position of the target window is at the beginning of the historical content.
[0112] Then, when the content processing model processes the input content, the original weight matrix is updated based on the current first trainable matrix, and then the reference state cached in advance is retrieved according to the input content and the first trainable matrix. Actually, the reference state is retrieved from multiple cached intermediate states in advance. The retrieval process of the reference state can be hierarchical. For example, first, the candidate state is retrieved from multiple intermediate states according to the first trainable matrix. It is required that the trainable matrix associated with the retrieved candidate state is the same as the current first trainable matrix, which is equivalent to retrieving the candidate state associated with the current task type first, and then the reference state is retrieved from the candidate state according to the input content. It is required that the historical prefix associated with the retrieved reference state matches the prefix of the input content. The historical prefix associated with the reference state is equivalent to the aforementioned reference content.
[0113] In addition, assuming that multiple reference states are retrieved, one of the reference states can be randomly selected as the final retrieval result, or one of the reference states can be selected as the final retrieval result by other means. The embodiments of the present disclosure do not limit this here; in addition, the reference state can also be retrieved from multiple intermediate states according to other retrieval processes. The embodiments of the present disclosure do not limit this here.
[0114] Based on this, since the original weight matrix is updated based on the same trainable matrix before the mapping layer maps the reference content and the input content respectively, and the prefix of the reference content matches the prefix of the input content, this indicates that the reference content and the input content correspond to the same task type. When the mapping layer maps the reference content and the input content respectively, the same supplementary knowledge needs to be referred to. The prefixes of the reference content and the input content are the same or similar. Therefore, the mapping result of the prefix of the input content is the same or similar to the reference state. Therefore, the retrieved reference state can be used as the mapping result of the prefix of the input content.
[0115] It should be noted that the mapping layer may include an attention sub-layer and a feed-forward sub-layer. The reference state may include the attention key-value pair determined by the attention sub-layer when the reference content is mapped, and the reference state may also include the hidden state determined by the feed-forward sub-layer when the reference content is mapped. In addition, the first trainable matrix can be inserted into the attention sub-layer or the feed-forward sub-layer in the mapping layer. The embodiments of the present disclosure do not limit this here.
[0116] Step 204: Invoke the updated mapping layer, reuse the reference state to map the input content to obtain mapping features, and invoke the first output layer corresponding to the task type information in the content processing model to output the processing result based on the mapping features.
[0117] Among them, the output layer refers to the last network layer in the content processing model. The output layer is responsible for converting the mapped features into the processing results output by the content processing model. The content processing model can set corresponding output layers for different task types. Therefore, the first output layer dedicated to the current task type can be determined based on the task type information. For example, both the generation task type and the classification task type have corresponding output layers. The generation task type can be used to indicate text generation, image generation, audio generation, etc. That is, the processing results corresponding to the generation task type can be data in modalities such as text, images, or audio. Taking the generation task type used to indicate text generation as an example, the output layer corresponding to the generation task type is used to output the probability distribution of each candidate token. Then, the predicted token at the current time step is determined through the probability distribution of each candidate token, and the processing result is determined through the predicted token. The output layer corresponding to the classification task type is used to output the probability distribution of each candidate category, and then the predicted category determined through the probability distribution of each candidate category is used as the processing result.
[0118] It should be noted that the content processing model is equivalent to a neural network model including a main network structure, multiple trainable matrices that can be loaded, and multiple output layers that can be called. The main network structure of the content processing model can directly adopt components such as the embedding layer and the Transformer layer in a large language model or a multi-modal large language model. That is, the overall structure formed by components such as the embedding layer and the Transformer layer is used as a mapping layer, enabling the content processing model to have strong semantic understanding ability and feature extraction ability, etc. Moreover, the main network structure can be shared among multiple processing tasks, achieving maximized knowledge transfer and parameter utilization.
[0119] It should be noted that since the reference state can be the state obtained by prefix mapping of the input content, reusing the reference state to map the input content to obtain the mapped features specifically means that when mapping the input content, there is no need to re-determine the state determined by the prefix of the input content during mapping. Only the result obtained by mapping other parts of the input content based on the reference state needs to be determined, so as to quickly map to obtain the mapped features, avoid repeated calculations for the same or similar prefixes, significantly improve the response speed of the model to long sequences or continuous conversations, and is particularly suitable for real-time applications that require fast responses.
[0120] Exemplarily, refer to Figure 4 , Figure 4 which is an optional process schematic diagram for obtaining the mapped features provided by the embodiments of the present disclosure.
[0121] Among them, the input content is input into the updated mapping layer. When the mapping layer maps the input content, it will reuse the reference state to map the input content to obtain the mapped features.
[0122] For example, when the task type represented by the task type information is a generation task type, the reference state may include the attention key-value pairs determined during mapping of the reference content. Since the prefix of the reference content matches the prefix of the input content, the attention key-value pairs determined during mapping of the prefix of the input content are the same as or similar to the attention key-value pairs determined during mapping of the reference content. By reusing the attention key-value pairs determined during mapping of the reference content, there is no need to recalculate the attention key-value pairs of the prefix of the input content. Only the query vector needs to be determined based on the other parts of the input content, and then the attention scores of the other parts of the input content are calculated through the query vector and the attention key-value pairs included in the reference state, which can effectively improve the response efficiency.
[0123] For another example, when the task type represented by the task type information is a classification task type, the reference state may include the hidden state determined during mapping of the reference content. Since the prefix of the reference content matches the prefix of the input content, the hidden state determined during mapping of the prefix of the input content is the same as or similar to the hidden state determined during mapping of the reference content. By reusing the hidden state determined during mapping of the reference content, there is no need to recalculate the hidden state of the prefix of the input content. Assuming that the length of the input content is short, then the entire input content can be used as the prefix of the input content. At this time, the hidden state determined during mapping of the reference content can be directly used as the mapping feature of the input content, realizing directly obtaining the mapping feature, which can effectively improve the response efficiency.
[0124] Based on this, by obtaining the input content and task type information, and then loading the first trainable matrix that matches the mapping layer of the content processing model according to the task type information. Since the task type information can represent the current task type, it is equivalent to loading the first trainable matrix dedicated to the current task type. Then, based on the first trainable matrix, the original weight matrix of the mapping layer is updated, so that the mapping layer can better adapt to the current task type. Moreover, for different task type information, only the corresponding first trainable matrix needs to be loaded based on the task type information, realizing the processing task of using one content processing model to process multiple task types, without the need to deploy multiple dedicated models, which can effectively reduce the development cost. Then, according to the input content and the first trainable matrix, the pre-cached reference state is retrieved. Since the mapping layer updates the original weight matrix based on the same first trainable matrix before mapping the reference content and the input content respectively, and the prefix of the reference content matches the input content, the mapping result of the prefix of the input content is the same as or similar to the reference state. Therefore, the retrieved reference state can be used as the mapping result of the prefix of the input content. Then, the updated mapping layer is called, and the reference state is reused to map the input content to obtain the mapping features. The first output layer corresponding to the task type information in the content processing model is called to output the processing result based on the mapping features, that is, the processing result is output based on the first output layer dedicated to the current task type, which can ensure the accuracy of the processing result. Moreover, the reference state is reused in the mapping process of the input content, without the need to remap the prefix of the input content, which can reduce repeated calculations, thereby effectively improving the response efficiency of the content processing model.
[0125] The processing process of the content processing model is described in detail below.
[0126] Among them, the mapping layer of the content processing model includes multiple stacked Transformer layers. Each Transformer layer includes an attention sub-layer for performing attention processing and a feed-forward sub-layer for performing feed-forward processing. The attention sub-layer can use the multi-head self-attention mechanism for attention processing, and both the attention sub-layer and the feed-forward sub-layer apply residual connections and layer normalization.
[0127] Suppose the mapping layer of the content processing model includes L stacked Transformer layers. The calculation formula of each Transformer layer is as follows:
[0128] h l ′ = LN(Attention(h l-l ) + h l-1 )
[0129] h l = LN(FFN(h l ′) + h l ′)
[0130] where l ≤ L, both l and L are positive integers, and h l-1 is the output of the (l - 1)-th Transformer layer, Attention is used to represent the attention sub-layer, and Attention(h l-1 ) means performing attention processing on h l-1 . Attention(h l-1 ) is specifically the output of the l-th attention sub-layer, and Attention(h l-1 ) + h l-1 means performing a residual connection on the output of the l-th attention sub-layer. The full name of LN is Layer Normalization, and LN() is used for layer normalization. h l ' is the output of the l-th attention sub-layer after residual connection and layer normalization;
[0131] The full name of FFN is Feed Forward Network, that is, FFN is used to represent the feed-forward sub-layer. FFN(h l ') means performing feed-forward processing on h l '. FFN(h l ') is specifically the output of the l-th feed-forward sub-layer, and FFN(h l ') + h l ' means performing a residual connection on the output of the l-th feed-forward sub-layer. h l is the output of the l-th feed-forward sub-layer after residual connection and layer normalization, that is, h l is the output of the l-th Transformer layer.
[0132] In addition, the output h L of the last Transformer layer can be used as the mapping feature. At this time, assuming that the task type information is used to characterize that the current task type is the generation task type, taking the generation task type used to indicate text generation as an example, the calculation formula of the first output layer corresponding to the generation task type is as follows:
[0133] y gen = softmax(W gen *h L + b gen )
[0134] where y gen is the probability distribution of each candidate token output by the first output layer corresponding to the generation task type, W gen is the weight matrix of the first output layer corresponding to the generation task type, b gen is the bias vector of the first output layer corresponding to the generation task type, and h LFor the mapping feature, softmax() is the normalization function. The probability distribution of each candidate token includes the probability values of each candidate token. The predicted token at the current time step is determined through the probability distribution of each candidate token. For example, the candidate token with the largest probability value is used as the predicted token.
[0135] Exemplarily, assume the input content is "What color is the sky". Multiple stacked Transformer layers are called to reuse the reference state, and each token in the input content is mapped to a corresponding feature representation. This is equivalent to mapping the input content through the updated mapping layer and reusing the reference state to obtain the mapping feature. The mapping feature includes the feature representations corresponding to each token in the input content. The feature representation corresponding to the last token in the input content is input to the first output layer corresponding to the generation task type. The predicted token at the current time step is determined through the probability distribution of each candidate token output by the first output layer. Assume that the predicted token at the current time step is not the end token. Then, the feature representation corresponding to the predicted token at the current time step is used as the input for the next time step for autoregressive generation to determine the predicted token for the next time step until the currently generated predicted token is the end token. The processing result is determined based on all the determined predicted tokens. For example, if the first determined predicted token is "blue", the second determined predicted token is "color", and the third determined predicted token is the end token, then the processing result is "blue color".
[0136] In addition, assume that the task type information is used to characterize that the current task type is a classification task type. The calculation formula for the first output layer corresponding to the classification task type is as follows:
[0137] y cls = softmax(W cls *h L + b cls )
[0138] where y cls is the probability distribution of each candidate category output by the first output layer corresponding to the classification task type, W cls is the weight matrix of the first output layer corresponding to the classification task type, b cls is the bias vector of the first output layer corresponding to the classification task type, h L is the mapping feature, softmax() is the normalization function. The probability distribution of each candidate category includes the probability values of each candidate category. The predicted category is determined through the probability distribution of each candidate token. For example, the candidate category with the largest probability value is used as the predicted category.
[0139] Exemplarily, the first output layer corresponding to the classification task type can determine the probability distributions of three candidate categories, namely positive, negative, and neutral. Suppose the input content is "I'm very happy today". By calling multiple stacked Transformer layers to reuse the reference state, each word segment in the input content is mapped to a corresponding feature representation. This is equivalent to mapping the input content through the updated mapping layer by reusing the reference state to obtain the mapping features. The mapping features include the feature representations corresponding to each word segment in the input content. The feature representation corresponding to the last word segment in the input content is input to the first output layer corresponding to the classification task type, and the prediction category is determined through the probability distributions of the three candidate categories output by the first output layer. The prediction category is used as the processing result. For example, among the probability distributions of the three candidate categories, if the candidate category with the largest probability value is "positive", then the prediction category is "positive".
[0140] It should be noted that there are various training methods for the output layer of the content processing model. One of the training methods for the output layer will be described in detail below.
[0141] In a possible implementation, when the first level of the task type includes the generation task type, the content processing model can set the output layer corresponding to the generation task type. The output layer corresponding to the generation task type can be called the generation head. The generation head is used to output the predicted token at the current time step and use the predicted token as the input of the content processing model at the next time step for autoregressive generation. That is, each output predicted token is used as the processing result, and then the second low-rank decomposition matrices corresponding to each generation sub-task type are deployed for the generation head. For example, both the translation generation sub-task type and the abstract generation sub-task type have corresponding second low-rank decomposition matrices.
[0142] Similarly, when the first level of the task type includes the classification task type, the content processing model can set the output layer corresponding to the classification task type. The output layer corresponding to the classification task type can be called the classification head. The classification head is used to output the probability distributions of each category. That is, the probability distribution is used as the processing result, and then the second low-rank decomposition matrices corresponding to each classification sub-task type are deployed for the classification head. For example, both the sentiment classification sub-task type and the topic classification sub-task type have corresponding second low-rank decomposition matrices.
[0143] In the fine-tuning training phase, it is necessary to adjust each of the second low-rank decomposition matrices separately. Before adjusting any one of the second low-rank decomposition matrices, it is necessary to freeze the weight matrix of the output layer, insert the second low-rank decomposition matrix into the corresponding output layer, and also obtain the training data corresponding to the second low-rank decomposition matrix. This training data is data related to the task type corresponding to the inserted second low-rank decomposition matrix. For example, when the task type is the abstract generation sub-task type, insert the second low-rank decomposition matrix into the output layer corresponding to the generation task type, and the training samples are samples related to the abstract generation sub-task type. Exemplarily, the training samples include the article text and the abstract text of the article text. Then, based on the training samples, train the loaded second low-rank decomposition matrix, which realizes the training of the output layer, can effectively reduce the number of training parameters, and thus effectively reduce the time and space for model training.
[0144] Then, in the inference phase, the process of determining the processing result can refer to Figure 5 , Figure 5 which is an optional flowchart for determining the processing result provided by the embodiments of the present disclosure.
[0145] Specifically, call the first output layer corresponding to the task type information in the content processing model to output the processing result based on the mapping features. Specifically, it can be to determine the first output layer in the generation head and the classification head according to the task type information, and load the target low-rank decomposition matrix corresponding to the task type information from multiple second low-rank decomposition matrices. For example, when the task type represented by the task type information is the abstract generation sub-task type, use the generation head as the first output layer, and use the second low-rank decomposition matrix corresponding to the abstract generation sub-task type as the target low-rank decomposition matrix. Then insert the target low-rank decomposition matrix into the first output layer, and then call the first output layer to output the processing result based on the mapping features, which can effectively reduce the development cost, system complexity, and maintenance cost. Moreover, when adding or reducing task types, only the management component of the second low-rank decomposition matrix needs to be independently updated, which can improve the maintainability and scalability of the system.
[0146] Next, another training method for the output layer will be described in detail.
[0147] In a possible implementation manner, the content processing model can set the output layers corresponding to each sub-task type. For example, assuming there are K sub-task types in total, then the content processing model can set K output layers, where K is a positive integer. Exemplarily, assuming there are 4 sub-task types including the translation generation sub-task type, the abstract generation sub-task type, the sentiment classification sub-task type, and the topic classification sub-task type, then the content processing model can set the output layers corresponding to the 4 sub-task types.
[0148] In the training phase, each output layer array needs to be trained separately. Before adjusting any output layer, the training data corresponding to the output layer needs to be obtained. This training data is data related to the subtask type corresponding to the output layer. For example, when the task type is the abstract generation subtask type, the training samples are samples related to the abstract generation subtask type. Exemplarily, the training samples include the article text and the abstract text of the article text. Then, the output layer is trained based on the training samples.
[0149] Then, in the inference phase, the process of determining the processing result can refer to Figure 6 , Figure 6 which is another optional flowchart for determining the processing result provided by the embodiments of the present disclosure.
[0150] Specifically, the first output layer corresponding to the task type information in the content processing model is called to output the processing result based on the mapping feature. Specifically, the first output layer can be determined from multiple output layers according to the task type information, and then the first output layer is called to output the processing result based on the mapping feature, which can improve the prediction accuracy.
[0151] In a possible implementation, the content processing method further includes: detecting performance metric information, generating a record log according to the performance metric information; and triggering an alarm when the performance metric information is greater than or equal to a third preset threshold. Among them, the performance metric information can include inference latency, throughput, resource utilization rate, or cache hit rate, etc. The resource utilization rate can include CPU utilization rate, GPU utilization rate, and memory utilization rate, etc.
[0152] In a possible implementation, the content processing model is further provided with at least one second output layer similar to the function of the first output layer. The original weight matrix of the mapping layer is updated based on the first trainable matrix. Specifically, the second trainable matrix corresponding to the second output layer can be loaded; based on the first trainable matrix and the second trainable matrix, an operation is performed with the original weight matrix of the mapping layer to obtain the updated original weight matrix.
[0153] Specifically, the content processing model can be provided with multiple output layers. Similar to the first output layer, the second output layer is also one of the multiple output layers. The second output layer being similar to the first output layer in function can mean that the information concerned by the second output layer is similar to the information concerned by the first output layer. For example, assuming that the first output layer is used to output the sentiment classification result, the output layer used to output the emotion generation result can be used as the second output layer. This is because both the first output layer and the second output layer are used to output emotion-related information, which means that both the first output layer and the second output layer need to understand and focus on emotion-related information. Therefore, the second output layer is similar to the first output layer in function.
[0154] It can be understood that corresponding trainable matrices can be deployed for each task type respectively. Similar to the first trainable matrix, the second trainable matrix is also one of the multiple trainable matrices. Since the second output layer is one of the multiple output layers, the second output layer also has a corresponding task type. Therefore, the corresponding second trainable matrix can be determined from the multiple trainable matrices according to the task type corresponding to the second output layer, and then the second trainable matrix corresponding to the second output layer can be loaded.
[0155] Based on this, the supplementary knowledge contained in the first trainable matrix can be used as the main supplementary knowledge, and the supplementary knowledge contained in the second trainable matrix can be used as the secondary supplementary knowledge. By loading the second trainable matrix corresponding to the second output layer, and then performing operations on the original weight matrix of the mapping layer based on the first trainable matrix and the second trainable matrix, an updated original weight matrix is obtained, realizing the combination of the main supplementary knowledge, the secondary supplementary knowledge, and the general knowledge. A mapping layer with general knowledge, main supplementary knowledge, and secondary supplementary knowledge can be obtained. Since the second output layer is similar in function to the first output layer, the task type corresponding to the second output layer is related to the task type corresponding to the first output layer. Therefore, the secondary supplementary knowledge has a promoting effect on the reasoning effect of the current processing task. On the basis of reasoning using general knowledge and main supplementary knowledge, by adding the secondary supplementary knowledge with a promoting effect, the task adaptation ability of the mapping layer can be further improved, thereby further improving the task processing effect.
[0156] Exemplarily, refer to Figure 7 , Figure 7 which is an optional flowchart showing the insertion of the first trainable matrix and the second trainable matrix provided by the embodiments of the present disclosure.
[0157] Among them, according to the task type information, the first trainable matrix matching the mapping layer is loaded from the multiple trainable matrices, then the second trainable matrix corresponding to the second output layer is loaded from the multiple trainable matrices, and then operations are performed on the original weight matrix of the mapping layer based on the first trainable matrix and the second trainable matrix to obtain an updated original weight matrix, realizing the insertion of the first trainable matrix and the second trainable matrix into the mapping layer.
[0158] It should be noted that there are various implementation manners for performing operations on the original weight matrix of the mapping layer based on the first trainable matrix and the second trainable matrix. For example, the first trainable matrix, the second trainable matrix, and the original weight matrix are summed, or the sum of the weighted sums of the first trainable matrix and the second trainable matrix is summed with the original weight matrix.
[0159] For another example, when the trainable matrix is the first low-rank factorization matrix, the trainable matrix can be expressed as the product of two low-rank matrices. Then, one of the low-rank matrices in each trainable matrix can be concatenated in sequence to obtain the first concatenated matrix, and the other low-rank matrix in each trainable matrix can be concatenated in sequence to obtain the second concatenated matrix. Finally, the sum of the product of the first concatenated matrix and the second concatenated matrix and the original weight matrix is calculated. Exemplarily, assume that the first trainable matrix is expressed as the product of two low-rank matrices A1 and B1, the second trainable matrix is expressed as the product of two low-rank matrices A 2(1) and B 2(1) , and the original weight matrix is W0. Then the calculated result is W = W0 + A concat ×B concat , where A concat is the concatenation result of A1 and A 2(1) , and B concat is the concatenation result of B1 and B 2(1) .
[0160] In addition to the above implementation methods, there can be other implementation methods. Another implementation method is described in detail below.
[0161] In a possible implementation method, the number of the second output layers is multiple. Based on the first trainable matrix and the second trainable matrix, operations are performed on the original weight matrix of the mapping layer to obtain the updated original weight matrix. Specifically, the first trainable matrix and the original weight matrix of the mapping layer can be operated on to obtain the first operation result; the multiple second output layers are sorted according to the functional similarity degree between the second output layer and the first output layer; based on the sorting result, the second trainable matrix corresponding to the second output layer is iteratively operated on the first operation result in sequence to obtain the updated original weight matrix.
[0162] Exemplarily, the content processing model can be set with 10 output layers, one of which is the first output layer, and at least two of the remaining output layers are the second output layers.
[0163] It should be noted that there are various implementation methods for operating on the first trainable matrix and the original weight matrix of the mapping layer. For example, adding the first trainable matrix and the original weight matrix, or performing weighted summation on the first trainable matrix and the original weight matrix, can both obtain the first operation result.
[0164] Among them, the degree of functional similarity can refer to the similarity between the information concerned by the second output layer and the information concerned by the first output layer. When the degree of functional similarity is higher, it means that the similarity between the information concerned by the second output layer and the information concerned by the first output layer is higher, and the secondary supplementary knowledge contained in the second trainable matrix corresponding to the second output layer has a greater promoting effect on the inference effect of the current processing task. When the degree of functional similarity is lower, it means that the similarity between the information concerned by the second output layer and the information concerned by the first output layer is lower, and the secondary supplementary knowledge contained in the second trainable matrix corresponding to the second output layer has a smaller promoting effect on the inference effect of the current processing task.
[0165] Based on this, first perform an operation on the first trainable matrix and the original weight matrix of the mapping layer to obtain a first operation result, which is equivalent to first obtaining a mapping layer with general knowledge and main supplementary knowledge through the operation. On this basis, sort multiple second output layers according to the degree of functional similarity between the second output layer and the first output layer. Specifically, the multiple second output layers can be sorted in descending order of the degree of functional similarity. Then, based on the sorting result, sequentially perform iterative operations on the second trainable matrix corresponding to the second output layer and the first operation result to obtain an updated original weight matrix, which is equivalent to performing a stacked update on the original weight matrix, realizing the operation based on the second trainable matrix corresponding to the second output layer with a larger degree of functional similarity first, that is, preferentially adding secondary supplementary knowledge with a greater promoting effect. By limiting the order of the second trainable matrix in the iterative operation process, the finally obtained mapping layer can have a high task adaptation ability, thereby further improving the task processing effect.
[0166] Exemplarily, refer to Figure 8 , Figure 8 which is an optional flowchart for updating the original weight matrix provided by an embodiment of the present disclosure.
[0167] Among them, according to the task type information, load the first trainable matrix matching the mapping layer from multiple trainable matrices, then load the second trainable matrix corresponding to each of the multiple second output layers from the multiple trainable matrices, then perform an operation on the first trainable matrix and the original weight matrix of the mapping layer to obtain a first operation result, and then based on the sorting result of the second output layer, sequentially perform iterative operations on the second trainable matrix corresponding to the second output layer and the first operation result to obtain an updated original weight matrix.
[0168] Specifically, there are various implementation manners for the iterative operation. One of the implementation manners of the iterative operation will be described in detail below.
[0169] In a possible implementation, based on the sorting result, the second trainable matrix corresponding to the second output layer is iteratively operated on the first operation result in sequence to obtain an updated original weight matrix. Specifically, the sorting result can be traversed in sequence. For the first second output layer in the sorting result, determine the element-wise multiplication result between the second trainable matrix and the first operation result, and sum the element-wise multiplication result and the first operation result to obtain the second operation result of the first round; for any other second output layer in the sorting result, determine the element-wise multiplication result between the second trainable matrix and the second operation result of the previous round, and sum the element-wise multiplication result and the second operation result of the previous round to obtain the second operation result of the current round; determine the second operation result obtained in the last round as the updated original weight matrix.
[0170] Among them, the multiple second output layers are sorted in descending order according to the degree of functional similarity. Therefore, the first second output layer in the sorting result has the greatest degree of functional similarity.
[0171] Based on this, traverse the sorting result in sequence. First, multiply the second trainable matrix corresponding to the first second output layer in the sorting result with the first operation result element-wise, and then sum the element-wise multiplication result and the first operation result to obtain the second operation result of the first round. The dimension of the second operation result is the same as that of the first operation result, realizing the insertion of the second trainable matrix corresponding to the first second output layer in the mapping layer while ensuring that the dimension of the weight matrix of the mapping layer remains unchanged. Similarly, each of the remaining second output layers in the sorting result is processed in this way. Multiply the second trainable matrix with the second operation result of the previous round element-wise, and then sum the element-wise multiplication result and the second operation result of the previous round to obtain the second operation result of the current round. The dimension of the second operation result obtained in each round is the same, realizing the insertion of the second trainable matrix corresponding to the remaining second output layers in the mapping layer while ensuring that the dimension of the weight matrix of the mapping layer remains unchanged. Finally, determine the second operation result obtained in the last round as the updated original weight matrix. Element-wise multiplication is performed in each operation round, which can capture complex patterns between the weight matrices and strengthen the secondary supplementary knowledge contained in the second trainable matrix, enabling the finally obtained mapping layer to have a high task adaptation ability, thereby further improving the task processing effect.
[0172] Exemplarily, assume that there are n second output layers in the sorting result, and the dimensions of the original weight matrix, the first trainable matrix, and the second trainable matrix are all N×N, where N is a positive integer. Since the first operation result can be obtained by summing the original weight matrix and the first trainable matrix, the dimension of the first operation result is N×N. Then, traverse the sorting result in sequence. For the first second output layer in the sorting result, determine the element-wise multiplication result between the second trainable matrix and the first operation result. The dimension of this element-wise multiplication result is N×N. Sum the element-wise multiplication result and the first operation result to obtain the second operation result of the first round. The dimension of this second operation result is N×N. Similarly, the dimensions of the second operation results obtained in other rounds are also N×N, such that the dimension of the updated original weight matrix is N×N.
[0173] Specifically, assume that the first trainable matrix is represented as the product of two low-rank matrices, and the second trainable matrix corresponding to each second output layer is represented as the product of two low-rank matrices. Then, the calculation formula for the first operation result is as follows:
[0174] W1 = W0 + A1B1
[0175] where W1 is the first operation result, W0 is the original weight matrix, A1 is one of the low-rank matrices of the first trainable matrix, and B1 is the other low-rank matrix of the first trainable matrix.
[0176] Then, the calculation formula for the second operation result of the first round is as follows:
[0177]
[0178] where, is the second operation result of the first round, W1 is the first operation result, A 2(1) is one of the low-rank matrices of the second trainable matrix corresponding to the first second output layer in the sorting result, B 2(1) is the other low-rank matrix of the second trainable matrix corresponding to the first second output layer in the sorting result, and ⊙ means calculating element-wise multiplication;
[0179] Then, the calculation formula for the second operation results of other rounds is as follows:
[0180]
[0181] where 1 < i ≤ n, both i and n are positive integers, and n is the number of second output layers, is the second operation result of the i-th round, is the second operation result of the (i - 1)-th round, A 2(i) is one of the low-rank matrices of the second trainable matrix corresponding to the i-th second output layer in the sorting result, B2(i) For another low-rank matrix of the second trainable matrix corresponding to the i-th second output layer in the sorting result, ⊙ refers to calculating element-wise multiplication.
[0182] A detailed description of another implementation manner of the iterative operation is given below.
[0183] In a possible implementation manner, based on the sorting result, the second trainable matrix corresponding to the second output layer is successively subjected to iterative operation with the first operation result to obtain an updated original weight matrix. Specifically, the sorting result can be traversed in sequence. For the first second output layer in the sorting result, the second trainable matrix is concatenated with the first trainable matrix, and the concatenated result is subjected to cross-attention with the first operation result and then summed with the first operation result to obtain the third operation result of the first round; for any other second output layer in the sorting result, all the second trainable matrices corresponding to the first second output layer to the current second output layer are concatenated with the first trainable matrix, and the concatenated result is subjected to cross-attention with the third operation result of the previous round and then summed with the third operation result of the previous round to obtain the third operation result of the current round; the third operation result obtained in the last round is determined as the updated original weight matrix.
[0184] Among them, the multiple second output layers are sorted in descending order according to the similarity of functions. Therefore, the function similarity corresponding to the first second output layer in the sorting result is the largest.
[0185] Based on this, traverse the sorting results in sequence. First, concatenate the second trainable matrix corresponding to the first second output layer in the sorting results with the first trainable matrix. Then, perform cross-attention on the concatenated result and the first operation result and sum them with the first operation result to obtain the third operation result of the first round. The dimension of the third operation result is the same as that of the first operation result, achieving the insertion and addition of the second trainable matrix corresponding to the first second output layer in the mapping layer while ensuring that the dimension of the weight matrix of the mapping layer remains unchanged. Similarly, each of the remaining second output layers in the sorting results is processed in this way. Concatenate all the second trainable matrices corresponding to the first second output layer to the current second output layer with the first trainable matrix. Perform cross-attention on the concatenated result and the third operation result of the previous round, and then sum them with the third operation result of the previous round to obtain the third operation result of the current round. The dimension of the third operation result obtained in each round is the same, achieving the insertion and addition of the second trainable matrices corresponding to the remaining second output layers in the mapping layer while ensuring that the dimension of the weight matrix of the mapping layer remains unchanged. Finally, determine the third operation result obtained in the last round as the updated original weight matrix. In each operation round, all the second trainable matrices corresponding to all the traversed second output layers will be concatenated with the first trainable matrix, which can integrate all the traversed supplementary knowledge. Additionally, through cross-attention processing, the supplementary knowledge is intelligently aggregated, further strengthening the main supplementary knowledge contained in the first trainable matrix and the secondary supplementary knowledge contained in the second trainable matrix, enabling the finally obtained mapping layer to have a high task adaptation ability, thereby further improving the task processing effect.
[0186] In a possible implementation manner, perform cross-attention on the concatenated result and the first operation result and sum them with the first operation result to obtain the third operation result of the first round. Specifically, the first operation result can be used as the query matrix, and the concatenated result can be used as the key matrix and value matrix; based on the query matrix, key matrix, and value matrix, perform cross-attention and then sum them with the first operation result to obtain the third operation result of the first round.
[0187] Based on this, by using the first operation result as the query matrix for the first round, and the concatenation result as the key matrix and value matrix for the first round, then, performing cross-attention based on the query matrix, key matrix, and value matrix of the first round to obtain the cross-attention result of the first round, and then summing the cross-attention result of the first round with the first operation result to obtain the third operation result of the first round, it is possible to selectively aggregate the supplementary knowledge in the value matrix according to the query matrix, further strengthening the main supplementary knowledge contained in the first trainable matrix and the secondary supplementary knowledge contained in the second trainable matrix, so that the finally obtained mapping layer can have a high task adaptation ability, thereby further improving the task processing effect.
[0188] It should be noted that, similar to the determination method of the third operation result of the first round, after performing cross-attention on the concatenation result and the third operation result of the previous round, and then summing it with the third operation result of the previous round to obtain the third operation result of the current round. Specifically, it can be to use the third operation result of the previous round as the query matrix of the current round, and the concatenation result as the key matrix and value matrix of the current round; perform cross-attention based on the query matrix, key matrix, and value matrix of the current round to obtain the cross-attention result of the current round, and then sum the cross-attention result of the current round with the third operation result of the previous round to obtain the third operation result of the current round.
[0189] Exemplarily, assume that there are n second output layers in the sorting result, and the dimensions of the original weight matrix, the first trainable matrix, and the second trainable matrix are all N×d k , N and d k are both positive integers. Since the first operation result can be obtained by summing the original weight matrix and the first trainable matrix, the dimension of the first operation result is N×d k , and then by concatenating the first trainable matrix and the second trainable matrix, a concatenation result with a dimension of 2N×d k can be obtained. Then, use the first operation result as the query matrix of the first round, and the concatenation result as the key matrix and value matrix of the first round, that is, the dimension of the query matrix of the first round is N×d k , and the dimensions of the key matrix and value matrix of the first round are both 2N×d k . Then, through cross-attention, the cross-attention result of the first round is obtained, and the dimension of the cross-attention result of the first round is N×d k . Then, the cross-attention result of the first round is summed with the first operation result to obtain the third operation result of the first round, and the dimension of the third operation result of the first round is N×d kSimilarly, the dimension of the third operation result of other rounds is also N×d k , so that the dimension of the updated original weight matrix is N×d k .
[0190] Specifically, the calculation formula for the third operation result of the first round is as follows:
[0191]
[0192] in, is the third operation result of the first round, W1 is the first operation result, Q (1) is the query matrix of the first round, K (1) is the key matrix of the first round, K (1) The transposed result, V (1) is the value matrix of the first round, Q (1) The dimension is N×d k , K (1) and V (1) The dimensions are all 2N×d k , N and d k are all positive integers, d k is the scaling factor, softmax() is the normalization function, This is the result of the first round of cross attention.
[0193] Then, the calculation formula for the third operation result of other rounds is as follows:
[0194]
[0195] Among them, 1 <i≤n,i和n均为正整数,n为第二输出层的数量, is the third operation result of the i-th round, is the result of the third operation in the i-1th round, Q (i) is the query matrix of the i-th round, K (i) is the key matrix of the ith round, K (i) The transposed result, V (i) is the value matrix of the i-th round, Q (i) The dimension is N×d k , K (i) and V (i) The dimensions are (i+1)N×d k , N and d k are all positive integers, d k is the scaling factor, softmax() is the normalization function, is the cross attention result of the i-th round.
[0196] In a possible implementation, a pre-cached reference state is retrieved based on input content and a first trainable matrix, specifically, a candidate state that matches the first trainable matrix and is pre-cached is retrieved based on the first trainable matrix; a first data feature of non-text prompt data is extracted, and a second data feature of non-text candidate data is extracted; a semantic match between the prompt text and the candidate text is determined, and a target similarity between the first data feature and the second data feature is determined; when the semantic match is greater than or equal to a first preset threshold, and the target similarity is greater than or equal to a second preset threshold, the candidate state is determined as the retrieved reference state.
[0197] Among them, the prefix of the input content includes prompt text and non-text prompt data, the candidate state is generated by calling the mapping layer to map the candidate content after updating the original weight matrix based on the first trainable matrix, and the candidate content includes candidate text and non-text candidate data.
[0198] It can be understood that since the candidate state is generated by calling the mapping layer to map the candidate content after updating the original weight matrix based on the first trainable matrix, the candidate state matches the first trainable matrix. Since the input content matches the first trainable matrix, the candidate state and the input content correspond to the same task type.
[0199] Specifically, the candidate content is the content that is processed before the content processing model processes the input content. The task type information corresponding to the candidate content is the same as the task type information corresponding to the input content. Similar to the process of processing the input content, when processing the candidate content, it is also necessary to load the first trainable matrix according to the task type information, update the original weight matrix of the mapping layer based on the first trainable matrix, and then call the mapping layer to map the candidate content, so as to generate a candidate state of the candidate content, and then cache the candidate state. The cached candidate state will be associated with the first trainable matrix loaded when processing the candidate content. Therefore, assuming that multiple aforementioned intermediate states are pre-cached, then the candidate state that matches the first trainable matrix and is pre-cached can be retrieved from multiple intermediate states according to the first trainable matrix.
[0200] Based on this, when the prefix of the input content includes the prompt text and non-text prompt data, it represents that the input content includes multiple modal data. Therefore, during the retrieval process, it is necessary to analyze based on different modal data respectively. Specifically, first retrieve the candidate states that match the first trainable matrix and are pre-cached. At this time, only the match with the first trainable matrix needs to be considered, without considering different modal data. Then, for the text modal data, it is necessary to determine the semantic matching degree between the prompt text in the prefix of the input content and the candidate text in the extracted candidate content. For the non-text modal data, it is necessary to extract the first data feature of the non-text prompt data in the prefix of the input content, and extract the second data feature of the non-text candidate data in the candidate content, and then determine the target similarity between the first data feature and the second data feature. When the semantic matching degree is greater than or equal to the first preset threshold, and the target similarity is greater than or equal to the second preset threshold, it represents that the reference content and the prefix of the input content have the same or similar text modal data, and have the same or similar non-text modal data, that is, the reference content matches the prefix of the input content. At this time, the candidate state is determined as the retrieved reference state, effectively improving the reliability of using the retrieved reference state as the mapping result of the prefix of the input content, thereby improving the accuracy of the processing result.
[0201] In a possible implementation manner, the number of non-text prompt data is multiple, and the number of the first data features and the number of the second data features are both the same as the number of non-text prompt data. To determine the target similarity between the first data feature and the second data feature, specifically, the feature weights of the first data features corresponding to each non-text prompt data can be determined according to the input order of each non-text prompt data in the input content; for each first data feature, determine the initial similarity between the first data feature and the corresponding second data feature; based on the feature weights, perform weighted summation on each initial similarity to obtain the target similarity.
[0202] It should be noted that when the number of non-text prompt data is multiple, it represents that the prefix of the input content includes multiple non-text prompt data. For example, the prefix of the input content includes multiple prompt images, and the first data features of each non-text prompt data can be extracted, that is, the number of the first data features is the same as the number of non-text prompt data. At the same time, only the candidate content with the same number of non-text candidate data needs to be considered, and the second data features of each non-text candidate data can also be extracted, that is, the number of the second data features is the same as the number of non-text prompt data.
[0203] Based on this, when the number of non-text prompt data is multiple, since the content with a more forward input order is usually more important, it is necessary to determine the feature weights of the first data features corresponding to each non-text prompt data according to the input order of each non-text prompt data in the input content, so that the feature weights of the first data features corresponding to the non-text prompt data with a more forward input order are larger, that is, the more forward the input order of the non-text prompt data, the higher its importance, and the feature weights of the first data features corresponding to the non-text prompt data with a more backward input order are smaller, that is, the more backward the input order of the non-text prompt data, the lower its importance. Then, determine the initial similarity between each first data feature and the corresponding second data feature, and finally perform weighted summation on each initial similarity based on the feature weights to obtain the target similarity. Through weighted summation, more attention can be paid to the non-text prompt data with higher importance in the prefix of the input content, which can improve the reliability of the target similarity and further improve the reliability of mapping the retrieved reference state as the prefix of the input content, thereby improving the accuracy of the processing result.
[0204] It should be noted that quantization processing can be used to reduce the model size and computational complexity. For example, the weight values of 32-bit floating-point numbers in the original weight matrix can be quantized to 8-bit integers, or the activation values of 32-bit floating-point numbers output by the normalization function can be quantized to 16-bit floating-point numbers or 16-bit brain floating-point numbers.
[0205] Specifically, the quantization formula for the weight values of 8-bit integers is as follows:
[0206] W_int8 = round(W_fp32 * scale)
[0207] where W_int8 is the weight value of 8-bit integers, W_fp32 is the weight value of 32-bit floating-point numbers, scale is the preset quantization scale factor, and round() is the rounding function;
[0208] Then, the quantization formula for the activation values of 16-bit floating-point numbers is as follows:
[0209] A_fp16 = cast_to_fp16(A_fp32)
[0210] where A_fp16 is the activation value of 16-bit floating-point numbers, A_fp32 is the activation value of 32-bit floating-point numbers, and cast_to_fp16() is the conversion function to convert to 16-bit floating-point numbers.
[0211] In addition to the above implementation methods, quantization processing can also be implemented in other ways. Another implementation method will be described in detail below.
[0212] In a possible implementation, the updated mapping layer is called to map the input content using the reference state to obtain mapping features. Specifically, the quantization parameters cached in advance can be retrieved according to the task type information; the original weight matrix updated based on the first trainable matrix is quantized according to the quantization parameters, the quantized mapping layer is called, and the input content is mapped using the reference state to obtain mapping features.
[0213] It can be understood that the corresponding quantization parameters can be cached in advance for each task type. Therefore, the quantization parameters cached in advance can be retrieved according to the task type information, that is, the quantization parameters dedicated to the current task type are determined. Then, the original weight matrix updated based on the first trainable matrix is quantized according to the quantization parameters. Then, the quantized mapping layer is called, and the input content is mapped using the reference state to obtain mapping features. The quantized original weight matrix can not only adapt to the current task type, but also effectively reduce the storage space required for each weight value in the original weight matrix with an acceptable accuracy loss, thereby saving storage space. It can also improve the computing efficiency, reduce the energy consumption, and achieve a significant reduction in the video memory occupancy while maintaining the model performance.
[0214] Exemplarily, refer to Fig. 9 , Fig. 9 which is another optional flowchart for obtaining the mapping features provided by the embodiments of the present disclosure.
[0215] Among them, the quantization parameters cached in advance are retrieved according to the task type information, then the updated original weight matrix is quantized according to the quantization parameters, and then the input content is input into the updated mapping layer. When the mapping layer maps the input content, it will reuse the reference state, so as to map the input content to obtain mapping features.
[0216] In a possible implementation, the quantization parameters include a scaling factor and a zero point. Quantizing the original weight matrix updated based on the first trainable matrix according to the quantization parameters can be specifically to perform a periodic mapping on the original weight matrix updated based on the first trainable matrix to obtain a periodic mapping result; compressing the original weight matrix to obtain a compression result, and determining the product result between the compression result and the periodic mapping result; performing a non-linear adjustment on the scaling factor, and adjusting the non-linearly adjusted scaling factor according to the zero point and a preset first offset parameter to obtain a first bias result, and determining the ratio between the product result and the first bias result; adjusting the ratio according to the zero point and a preset second offset parameter to obtain a second bias result, and rounding the second bias result to obtain the quantization result of the original weight matrix updated based on the first trainable matrix.
[0217] It should be noted that the periodic mapping of the original weight matrix can be to input the weight values of the original weight matrix into a periodic function. For example, the periodic function uses the sine function sin(). The value range of the sine function sin() is [-1, 1]. Assuming the weight value r is very large, after being processed by the sine function, the value can be limited within [-1, 1], ensuring that the quantization result will not overflow due to the excessive weight value in the subsequent quantization calculation.
[0218] It can be understood that the quantization process can be explained in three processing steps. In the first processing step, by performing periodic mapping on the original weight matrix, a periodic mapping result is obtained, which realizes mapping the weight values of the original weight matrix to a finite and periodic range, enabling the weight values that might originally be widely distributed to fluctuate within a relatively concentrated interval. This helps to better handle the distribution of weight values during the quantization process and avoid the excessive influence of some overly large or small weight values on the quantization result. Then, the original weight matrix is compressed so that weight values of different magnitudes can be more evenly processed during the quantization process, which helps to improve the quantization accuracy. Then, the compressed result is multiplied by the periodic mapping result to determine the product result between the compressed result and the periodic mapping result. The product result is equivalent to combining the periodic fluctuations and the smoothly changing weight values.
[0219] Then, in the second processing step, the ratio between the product result and the first bias result is determined, which is equivalent to further scaling and bias-adjusting the weight values to make the effect of the weight values better. Specifically, the first bias result is obtained by adjusting the non-linearly adjusted scaling factor according to the zero point and the first offset parameter. Non-linearly adjusting the scaling factor can effectively control the mapping ratio from floating-point values to integer values. The zero point is used to determine the zero position of the quantization value. By adjusting through the zero point and the first offset parameter, the possible deviation during the quantization process can be compensated, enabling the quantization result to better reflect the distribution characteristics of the original weights.
[0220] Then, in the third processing step, the ratio is adjusted according to the zero point and the preset second offset parameter to obtain the second bias result, and the second bias result is rounded to obtain the quantization result of the original weight matrix updated based on the first trainable matrix, making the weight values of the original weight matrix integer values. By adjusting through the zero point and the second offset parameter, the quantized integer values can be more reasonably distributed within the target quantization interval.
[0221] In a possible implementation manner, to obtain the compressed result by compressing the original weight matrix, specifically, for each weight value of the original weight matrix, the summation result between the absolute value of the weight value and the preset offset is determined; the summation result is input into the logarithmic function for operation to obtain the compressed result.
[0222] Among them, the offset is used to represent the starting point of the logarithmic function. For example, the offset can take values of 1, 2, or other numerical values, which are not limited in the embodiments of the present disclosure.
[0223] Based on this, the logarithmic function has the effect of compressing the data range. For example, assuming the offset is 1, when |r| is small, log(1 + |r|) is approximately equal to |r|, and when |r| is large, the growth rate of log(1 + |r|) is much smaller than that of |r|. It can compress the dynamic range of the weight values, so that weight values of different magnitudes can be more evenly processed during the quantization process. For example, when r = 1, log(1 + 1) ≈ 0.3, and when r = 100, log(1 + 100) ≈ 2, which can narrow the gap between the weight values and help improve the quantization accuracy.
[0224] Specifically, the calculation formula for the quantization result is as follows:
[0225]
[0226] Among them, q is the quantization result, r is the weight value of the original weight matrix, sin() is the sine function, sin(r) is used to perform a periodic mapping on the weight value, |r| is the absolute value of the weight value, 1 is the offset, 1 + |r| is the summation result between the absolute value and the preset offset, log() is the logarithmic function, log(1 + |r|) is used to compress the original weight matrix, s is the scaling factor, s 2 refers to a non-linear adjustment of the scaling factor, z0 is the zero point, b1 is the first offset parameter, b2 is the second offset parameter, s 2 +z0×b1 is the first bias result, is the ratio between the product result and the first bias result, is the second bias result, round() is the rounding function, is used to round the second bias result.
[0227] It should be noted that in addition to model compression through quantization processing, model compression can also be achieved through structured pruning or knowledge distillation, etc. And dedicated hardware accelerators can be designed for the model architecture to maximize the energy efficiency ratio, reduce the inference latency, and improve the system throughput. Also, according to the characteristics of large-scale models, the memory access mode and cache policy can be optimized to reduce the memory bandwidth bottleneck and improve the data access efficiency.
[0228] The complete process of the content processing method will be described in detail below.
[0229] Refer to Fig.10 , Fig.10 An optional architecture diagram of the content processing method provided by the embodiments of the present disclosure.
[0230] First, obtain input content and task type information; wherein, the content processing model is further provided with at least one second output layer that is similar in function to the first output layer, and the number of second output layers is multiple.
[0231] Then, load a first trainable matrix that matches the mapping layer of the content processing model according to the task type information; load the second trainable matrix corresponding to the second output layer; perform an operation on the first trainable matrix and the original weight matrix of the mapping layer to obtain a first operation result; sort the multiple second output layers according to the similarity in function between the second output layer and the first output layer; based on the sorting result, sequentially perform iterative operations on the second trainable matrix corresponding to the second output layer and the first operation result to obtain an updated original weight matrix; which is equivalent to inserting the first trainable matrix and the second trainable matrix into the mapping layer.
[0232] Then, retrieve a pre-cached reference state according to the input content and the first trainable matrix, wherein the reference state is generated by calling the mapping layer to map the reference content after updating the original weight matrix based on the first trainable matrix, and the reference content matches the prefix of the input content.
[0233] Then, retrieve pre-cached quantization parameters according to the task type information; quantize the original weight matrix updated based on the first trainable matrix according to the quantization parameters, call the quantized mapping layer, and reuse the reference state to map the input content to obtain mapping features.
[0234] Then, call the first output layer corresponding to the task type information in the content processing model to output a processing result based on the mapping features.
[0235] Based on this, by obtaining the input content and task type information, and then loading the first trainable matrix that matches the mapping layer of the content processing model according to the task type information. Since the task type information can represent the current task type, it is equivalent to loading the first trainable matrix dedicated to the current task type. Then, based on the first trainable matrix, the original weight matrix of the mapping layer is updated, enabling the mapping layer to better adapt to the current task type. Moreover, for different task type information, only the corresponding first trainable matrix needs to be loaded based on the task type information, achieving the processing task of using one content processing model to handle multiple task types, without the need to deploy multiple dedicated models, which can effectively reduce the development cost. Then, according to the input content and the first trainable matrix, the pre-cached reference state is retrieved. Since the mapping layer updates the original weight matrix based on the same first trainable matrix before mapping the reference content and the input content respectively, and the prefix of the reference content matches that of the input content, the mapping result of the prefix of the input content is the same as or similar to the reference state. Therefore, the retrieved reference state can be used as the mapping result of the prefix of the input content. Then, the updated mapping layer is called, and the reference state is reused to map the input content to obtain the mapping features. The first output layer corresponding to the task type information in the content processing model is called to output the processing result based on the mapping features, that is, the processing result is output based on the first output layer dedicated to the current task type, which can ensure the accuracy of the processing result. Moreover, the reference state is reused during the mapping process of the input content, without the need to remap the prefix of the input content, which can reduce duplicate calculations and thus effectively improve the response efficiency of the content processing model.
[0236] It can be seen that the content processing method provided by the embodiments of the present application can be applied to multiple scenarios.
[0237] For example, the content processing method can be applied to an intelligent question answering assistant. The intelligent question answering assistant can be an in-vehicle intelligent question answering assistant, a home intelligent question answering assistant, a medical intelligent question answering assistant, etc. The intelligent question answering assistant can obtain the input content and task type information input by the user, and then quickly output an accurate processing result through the content processing method, which can reduce the development cost of the intelligent question answering assistant and improve the response efficiency of the intelligent question answering assistant.
[0238] Again, for example, the content processing method can be applied to an in-vehicle decision-making system. The in-vehicle decision-making system can obtain the environmental data collected by sensors such as cameras, lidar, and ultrasonic sensors, and use the environmental data as the input content. Then, the in-vehicle decision-making system can determine the task type information according to the input content, and then quickly output an accurate processing result through the content processing method, which can reduce the development cost of the in-vehicle decision-making system and improve the response efficiency of the in-vehicle decision-making system.
[0239] It can be understood that although the steps in each of the above flowcharts are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear description in this embodiment, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above flowchart may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0240] Referring to Fig.11 , Fig.11 FIG. is an optional structural schematic diagram of a content processing device provided by an embodiment of the present disclosure. The content processing device 1100 includes:
[0241] An acquisition module 1101, configured to acquire input content and task type information;
[0242] An update module 1102, configured to load a first trainable matrix matching the mapping layer of the content processing model according to the task type information, and update the original weight matrix of the mapping layer based on the first trainable matrix;
[0243] A retrieval module 1103, configured to retrieve a pre-cached reference state according to the input content and the first trainable matrix, where the reference state is generated by calling the mapping layer to map the reference content after updating the original weight matrix based on the first trainable matrix, and the reference content matches the prefix of the input content;
[0244] An output module 1104, configured to call the updated mapping layer, reuse the reference state to map the input content to obtain mapping features, and call a first output layer corresponding to the task type information in the content processing model to output a processing result based on the mapping features.
[0245] Furthermore, the content processing model is further provided with at least one second output layer having a function similar to that of the first output layer. The update module 1102 is specifically configured to:
[0246] Load a second trainable matrix corresponding to the second output layer;
[0247] Based on the first trainable matrix and the second trainable matrix, perform an operation with the original weight matrix of the mapping layer to obtain an updated original weight matrix.
[0248] Furthermore, the number of the second output layers is multiple. The update module 1102 is specifically configured to:
[0249] Operate the first trainable matrix with the original weight matrix of the mapping layer to obtain a first operation result;
[0250] Sort the multiple second output layers according to the functional similarity degree between the second output layer and the first output layer;
[0251] Based on the sorting result, sequentially perform iterative operations on the second trainable matrix corresponding to the second output layer and the first operation result to obtain an updated original weight matrix.
[0252] Furthermore, the above update module 1102 is specifically configured to:
[0253] Traverse the sorting result sequentially. For the first second output layer in the sorting result, determine the element-wise multiplication result between the second trainable matrix and the first operation result, and sum the element-wise multiplication result and the first operation result to obtain a second operation result for the first round;
[0254] For any other second output layer in the sorting result, determine the element-wise multiplication result between the second trainable matrix and the second operation result of the previous round, and sum the element-wise multiplication result and the second operation result of the previous round to obtain a second operation result for the current round;
[0255] Determine the second operation result obtained in the last round as the updated original weight matrix.
[0256] Furthermore, the above update module 1102 is specifically configured to:
[0257] Traverse the sorting result sequentially. For the first second output layer in the sorting result, splice the second trainable matrix and the first trainable matrix, perform cross-attention on the splicing result and the first operation result, and then sum the result with the first operation result to obtain a third operation result for the first round;
[0258] For any other second output layer in the sorting result, splice all the second trainable matrices corresponding to the first second output layer to the current second output layer and the first trainable matrix, perform cross-attention on the splicing result and the third operation result of the previous round, and then sum the result with the third operation result of the previous round to obtain a third operation result for the current round;
[0259] Determine the third operation result obtained in the last round as the updated original weight matrix.
[0260] Furthermore, the above update module 1102 is specifically configured to:
[0261] Use the first operation result as the query matrix, and use the splicing result as the key matrix and value matrix;
[0262] After performing cross-attention on the query matrix, key matrix, and value matrix and summing with the first operation result, the third operation result of the first round is obtained.
[0263] Furthermore, the prefix of the input content includes prompt text and non-text prompt data. The above-mentioned retrieval module 1103 is specifically used for:
[0264] Retrieve the candidate states that match the first trainable matrix and are pre-cached according to the first trainable matrix. Among them, the candidate states are generated by updating the original weight matrix based on the first trainable matrix and then calling the mapping layer to map the candidate content. The candidate content includes candidate text and non-text candidate data;
[0265] Extract the first data feature of the non-text prompt data and the second data feature of the non-text candidate data;
[0266] Determine the semantic matching degree between the prompt text and the candidate text, and determine the target similarity between the first data feature and the second data feature;
[0267] When the semantic matching degree is greater than or equal to the first preset threshold and the target similarity is greater than or equal to the second preset threshold, determine the candidate state as the retrieved reference state.
[0268] Furthermore, the number of non-text prompt data is multiple, and the number of the first data features and the number of the second data features are both the same as the number of non-text prompt data. The above-mentioned retrieval module 1103 is specifically used for:
[0269] Determine the feature weights of the first data features corresponding to each non-text prompt data according to the input order of each non-text prompt data in the input content;
[0270] For each first data feature, determine the initial similarity between the first data feature and the corresponding second data feature;
[0271] Perform weighted summation on each initial similarity based on the feature weights to obtain the target similarity.
[0272] Furthermore, the above-mentioned output module 1104 is specifically used for:
[0273] Retrieve the pre-cached quantization parameters according to the task type information;
[0274] Quantize the original weight matrix updated based on the first trainable matrix according to the quantization parameters, call the quantized mapping layer, and reuse the reference state to map the input content to obtain the mapping features.
[0275] Furthermore, the quantization parameters include a scaling factor and a zero point. The above-mentioned output module 1104 is specifically used for:
[0276] Perform a periodic mapping on the original weight matrix updated based on the first trainable matrix to obtain a periodic mapping result;
[0277] Compress the original weight matrix to obtain a compression result, and determine the product result between the compression result and the periodic mapping result;
[0278] Perform a non - linear adjustment on the scaling factor, adjust the non - linearly adjusted scaling factor according to the zero point and a preset first offset parameter to obtain a first bias result, and determine the ratio between the product result and the first bias result;
[0279] Adjust the ratio according to the zero point and a preset second offset parameter to obtain a second bias result, and round the second bias result to obtain the quantization result of the original weight matrix updated based on the first trainable matrix.
[0280] Furthermore, the above - mentioned output module 1104 is specifically used for:
[0281] For each weight value of the original weight matrix, determine the sum result between the absolute value of the weight value and a preset offset;
[0282] Input the sum result into a logarithmic function for operation to obtain a compression result.
[0283] The above-mentioned content processing device 1100 and the content processing method are based on the same inventive concept. By obtaining input content and task type information, and then loading a first trainable matrix that matches the mapping layer of the content processing model according to the task type information. Since the task type information can represent the current task type, it is equivalent to loading the first trainable matrix dedicated to the current task type. Then, based on the first trainable matrix, the original weight matrix of the mapping layer is updated, so that the mapping layer can better adapt to the current task type. Moreover, for different task type information, only the corresponding first trainable matrix needs to be loaded based on the task type information, realizing the processing task of using one content processing model to process multiple task types, without deploying multiple dedicated models, which can effectively reduce the development cost. Then, according to the input content and the first trainable matrix, a pre-cached reference state is retrieved. Since the mapping layer updates the original weight matrix based on the same first trainable matrix before mapping the reference content and the input content respectively, and the prefix of the reference content matches the input content, the mapping result of the prefix of the input content is the same as or similar to the reference state. Therefore, the retrieved reference state can be used as the mapping result of the prefix of the input content. Then, the updated mapping layer is called, and the reference state is reused to map the input content to obtain mapping features. The first output layer corresponding to the task type information in the content processing model is called to output a processing result based on the mapping features, that is, the processing result is output based on the first output layer dedicated to the current task type, which can ensure the accuracy of the processing result. Moreover, the reference state is reused during the mapping process of the input content, and there is no need to remap the prefix of the input content, which can reduce repeated calculations, thereby effectively improving the response efficiency of the content processing model.
[0284] The electronic device for executing the above content processing method provided by the embodiments of the present disclosure may be a terminal. Referring to Fig.12 , Fig.12 is a partial structural block diagram of the terminal provided by the embodiments of the present disclosure. The terminal includes: a camera component 1210, a first memory 1220, an input unit 1230, a display unit 1240, a sensor 1250, an audio circuit 1260, a wireless fidelity (WiFi) module 1270, a first processor 1280, and a first power supply 1290, etc. Those skilled in the art can understand that Fig.12 the terminal structure shown in
[0285] The camera assembly 1210 can be used to collect images or videos. Optionally, the camera assembly 1210 includes a front camera and a rear camera. Generally, the front camera is disposed on the front panel of the terminal, and the rear camera is disposed on the back of the terminal. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth camera, a wide-angle camera, and a telephoto camera, so as to implement the function of background blurring by fusing the main camera and the depth camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting function or other fused shooting functions.
[0286] The first memory 1220 can be used to store software programs and modules. The first processor 1280 executes various functional applications and data processing of the terminal by running the software programs and modules stored in the first memory 1220.
[0287] The input unit 1230 can be used to receive input digital or character information, and generate key signal inputs related to the settings and function controls of the terminal. Specifically, the input unit 1230 can include a touch panel 1231 and other input devices 1232.
[0288] The display unit 1240 can be used to display the input information or provided information and various menus of the terminal. The display unit 1240 can include a display panel 1241.
[0289] The audio circuit 1260, the speaker 1261, and the microphone 1262 can provide an audio interface.
[0290] The first power supply 1290 can be alternating current, direct current, a disposable battery, or a rechargeable battery.
[0291] The number of sensors 1250 can be one or more. The one or more sensors 1250 include, but are not limited to: an acceleration sensor, a gyroscope sensor, a pressure sensor, an optical sensor, and so on. Among them:
[0292] The acceleration sensor can detect the magnitudes of accelerations on the three coordinate axes of the coordinate system established with the terminal. For example, the acceleration sensor can be used to detect the components of the gravitational acceleration on the three coordinate axes. The first processor 1280 can control the display unit 1240 to display the user interface in a landscape view or a portrait view according to the gravitational acceleration signal collected by the acceleration sensor. The acceleration sensor can also be used for game or user motion data collection.
[0293] The gyroscope sensor can detect the body direction and rotation angle of the terminal, and the gyroscope sensor can cooperate with the acceleration sensor to collect the 3D actions of the user on the terminal. According to the data collected by the gyroscope sensor, the first processor 1280 can implement the following functions: motion sensing (such as changing the UI according to the user's tilting operation), image stabilization during shooting, game control, and inertial navigation.
[0294] The pressure sensor can be disposed on the side frame of the terminal and / or the lower layer of the display unit 1240. When the pressure sensor is disposed on the side frame of the terminal, it can detect the holding signal of the user on the terminal, and the first processor 1280 can perform left / right hand recognition or quick operation according to the holding signal collected by the pressure sensor. When the pressure sensor is disposed on the lower layer of the display unit 1240, the first processor 1280 can control the operable controls on the UI interface according to the pressure operation of the user on the display unit 1240. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0295] The optical sensor is used to collect the ambient light intensity. In one embodiment, the first processor 1280 can control the display brightness of the display unit 1240 according to the ambient light intensity collected by the optical sensor. Specifically, when the ambient light intensity is high, the display brightness of the display unit 1240 is increased; when the ambient light intensity is low, the display brightness of the display unit 1240 is decreased. In another embodiment, the first processor 1280 can also dynamically adjust the shooting parameters of the camera assembly 1210 according to the ambient light intensity collected by the optical sensor.
[0296] In this embodiment, the first processor 1280 included in the terminal can execute the content processing method of the previous embodiment.
[0297] The electronic device for executing the above content processing method provided by the embodiments of the present disclosure can also be a server. Refer to Fig.13 , Fig.13 which is a partial structural block diagram of the server provided by the embodiments of the present disclosure. The server may vary greatly due to configuration or performance, and may include one or more second processors 1310 and second memories 1330, and one or more storage media 1340 (such as one or more mass storage devices) for storing application programs 1343 or data 1342. Among them, the second memory 1330 and the storage medium 1340 can be transient storage or persistent storage. The program stored in the storage medium 1340 can include one or more modules (not shown in the figure), and each module can include a series of instruction operations on the server. Further, the second processor 1310 can be configured to communicate with the storage medium 1340 and execute a series of instruction operations in the storage medium 1340 on the server.
[0298] The server may also include one or more second power supplies 1320, one or more wired or wireless network interfaces 1350, one or more input / output interfaces 1360, and / or one or more operating systems 1341, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM, and so on.
[0299] The second processor 1310 in the server can be used to execute the content processing method.
[0300] The embodiments of the present disclosure also provide a computer-readable storage medium for storing a computer program for executing the content processing method of each of the foregoing embodiments.
[0301] The embodiments of the present disclosure also provide a computer program product including a computer program stored in a computer-readable storage medium. The processor of the computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, so that the computer device executes the content processing method described above.
[0302] Terms such as "first", "second", "third", "fourth", etc. (if any) in the specification of the present disclosure and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0303] It should be understood that in this disclosure, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the relationship between associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist simultaneously. Here, A and B can be singular or plural. The character " / " generally indicates an "or" relationship between the associated objects before and after. "At least one (item) of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one (item) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0304] It should be understood that in the description of the embodiments of this disclosure, the meaning of "a plurality (or multiple items)" is more than two. Understandings such as "greater than", "less than", and "exceeding" do not include the corresponding number, while understandings such as "above", "below", and "within" include the corresponding number.
[0305] In several embodiments provided in this disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or units can be in electrical, mechanical, or other forms.
[0306] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0307] In addition, the functional units in each embodiment of this disclosure can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0308] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present disclosure, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of various embodiments of the present disclosure. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical discs that can store program codes.
[0309] It should also be understood that the various embodiments provided in the embodiments of the present disclosure can be combined arbitrarily to achieve different technical effects.
[0310] The above is a specific description of the preferred embodiments of the present disclosure, but the present disclosure is not limited to the above embodiments. Those skilled in the art can also make various equivalent deformations or substitutions without departing from the spirit of the present disclosure, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present disclosure.
Claims
1. A content processing method, characterized in that: include: Get input content and task type information; Loading a first trainable matrix matching a mapping layer of a content processing model according to the task type information, and updating an original weight matrix of the mapping layer based on the first trainable matrix; Retrieving a pre-cached reference state according to the input content and the first trainable matrix, wherein the reference state is generated by calling the mapping layer to map the reference content after updating the original weight matrix based on the first trainable matrix, and the reference content matches a prefix of the input content; The updated mapping layer is called, the reference state is reused to map the input content to obtain mapping features, and the first output layer corresponding to the task type information in the content processing model is called to output a processing result based on the mapping features.
2. The content processing method according to claim 1, characterized in that: The content processing model is further provided with at least one second output layer having a function similar to that of the first output layer, and the updating of the original weight matrix of the mapping layer based on the first trainable matrix comprises: Loading a second trainable matrix corresponding to the second output layer; Based on the first trainable matrix and the second trainable matrix, an operation is performed with the original weight matrix of the mapping layer to obtain the updated original weight matrix.
3. The content processing method according to claim 2, characterized in that: The number of the second output layers is multiple, and the first trainable matrix and the second trainable matrix are operated on the original weight matrix of the mapping layer to obtain the updated original weight matrix, including: Performing an operation on the first trainable matrix and the original weight matrix of the mapping layer to obtain a first operation result; sorting the plurality of second output layers according to the functional similarity between the second output layers and the first output layers; Based on the sorting result, the second trainable matrix corresponding to the second output layer is iteratively operated with the first operation result in turn to obtain the updated original weight matrix.
4. The content processing method according to claim 3, characterized in that: The iterative operation of the second trainable matrix corresponding to the second output layer and the first operation result in sequence based on the sorting result to obtain the updated original weight matrix includes: Traversing the sorting results in sequence, for the first second output layer in the sorting results, determining the element-by-element multiplication result between the second trainable matrix and the first operation result, and summing the element-by-element multiplication result and the first operation result to obtain the second operation result of the first round; For any remaining second output layer in the sorting result, determine the element-by-element multiplication result between the second trainable matrix and the second operation result of the previous round, and sum the element-by-element multiplication result with the second operation result of the previous round to obtain the second operation result of the current round; The second operation result obtained in the last round is determined as the updated original weight matrix.
5. The content processing method according to claim 3, characterized in that: The iterative operation of the second trainable matrix corresponding to the second output layer and the first operation result in sequence based on the sorting result to obtain the updated original weight matrix includes: Traversing the sorting results in sequence, for the first second output layer in the sorting results, concatenating the second trainable matrix with the first trainable matrix, performing cross-attention on the concatenated result and the first operation result, and then summing the result with the first operation result to obtain the third operation result of the first round; For any remaining second output layer in the sorting result, concatenate all the second trainable matrices corresponding to the first second output layer to the current second output layer with the first trainable matrix, cross-attend the concatenation result with the third operation result of the previous round, and sum them with the third operation result of the previous round to obtain the third operation result of the current round; The third operation result obtained in the last round is determined as the updated original weight matrix.
6. The content processing method according to claim 5, characterized in that: The step of performing cross attention on the splicing result and the first operation result and summing the result with the first operation result to obtain the third operation result of the first round includes: Using the first operation result as a query matrix and the concatenation result as a key matrix and a value matrix; A third operation result of the first round is obtained by performing cross attention on the query matrix, the key matrix and the value matrix and summing the results with the first operation result.
7. The content processing method according to claim 1, characterized in that: The prefix of the input content includes prompt text and non-text prompt data, and the retrieving a pre-cached reference state according to the input content and the first trainable matrix includes: Retrieving a candidate state that matches the first trainable matrix and is pre-cached according to the first trainable matrix, wherein the candidate state is generated by calling the mapping layer to map candidate content after the original weight matrix is updated based on the first trainable matrix, and the candidate content includes candidate text and non-text candidate data; Extracting a first data feature of the non-text prompt data, and extracting a second data feature of the non-text candidate data; Determine the semantic matching degree between the prompt text and the candidate text, and determine the target similarity between the first data feature and the second data feature; When the semantic matching degree is greater than or equal to a first preset threshold, and the target similarity is greater than or equal to a second preset threshold, the candidate state is determined as the retrieved reference state.
8. The content processing method according to claim 7, characterized in that: The number of the non-text prompt data is multiple, the number of the first data features and the number of the second data features are both the same as the number of the non-text prompt data, and determining the target similarity between the first data features and the second data features includes: Determining a feature weight of the first data feature corresponding to each of the non-text prompt data according to an input order of each of the non-text prompt data in the input content; For each of the first data features, determining an initial similarity between the first data feature and the corresponding second data feature; The initial similarities are weighted and summed based on the feature weights to obtain the target similarity.
9. The content processing method according to claim 1, characterized in that: The calling the updated mapping layer and reusing the reference state to map the input content to obtain a mapping feature includes: Retrieving pre-cached quantization parameters according to the task type information; The original weight matrix updated based on the first trainable matrix is quantized according to the quantization parameter, the quantized mapping layer is called, and the reference state is reused to map the input content to obtain a mapping feature.
10. The content processing method according to claim 9, characterized in that: The quantization parameter includes a scaling factor and a zero point, and quantizing the original weight matrix updated based on the first trainable matrix according to the quantization parameter includes: Performing periodic mapping on the original weight matrix updated based on the first trainable matrix to obtain a periodic mapping result; Compressing the original weight matrix to obtain a compression result, and determining a product result between the compression result and the periodic mapping result; Performing nonlinear adjustment on the scaling factor, adjusting the scaling factor after the nonlinear adjustment according to the zero point and a preset first offset parameter to obtain a first offset result, and determining a ratio between the product result and the first offset result; The ratio is adjusted according to the zero point and a preset second offset parameter to obtain a second offset result, and the second offset result is rounded to obtain a quantization result of the original weight matrix after updating based on the first trainable matrix.
11. The content processing method according to claim 10, characterized in that: The compressing the original weight matrix to obtain a compression result includes: For each weight value of the original weight matrix, determining a summation result between an absolute value of the weight value and a preset offset; The summation result is input into a logarithmic function for operation to obtain a compression result.
12. A content processing device, characterized in that: include: The acquisition module is used to obtain input content and task type information; An updating module, configured to load a first trainable matrix matching a mapping layer of a content processing model according to the task type information, and update an original weight matrix of the mapping layer based on the first trainable matrix; a retrieval module, configured to retrieve a pre-cached reference state according to the input content and the first trainable matrix, wherein the reference state is generated by calling the mapping layer to map the reference content after updating the original weight matrix based on the first trainable matrix, and the reference content matches a prefix of the input content; The output module is used to call the updated mapping layer, reuse the reference state to map the input content to obtain mapping features, and call the first output layer corresponding to the task type information in the content processing model to output the processing result based on the mapping features.
13. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the content processing method described in any one of claims 1 to 11 is implemented.
14. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the content processing method described in any one of claims 1 to 11 is implemented.
15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the content processing method described in any one of claims 1 to 11 is implemented.