Task processing method and device based on large language model, equipment and storage medium
By analyzing the semantic importance of the network layers of a large language model and performing dynamic rank allocation and sparse regularization training, the problem of insufficient accuracy of the traditional LoRA method in financial tasks is solved, and the accuracy and reliability of the model after efficient fine-tuning are improved in financial tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PING AN TECH (SHENZHEN) CO LTD
- Filing Date
- 2026-01-13
- Publication Date
- 2026-05-12
AI Technical Summary
Traditional LoRA methods cannot adapt to the hierarchical semantic features of large language models in financial scenarios, resulting in low accuracy when the model processes financial tasks, especially in insurance claims review scenarios where important information may be missed or the understanding may be inaccurate.
By acquiring financial corpus data, analyzing the semantic importance of network layers, training with a dynamic rank allocation engine and a sparse regularization constraint module, and combining semantic fidelity verification and iterative optimization, a target rank allocation scheme is generated to optimize the parameter configuration of the large language model.
It improves the accuracy and reliability of large language models under limited parameter budgets, thereby enhancing the accuracy and efficiency of financial task processing, such as insurance claims automation and insurance product recommendation.
Smart Images

Figure CN122021791A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology and can be applied to the financial technology field, particularly to task processing methods, devices, computer equipment and storage media based on large language models. Background Technology
[0002] With the rapid development of artificial intelligence technology, large language models, with their powerful language understanding and generation capabilities, are increasingly widely used in the financial sector, covering multiple key business scenarios such as financial risk assessment, market trend prediction, and customer service interaction. Efficient parameter fine-tuning techniques, as a key means to adapt large language models to specific financial tasks, are crucial for improving the model's performance in financial scenarios.
[0003] Currently, the LoRA (Low-Rank Adaptation) method, as a typical representative of efficient parameter fine-tuning techniques, has demonstrated certain advantages in many fields. However, in financial scenarios, the traditional LoRA method has significant drawbacks. Specifically, the traditional LoRA method uses a fixed-rank configuration, applying the same rank value to all layers of a large language model. However, financial semantics have distinct hierarchical characteristics, with different network layers processing semantic information of varying complexity and importance. This fixed-rank configuration cannot adapt to the hierarchical nature of financial semantics, making it difficult for the model to accurately capture key semantic information when processing financial tasks, thus resulting in lower accuracy of large language models in handling financial tasks.
[0004] For example, in the claims review process within the financial and insurance sector, insurance claim materials contain rich and complex textual information such as accident descriptions, medical diagnoses, and expense lists. Different network layers have different processing requirements for this information; some layers need to process basic semantic information, while others need to delve deeper into key details to determine the compliance of the claim. The fixed-rank configuration of traditional LoRA methods cannot flexibly adjust to these different needs, which may cause the model to miss important information or misunderstand the information during the review process, ultimately leading to biased claims review results and affecting the operational efficiency and customer satisfaction of insurance companies.
[0005] Therefore, there is an urgent need for an intelligent and efficient parameter fine-tuning technology to improve the accuracy of large language models in financial task processing. Summary of the Invention
[0006] The purpose of this application is to propose a task processing method, apparatus, computer device, and storage medium based on a large language model, so as to solve the technical problem that existing large language models have low accuracy in financial task processing.
[0007] Firstly, a task processing method based on a large language model is provided, including: Acquire pre-collected financial corpus data; The financial corpus is processed based on a pre-defined large language model, and the semantic importance information of each network layer contained in the large language model is analyzed based on the obtained model output data. Based on a preset dynamic rank allocation engine, the semantic importance information and the preset parameter budget are used to generate a scheme to obtain the corresponding specified rank allocation scheme. Based on the preset training dataset and the specified rank allocation scheme, the large language model is trained using the preset sparse regularization constraint module to obtain the trained first large language model. The first language model is fine-tuned based on a pre-set test dataset to obtain the corresponding second language model, and the semantic fidelity of the second language model is verified. If the second largest language model fails the semantic fidelity verification, the specified rank allocation scheme is adjusted based on the preset adjustment strategy to obtain the corresponding target rank allocation scheme. Based on a preset iterative processing strategy, the target rank allocation scheme is used to train, fine-tune, and evaluate the large language model until the preset iterative termination condition is met and the corresponding target large language model is obtained. The target large language model is used to process the task data to be processed.
[0008] Secondly, a task processing device based on a large language model is provided, including: The first acquisition module is used to acquire pre-collected financial corpus data; The first processing module is used to process the financial corpus based on a preset large language model, and to analyze the semantic importance information of each network layer contained in the large language model based on the obtained model output data. The first generation module is used to perform scheme generation processing on the semantic importance information and the preset parameter budget based on the preset dynamic rank allocation engine to obtain the corresponding specified rank allocation scheme. The training module is used to train the large language model based on a preset training dataset and the specified rank allocation scheme, using a preset sparse regularization constraint module, to obtain the trained first large language model. The verification module is used to fine-tune the first large language model based on a preset test dataset to obtain the corresponding second large language model, and to verify the semantic fidelity of the second large language model. The adjustment module is used to adjust the specified rank allocation scheme based on a preset adjustment strategy to obtain the corresponding target rank allocation scheme if the second largest language model fails the semantic fidelity verification. The second processing module is used to perform model training, fine-tuning and evaluation of the large language model using the target rank allocation scheme based on a preset iterative processing strategy, until the preset iterative termination condition is met and the corresponding target large language model is obtained. The third processing module is used to process the task data to be processed based on the target large language model.
[0009] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described task processing method based on a large language model.
[0010] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described task processing method based on a large language model.
[0011] In the above-mentioned task processing method, apparatus, computer equipment, and storage medium based on a large language model, the following steps are taken: First, pre-collected financial corpus data is acquired; then, the financial corpus is processed based on a preset large language model, and the semantic importance information of each network layer in the large language model is analyzed based on the obtained model output data; subsequently, a scheme generation process is performed on the semantic importance information and preset parameter budget based on a preset dynamic rank allocation engine to obtain a corresponding specified rank allocation scheme; and based on a preset training dataset and the specified rank allocation scheme, the large language model is trained using a preset sparse regularization constraint module to obtain a trained first... A large language model is first established; subsequently, a second large language model is obtained by fine-tuning the first large language model based on a preset test dataset, and the semantic fidelity of the second large language model is verified; if the second large language model fails the semantic fidelity verification, the specified rank allocation scheme is adjusted based on a preset adjustment strategy to obtain the corresponding target rank allocation scheme; further, based on a preset iterative processing strategy, the target rank allocation scheme is used to train, fine-tune, and evaluate the large language model until a preset iteration termination condition is met and the corresponding target large language model is obtained; finally, the task data to be processed is processed based on the target large language model. Unlike existing methods that use fixed-rank configuration, this application employs an automated processing flow centered around four core steps: semantic importance assessment, dynamic rank allocation, sparse regularization constraints, and financial semantic fidelity verification. These steps collaborate to optimize and adapt the large language model, achieving efficient fine-tuning of the large language model within a limited parameter budget. This balances parameter efficiency with semantic fidelity, enabling the subsequent use of the efficiently fine-tuned target large language model to process task data. This effectively improves the accuracy and reliability of the target large language model in relevant task processing. Attached Figure Description
[0012] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0013] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the task processing method based on a large language model according to this application; Figure 3 This is a schematic diagram of a structure of an embodiment of a task processing device based on a large language model according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation
[0014] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.
[0015] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.
[0016] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.
[0017] like Figure 1 As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0018] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.
[0019] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.
[0020] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.
[0021] It should be noted that the task processing method based on a large language model provided in this application is generally executed by a server / terminal device, and correspondingly, the task processing device based on a large language model is generally located in the server / terminal device.
[0022] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.
[0023] Continue to refer to Figure 2 This document illustrates a flowchart of an embodiment of the task processing method based on a large language model according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different requirements. The task processing method based on a large language model provided in this application can be applied to any scenario requiring task processing, and therefore can be applied to products in these scenarios, such as task processing products in the financial insurance field. The task processing method based on a large language model includes the following steps: Step S201: Obtain pre-collected financial corpus data.
[0024] In this embodiment, the task processing method based on a large language model runs on an electronic device (e.g., Figure 1The server / terminal device shown can acquire financial corpus data via wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra-wideband) connections, and other currently known or future known wireless connection methods. The executing entity of this application is specifically a task processing system, which may be simply referred to as the system.
[0025] This involves pre-collecting a large-scale financial corpus, covering various sources such as news reports, company financial reports, industry research reports, and financial forum discussions, to ensure the diversity and comprehensiveness of the corpus and to cover various semantic information in the financial field.
[0026] Step S202: Process the financial corpus based on a preset large language model, and analyze the semantic importance information of each network layer contained in the large language model based on the obtained model output data.
[0027] In this embodiment, the specific implementation process of processing the financial corpus based on the preset large language model and analyzing the semantic importance information of each network layer contained in the large language model based on the obtained model output data will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0028] Step S203: Based on the preset dynamic rank allocation engine, the semantic importance information and the preset parameter budget are processed to generate a scheme to obtain the corresponding specified rank allocation scheme.
[0029] In this embodiment, the specific implementation process of generating a scheme by performing scheme generation on the semantic importance information and the preset parameter budget based on the preset dynamic rank allocation engine to obtain the corresponding specified rank allocation scheme will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0030] Step S204: Based on the preset training dataset and the specified rank allocation scheme, the large language model is trained using the preset sparse regularization constraint module to obtain the trained first large language model.
[0031] In this embodiment, the sparse regularization constraint module is a functional module used to control parameter growth and improve parameter efficiency during model training. The functional implementation process of the sparse regularization constraint module includes: 1) Regularization method selection: Selecting a suitable sparse regularization method, such as L1 regularization and L0 regularization. L1 regularization promotes the sparsity of model parameters by adding the L1 norm of the model parameters as a regularization term to the loss function; L0 regularization directly constrains the number of non-zero parameters in the model, but since the L0 norm is not differentiable, approximation methods are usually used in practical applications, such as using the L1 norm as an approximation of the L0 norm. Here, L1 regularization is used as an example for explanation.
[0032] 2) Loss Function Construction: Based on the original model training loss function (such as the cross-entropy loss function for classification tasks), an L1 regularization term is added. The new loss function can be expressed as: ,in λ is the original loss function, λ is the regularization coefficient used to balance the weights of the original loss and the regularization term, and n is the total number of model parameters. It is the i-th parameter of the model.
[0033] 3) Training Process Control: During model training, optimization algorithms (such as stochastic gradient descent (SGD), Adam, etc.) are used to optimize the new loss function. In each iteration updating the model parameters, not only the gradient of the original loss function with respect to the parameters but also the gradient of the L1 regularization term with respect to the parameters are considered. The gradient of the L1 regularization term pushes the parameters towards zero, thus gradually sparsifying the model parameters. Based on the generated optimal rank allocation scheme (specified rank allocation scheme), different degrees of regularization constraints are applied to the parameters of different network layers during training. For network layers with low semantic importance and low rank allocation, a larger regularization coefficient λ can be set to accelerate the sparsification of these network layer parameters; for network layers with high semantic importance and high rank allocation, a smaller regularization coefficient λ is set to preserve the important parameters of these network layers as much as possible.
[0034] 4) Parameter efficiency evaluation: During training, the parameter efficiency of the model should be evaluated periodically. Parameter efficiency can be assessed by calculating metrics such as the proportion of non-zero parameters in the model and the ratio of the total number of parameters to performance metrics (such as accuracy and recall). If the parameter efficiency does not meet expectations, the value of the regularization coefficient λ can be adjusted, and training can be repeated.
[0035] The sparse regularization constraint module introduces regularization terms during model training to encourage model parameters to become sparser, thereby reducing redundant parameters and improving parameter efficiency. Combined with the optimal rank allocation scheme generated in the preceding steps, targeted parameter control can be applied to network layers of varying importance, compressing the model size as much as possible while maintaining performance.
[0036] In addition, the specific model training process includes: Input data: The training dataset (features X and labels Y) is fed into the large language model.
[0037] Forward propagation: The model calculates the predicted value Y^ and calculates the loss based on the original loss function (such as cross-entropy). .
[0038] Add regularization terms: in Adding an L1 regularization term to the above, we obtain the total loss. .
[0039] Backpropagation and parameter update: computation For parameters The gradients (including the gradients of the original loss and the regularization term) are used to update the parameters using an optimization algorithm (such as SGD or Adam), where the gradients of the L1 regularization push the parameters toward zero sparsity.
[0040] Hierarchical regularization control: Based on the optimal rank allocation scheme in the previous steps, different λ values are set for different layers (e.g., a larger λ is used for lower rank layers to accelerate sparsity).
[0041] Iterative training: Repeat the above process until the model converges or reaches the preset number of iterations, thus obtaining the first trained language model.
[0042] Step S205: Based on the preset test dataset, the first large language model is fine-tuned to obtain the corresponding second large language model, and the semantic fidelity of the second large language model is verified.
[0043] In this embodiment, after the model is trained, it enters the fine-tuning stage. During the fine-tuning stage, a portion of labeled financial corpus is used to further train the model to improve its performance on specific financial tasks, resulting in a fine-tuned second language model. During fine-tuning, a mini-batch gradient descent method is employed, updating the model parameters with a small subset of samples each time. After each batch of training in the fine-tuning stage, the model is evaluated using pre-designed financial semantic fidelity metrics (such as accuracy, business logic, and performance metrics). The model trained in the current batch is then tested on a validation set, and the values of each evaluation metric are calculated.
[0044] One approach is to design a set of semantic fidelity evaluation metrics suitable for the financial sector, tailored to specific business needs. This can be considered from multiple perspectives. For example, for financial text classification tasks, the evaluation can assess whether the model accurately captures key semantic information within the financial text, such as a company's financial situation and industry trends. Traditional metrics like precision, recall, and F1 score can be combined with specific financial sector metrics, such as accuracy in identifying financial entities or classifying financial events. Alternatively, manual evaluation can be employed, inviting financial experts to assess the semantic consistency of the model's output, determining whether the model accurately understands the semantics of the input financial text.
[0045] Furthermore, the specific implementation process of semantic fidelity verification of the second language model will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0046] Step S206: If the second language model fails the semantic fidelity verification, the specified rank allocation scheme is adjusted based on a preset adjustment strategy to obtain the corresponding target rank allocation scheme.
[0047] In this embodiment, the process of adjusting the specified rank allocation scheme includes: determining whether the semantic expressive power of the current model meets the requirements based on the semantic fidelity evaluation results obtained from real-time verification. If the model is found to perform poorly in certain financial semantic aspects, such as having a low accuracy rate in identifying specific financial entities, the possible reasons are analyzed. If it is because the rank assigned to a network layer with high semantic importance is too low, causing the layer to be unable to fully capture relevant semantic information, the rank value of the network layer is appropriately increased; conversely, if the rank assigned to a network layer with low semantic importance is too high and does not contribute much to improving model performance, the rank value of the network layer is appropriately decreased.
[0048] During model fine-tuning, the fidelity of financial semantics is verified in real time to promptly identify problems in the model's semantic expression, and the rank allocation strategy is dynamically adjusted based on these problems. This ensures that the model maintains its accurate understanding and expression of financial semantics throughout the compression and optimization process, avoiding semantic loss due to over-compression.
[0049] Step S207: Based on the preset iterative processing strategy, the target rank allocation scheme is used to perform model training, fine-tuning and evaluation on the large language model until the preset iterative termination condition is met and the corresponding target large language model is obtained.
[0050] In this embodiment, the iterative processing strategy includes: 1) Iterative optimization process design: Steps S203-S206 are combined into a complete iterative optimization process. In the first iteration, step S203 is executed first to generate an initial optimal rank allocation scheme. Then, step S204 is executed according to this scheme for model training and parameter control. Next, step S205 is executed for fine-tuning and semantic fidelity verification. Finally, step S206 is executed to adjust the rank allocation strategy. In each subsequent iteration, the optimization result obtained in the previous iteration is used as the basis to re-execute steps S203-S206. For example, in the second iteration, the optimal rank allocation scheme is regenerated according to the rank allocation strategy adjusted in the first iteration, and then model training, fine-tuning, and semantic fidelity verification are performed again.
[0051] 2) Setting Balanced Evaluation Metrics: Establish a set of metrics to evaluate the balance between parameter efficiency and semantic fidelity. This can comprehensively consider the number of model parameters, the model's performance metrics on financial tasks (such as accuracy and recall), and financial semantic fidelity evaluation metrics. For example, a comprehensive score can be calculated by weighting and summing the percentage reduction in the number of parameters, the percentage improvement in performance metrics, and the percentage change in semantic fidelity metrics to obtain a comprehensive balanced score.
[0052] 3) Setting the iteration termination condition: Set the iteration termination condition, such as stopping the iterative optimization process when the change of the comprehensive balance score in several consecutive iterations is less than a preset threshold, or when the preset maximum number of iterations is reached.
[0053] 4) Determining the optimal configuration: After the iterative optimization process is completed, the configuration with the highest overall balance score is selected from all iteration results as the optimal configuration that balances parameter efficiency and semantic fidelity. This configuration includes the optimal rank allocation scheme, model parameters, and corresponding training and fine-tuning parameter settings.
[0054] Specifically, based on the iterative processing strategy described above, the target rank allocation scheme can be used to train, fine-tune, and evaluate the large language model until the preset iteration termination condition is met and the corresponding target large language model is obtained. Furthermore, through multiple rounds of iterative optimization, the model can continuously find the optimal balance between parameter efficiency and semantic fidelity. Each iteration improves upon the results of the previous round, gradually optimizing the rank allocation scheme and model parameters, so that the model can express financial semantics as accurately as possible while meeting parameter budget requirements.
[0055] Step S208: Process the task data to be processed based on the target large language model.
[0056] In this embodiment, the final target large language model, after fine-tuning, can be applied to task processing scenarios in the fintech field. Here are two specific examples: 1. Automated Insurance Claims Processing. Task Description: During the insurance claims process, customers need to submit a claim application and relevant supporting documents. The insurance company then needs to review these documents to determine the claim amount and whether to approve the claim. This process often involves processing a large amount of textual materials, such as claim forms, medical reports, accident certificates, etc.
[0057] Applications of the Large Language Model: Text Parsing and Information Extraction: The finely tuned large language model can automatically parse claim applications and related supporting documents, extracting key information such as the time and location of the accident, the extent of the loss, and medical expenses. Intelligent Review and Decision Support: Based on the extracted information, the model can automatically determine the compliance of the claim application by comparing it with insurance terms and claim rules, and provide preliminary claim amount suggestions. This helps reduce the workload of manual review and improves the efficiency of claims processing. Customer Communication and Explanation: The model can also generate detailed claim reports to explain the reasons and basis for the claim decision to customers, improving customer satisfaction.
[0058] Real-world example: An insurance company used a finely tuned large language model to automate the processing of claims. Through the model's intelligent parsing and review of claim materials, claim processing time was significantly reduced, while also minimizing human error and fraud risks.
[0059] 2. Insurance Product Recommendation and Personalized Service. Task Description: In the insurance sales process, sales personnel need to recommend suitable insurance products based on the client's specific needs and risk profile. This process requires sales personnel to possess extensive product knowledge and excellent communication skills.
[0060] Applications of the Large Language Model: Customer Needs Analysis: The finely tuned large language model can gain a deeper understanding of customers' insurance needs, risk preferences, and financial situations through interaction with their natural language. Intelligent Product Recommendation: Based on the results of customer needs analysis, the model can automatically filter the most suitable insurance products for the customer and generate a detailed recommendation report. The report can include information such as product features, coverage, and premium calculations to help customers make informed purchasing decisions. Personalized Service and Support: The model can also provide continuous personalized service and support based on customer feedback and purchase history, such as renewal reminders and claims assistance.
[0061] Real-world case study: An insurance company built an intelligent insurance recommendation system using a finely tuned large language model. This system can automatically recommend suitable insurance products based on customers' specific needs and risk profiles, and provide personalized services and support. Through this system, the insurance company's sales efficiency has significantly improved, while customer satisfaction has also increased substantially.
[0062] This application first acquires pre-collected financial corpus data; then, it processes the financial corpus based on a pre-defined large language model, and analyzes the semantic importance information of each network layer contained in the large language model based on the obtained model output data; subsequently, it uses a pre-defined dynamic rank allocation engine to generate a scheme based on the semantic importance information and a pre-defined parameter budget, obtaining a corresponding specified rank allocation scheme; and then, based on a pre-defined training dataset and the specified rank allocation scheme, it trains the large language model using a pre-defined sparse regularization constraint module to obtain a trained first large language model; subsequently, it fine-tunes the first large language model based on a pre-defined test dataset to obtain a corresponding second large language model, and performs semantic fidelity verification on the second large language model; if the second large language model fails the semantic fidelity verification, it adjusts the specified rank allocation scheme based on a pre-defined adjustment strategy to obtain a corresponding target rank allocation scheme; further, based on a pre-defined iterative processing strategy, it uses the target rank allocation scheme to train, fine-tune, and evaluate the large language model until a pre-defined iteration termination condition is met and a corresponding target large language model is obtained; finally, it processes the task data to be processed based on the target large language model. Unlike existing methods that use fixed-rank configuration, this application employs an automated processing flow centered around four core steps: semantic importance assessment, dynamic rank allocation, sparse regularization constraints, and financial semantic fidelity verification. These steps collaborate to optimize and adapt the large language model, achieving efficient fine-tuning of the large language model within a limited parameter budget. This balances parameter efficiency with semantic fidelity, enabling the subsequent use of the efficiently fine-tuned target large language model to process task data. This effectively improves the accuracy and reliability of the target large language model in relevant task processing.
[0063] In some alternative implementations, step S202 includes the following steps: The financial corpus data is cleaned to obtain the corresponding target corpus data.
[0064] In this embodiment, the aforementioned financial corpus data is cleaned to remove noise data, such as special characters and irrelevant advertising information. Then, word segmentation is performed. For financial terminology, such as company names and financial terms, a specialized financial dictionary is used for accurate segmentation to ensure accuracy. Furthermore, the segmented corpus data is converted into a format suitable for input to a large language model to obtain the target corpus data. For example, text sequences are converted into numerical word vector representations. Pre-trained financial domain word vector models, such as Word2Vec or GloVe models trained on financial corpora, can be used to map each word to a fixed-dimensional vector space.
[0065] The target corpus data is processed based on the large language model, and the output features of each network layer in the large language model are recorded.
[0066] In this embodiment, the aforementioned large language model can be a pre-trained language model, such as a general-purpose language model like BERT or GPT, or a specialized language model pre-trained for the financial field. These models have already learned rich linguistic knowledge on large-scale text data. The pre-trained model weights are loaded to facilitate subsequent analysis of the semantic information of each network layer.
[0067] The cleaned target corpus data is input into the aforementioned large language model (referred to as the model for short), which then performs forward propagation computation on each input data. During the computation, the output features of each network layer of the model are recorded. For the Transformer architecture model, each network layer (such as the encoder layer of the Transformer) outputs a feature matrix containing rich semantic information.
[0068] Information extraction processing is performed based on the output features of each network layer to obtain the semantic information of each network layer.
[0069] In this embodiment, semantic information can be extracted from the output features of each network layer using various feature analysis methods. For example, Principal Component Analysis (PCA) can be used to reduce the dimensionality of the feature matrix and extract the principal components, which can be considered as the important semantic features captured by the network layer. Alternatively, attention weight analysis can be used. For attention-based models, the distribution of attention weights at different positions of each attention head can be analyzed. Positions with larger attention weights indicate that the model pays more attention to the semantic information at those positions in that layer.
[0070] The semantic importance of each network layer is evaluated to obtain the semantic importance information of each network layer.
[0071] In this embodiment, the specific implementation process of performing semantic importance evaluation on the semantic information of each network layer to obtain the semantic importance information of each network layer will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.
[0072] This application cleanses financial corpus data to obtain corresponding target corpus data. Then, it processes the target corpus data using a large language model and records the output features of each network layer in the model. Next, it extracts information based on the output features of each network layer to obtain the semantic information of that layer. Finally, it evaluates the semantic importance of each network layer to obtain its semantic importance information. In this way, this application can gain a deeper understanding of the ability of each network layer in the large language model to capture and express semantic information when processing financial corpus data, ensuring the accuracy of the obtained semantic importance information. Furthermore, by analyzing the semantic importance of each network layer, a basis can be provided for subsequent dynamic rank allocation, because network layers with different semantic importance may require different processing methods during model compression and optimization. Network layers with high semantic importance may need to retain more information, while network layers with low semantic importance can be compressed to a certain extent.
[0073] In some optional implementations of this embodiment, the step of performing semantic importance evaluation on the semantic information of each network layer to obtain the semantic importance information of each network layer includes the following steps: Call the preset scoring function.
[0074] In this embodiment, the scoring function is a pre-constructed function used to calculate the semantic importance of each network layer. The scoring function can consider multiple factors, such as the principal component variance contribution rate of features (summing the variance contribution rates of all principal components of a network layer to obtain the proportion of total variance captured by that layer), and the concentration of attention weights (calculating the distribution concentration of attention weights, such as the Gini coefficient or the reciprocal of entropy). For example, for a network layer, the proportion of the sum of the variances of all its principal components to the total variance is calculated; a higher proportion indicates that the layer extracts richer semantic information and has higher semantic importance. Preferably, the scoring function is used to weightedly combine multiple indicators (such as principal component variance contribution rate and attention score) to obtain the final semantic importance score.
[0075] Obtain the specified semantic information of a specified network layer.
[0076] In this embodiment, the specified network layer is any one of all network layers contained in the large language model.
[0077] The specified semantic information is calculated and processed based on the scoring function to obtain the corresponding scoring data.
[0078] In this embodiment, the semantic importance score of the specified semantic information can be calculated based on the above scoring function to obtain the corresponding score data.
[0079] The scoring data is normalized to obtain the corresponding target scoring data.
[0080] In this embodiment, the calculated score data is normalized to map the score range to the [0, 1] interval, so that the semantic importance of different network layers is comparable, and finally a normalized semantic importance score vector is generated to represent the relative importance of each network layer (range [0, 1]).
[0081] The target score data is used as the specified semantic importance information for the specified network layer.
[0082] Based on the above processing flow, this application obtains specified semantic information of a specified network layer, then calculates and processes the specified semantic information based on the use of a scoring function to obtain scoring data, then normalizes the scoring data, and uses the obtained target scoring data as the specified semantic importance information of the specified network layer. This allows for efficient and accurate evaluation of the semantic importance of the semantic information of the network layer, ensuring the accuracy of the generated semantic importance information.
[0083] In some alternative implementations, step S203 includes the following steps: The dynamic rank allocation engine calls a preset mapping table.
[0084] In this embodiment, the aforementioned dynamic rank allocation engine is a pre-built automated tool used to assist in the generation of rank allocation schemes. The aforementioned mapping table is a pre-built data table based on semantic importance-rank mapping relationships. The construction process of the semantic importance-rank mapping relationships in the aforementioned mapping table includes: 1. Data Collection. Determine the semantic importance range: Define the semantic importance score range as 0-100. This range can be set according to actual needs and model characteristics, aiming to comprehensively cover situations with different levels of semantic sensitivity. Design experiments: Construct a series of test scenarios with different semantic importance. For example, in the financial field, financial texts and tasks of varying complexity and professionalism can be selected as experimental objects. For each experimental object, set different semantic importance levels, which can be determined through manual annotation, expert evaluation, etc. Collect model performance data: For each semantic importance level, conduct multiple experiments and record the model's performance indicators under different rank configurations, such as task accuracy and semantic fidelity. For example, in a securities investment consulting model, for a financial text understanding task with a specific semantic importance, set different rank values such as 16, 32, 48, and 64, run the model multiple times, calculate the task accuracy and semantic fidelity of each run, and take the average as the model performance data under that rank configuration.
[0085] 2. Data Analysis and Fitting. Data Preparation: Organize the collected data on model performance and rank configuration under different semantic importance levels into a dataset. Each row of the dataset records a semantic importance level, its corresponding rank configuration, and a model performance metric. Selection of Fitting Method: Based on the characteristics and distribution of the data, select an appropriate regression analysis method, such as linear regression or multinomial regression. If the data exhibits a clear linear relationship, linear regression can be chosen; if the data relationship is more complex, multinomial regression or other more complex regression methods may be necessary. Fitting the Functional Relationship: Using the selected regression analysis method, with semantic importance as the independent variable and optimal rank as the dependent variable, fit the dataset to obtain the functional relationship between semantic importance and optimal rank. For example, if analysis reveals that the data roughly exhibits piecewise linear characteristics, a piecewise linear function can be fitted. For instance, when the semantic importance score is between 0 and 30, the corresponding optimal rank is 16; when the semantic importance score is between 31 and 70, the optimal rank is 48; and when the semantic importance score is between 71 and 100, the optimal rank is 64.
[0086] The rank of each network layer corresponding to the semantic importance information of each network layer is retrieved from the mapping relationship, and an initial rank allocation scheme is constructed based on the rank of each network layer.
[0087] In this embodiment, the semantic importance-rank mapping relationship in the above mapping table can be queried based on the semantic importance information of each network layer, so as to assign a corresponding rank to each network layer. For example, the semantic importance score of the middle layer is 75, which belongs to the range of 71-100, and the rank is assigned as 64; the semantic importance score of the top layer is 60, which belongs to the range of 31-70, and the rank is assigned as 48; the semantic importance score of the bottom layer is 20, which belongs to the range of 0-30, and the rank is assigned as 16.
[0088] The total number of parameters is obtained by calculating the parameters based on the rank of each network layer.
[0089] In this embodiment, the calculation process for the above parameters includes: determining the parameter quantity calculation formula: assuming the number of parameters in each layer is proportional to the square of the rank, let the rank of the i-th layer be ri, then the number of parameters in that layer is pi = k × ri^2, where k is a constant that can be determined experimentally or empirically. Calculating the total number of parameters under the current rank allocation scheme: based on the rank of the previous layered allocation, calculate the number of parameters in each layer, and then add the number of parameters in all layers to obtain the total number of parameters under the current rank allocation scheme.
[0090] Obtain the parameter budget of the large language model, and analyze the total number of parameters and the parameter budget to obtain the corresponding analysis results.
[0091] In this embodiment, the aforementioned parameter budget refers to the available parameter budget of the large language model, which may be, for example, 5% of the original model. The calculated total number of parameters can be compared and analyzed to obtain corresponding analysis results. These results may include whether the total number of parameters exceeds the parameter budget or is lower than the parameter budget.
[0092] Based on the analysis results, the initial rank allocation scheme is adjusted to obtain a first rank allocation scheme that meets the budget requirements.
[0093] In this embodiment, if the analysis result indicates that the total number of parameters exceeds the parameter budget, the adjustment process includes: reducing the rank of semantically less important layers while ensuring that the rank of layers with higher semantic importance remains unchanged or changes as little as possible. For example, first try reducing the rank of the lower layers, gradually decreasing the rank value by a certain step size, and recalculating the total number of parameters after each reduction until the total number of parameters meets the budget requirement. If the total number of parameters is lower than the parameter budget, the adjustment process includes: appropriately increasing the rank of some key layers to further improve model performance. For example, select layers with high semantic sensitivity and a significant impact on model performance, gradually increasing the rank value by a certain step size, and recalculating the total number of parameters after each increase until the total number of parameters approaches the budget limit or achieves the expected performance improvement.
[0094] The first rank allocation scheme is used as the designated rank allocation scheme.
[0095] This application utilizes a dynamic rank allocation engine to call a pre-defined mapping table. It then retrieves the rank of each network layer corresponding to its semantic importance information from the mapping table, and constructs an initial rank allocation scheme based on the ranks of each network layer. Next, it calculates the total number of parameters based on the ranks of each network layer. Subsequently, it obtains the parameter budget of the large language model and analyzes the total number of parameters against the parameter budget to obtain corresponding analysis results. Further, it adjusts the initial rank allocation scheme based on the analysis results to obtain a first rank allocation scheme that meets the budget requirements. Finally, it uses the first rank allocation scheme as the designated rank allocation scheme. Based on this process, this application calculates the total number of parameters under the current rank allocation scheme and compares it with the parameter budget of the large language model. Adjusting the ranks of each layer based on the analysis results allows for maximizing model performance while meeting resource constraints, ensuring the accuracy and adaptability of the obtained designated rank allocation scheme. Furthermore, this optimization method avoids resource waste due to excessive model parameters and ensures optimal model performance with limited resources.
[0096] In some alternative implementations, step S203 includes the following steps: The dynamic rank allocation engine invokes a preset dynamic rank allocation model.
[0097] In this embodiment, the aforementioned dynamic rank allocation engine is a pre-built automated tool used to assist in the generation of rank allocation schemes.
[0098] This involves constructing a dynamic rank allocation model. This model takes the semantic importance score and parameter budget of each network layer as input and outputs the optimal rank allocation for each network layer. An optimization algorithm can be used to construct this model, with a genetic algorithm being preferred. The specific implementation process includes: 1) Population Initialization: Randomly generate a set of initial rank allocation schemes, each represented as a vector, where each element corresponds to the rank value of a network layer. 2) Fitness Function Design: Design a fitness function to evaluate the quality of each rank allocation scheme. The fitness function can consider two factors: first, the closeness between the total number of model parameters calculated based on the current rank allocation scheme and the parameter budget; second, the matching degree between the semantic importance of each network layer and the rank allocation, for example, network layers with high semantic importance are assigned relatively higher ranks. The higher the fitness function value, the better the rank allocation scheme. 3) Selection Operation: Select a subset of excellent rank allocation schemes as parents based on the fitness function values to generate the next generation of schemes. Selection operations can use methods such as roulette wheel selection or tournament selection. 4) Crossover Operation: Perform a crossover operation on the selected parent schemes, randomly selecting some rank values from two parent schemes and exchanging them to generate new child schemes. 5) Mutation Operation: Perform a mutation operation on the child schemes, randomly changing the rank value of a certain network layer with a certain probability to increase the diversity of the population. 6) Iterative optimization: Repeat the selection, crossover and mutation operations until the preset number of iterations is reached or the convergence condition is met. The optimal individual obtained at this time is the optimal rank allocation scheme.
[0099] Obtain the parameter budget corresponding to the large language model.
[0100] In this embodiment, the parameter budget, or total parameter budget, of the large language model can be determined based on the actual application scenario and hardware resource limitations. The parameter budget can be measured by the total number of parameters of the large language model or the amount of memory occupied by the model. For example, if the goal is to deploy the model on a mobile device, a smaller parameter budget may be needed due to the limited memory of mobile devices.
[0101] The parameter budget and semantic importance information are allocated based on the dynamic rank allocation model to obtain the corresponding allocation results.
[0102] In this embodiment, the semantic importance score and parameter budget of each network layer are used as model input, and a dynamic rank allocation model is run. After multiple iterations of optimization, the optimal rank allocation scheme (allocation result) that satisfies the parameter budget and comprehensively considers semantic importance is obtained. This optimal rank allocation scheme assigns a specific rank value to each network layer for subsequent model compression and optimization.
[0103] The allocation result is used as the specified rank allocation scheme.
[0104] This application utilizes a dynamic rank allocation engine to invoke a pre-defined dynamic rank allocation model; then, it obtains the parameter budget corresponding to the large language model; subsequently, it allocates the parameter budget and semantic importance information based on the dynamic rank allocation model to obtain the corresponding allocation result; finally, it uses the allocation result as the specified rank allocation scheme. Based on the above processing flow, this application, by using the dynamic rank allocation model, can reasonably allocate the rank of each network layer according to the semantic importance of each network layer under a given parameter budget, effectively ensuring the accuracy and adaptability of the obtained specified rank allocation scheme.
[0105] In some optional implementations of this embodiment, the semantic fidelity verification of the second large language model in step S204 includes the following steps: The second language model is subjected to terminology consistency verification.
[0106] In this embodiment, the terminology consistency verification includes: constructing a test set: collecting a large number of financial professional terms and their standard explanations to construct a test set. The test set should cover financial terms of different fields and varying degrees of complexity to ensure a comprehensive evaluation of the model's accuracy in understanding financial professional terms. Model output and comparison: inputting the terms from the test set into the fine-tuned model, allowing the model to output its understanding of the term, and then comparing it with the standard explanation. The accuracy of the model's understanding can be calculated using manual or automatic comparison. For example, for each term, if the model's output explanation is completely consistent with the standard explanation, the term is considered correctly understood; if there is some deviation, a score is given based on the degree of deviation, and finally, the accuracy of all term understandings is calculated. If the accuracy rate is greater than or equal to a preset accuracy threshold, the second language model is deemed to have passed the terminology consistency verification; if the accuracy rate is less than the accuracy threshold, the second language model is deemed to have failed the terminology consistency verification and directly failed the semantic fidelity verification. The value of the aforementioned accuracy threshold is not limited and can be set according to actual business needs.
[0107] If the second language model passes the terminology consistency verification, then the business logic verification is performed on the second language model.
[0108] In this embodiment, the above-mentioned business logic verification includes: designing financial inference tasks: designing a series of financial inference tasks, such as portfolio optimization and risk assessment. These tasks should have practical application scenarios and be able to comprehensively test the logical correctness of the model in the financial inference tasks. Expert evaluation: inputting the financial inference tasks into the fine-tuned model and letting the model output inference results. Inviting experts in the financial field to evaluate the inference results of the model and determine whether its business logic is correct. For example, for a portfolio optimization task, experts evaluate whether the portfolio scheme given by the model is reasonable based on logical requirements such as investment objectives and risk tolerance. Wherein, if the business logic is correct, the second language model is determined to have passed the business logic verification; otherwise, if the business logic is correct, the second language model is determined to have failed the business logic verification and directly fails the semantic fidelity verification.
[0109] If the second largest language model passes the business logic verification, then the performance of the second largest language model will be verified.
[0110] In this embodiment, the performance verification includes: setting key financial tasks and performance indicators: setting a series of key financial tasks, such as securities price prediction and investment suggestion generation, and determining corresponding performance indicators, such as accuracy, recall, and F1 score. Real-time monitoring of performance changes: monitoring the performance indicators of these key tasks in each iteration of model fine-tuning. Recording the performance indicator values of the model after each iteration and comparing them with the values of the previous iteration. If the performance indicator of a certain task shows a significant decline, such as an accuracy drop exceeding a certain threshold (e.g., 5%), the second-largest language model is determined to have failed the performance verification, and a performance degradation warning can be issued, indicating that the rank allocation strategy needs to be adjusted. If the performance indicators of the task are normal, the second-largest language model is determined to have passed the performance verification.
[0111] If the second largest language model passes the performance verification, then the second largest language model is deemed to have failed the semantic fidelity verification; otherwise, the second largest language model is deemed to have failed the semantic fidelity verification.
[0112] In this embodiment, the second language model is determined to have failed semantic fidelity verification only if it is detected that the second language model has passed the terminology consistency verification, business logic verification, and performance verification; otherwise, the second language model is determined to have failed semantic fidelity verification.
[0113] Based on the above processing flow, this application verifies the model's accuracy in understanding financial terminology through terminology consistency verification; business logic verification tests the model's logical correctness in financial reasoning tasks; and performance verification monitors performance changes in key financial tasks in real time, promptly identifying performance degradation. These verification methods help dynamically adjust the rank allocation strategy during model fine-tuning, ensuring the model's accuracy and reliability in financial semantic processing.
[0114] In some optional implementations of this embodiment, after step S207, the electronic device may further perform the following steps: Obtain the optimal second-rank allocation scheme corresponding to the target large language model.
[0115] In this embodiment, the target large language model is the final model obtained after iteration based on the optimal rank allocation scheme (i.e., the second rank allocation scheme). This is because the optimal rank allocation scheme provides a good structural foundation for the model, and the fine-tuning process further optimizes the model's parameters based on this foundation. In most cases, the final model obtained after fine-tuning can be considered consistent with the model obtained after iteration based on the optimal rank allocation scheme.
[0116] The second rank allocation scheme is processed to generate a report based on a preset report generation strategy, resulting in a corresponding rank allocation analysis report.
[0117] In this embodiment, the report generation strategy includes the following: Rank Allocation Scheme Summary: A detailed record of the final optimal rank allocation scheme, including the rank value of each network layer, how the rank allocation scheme was generated, and which semantic importance scores and parameter budgets were referenced. Performance Comparison Analysis: A comparison of the model's performance metrics before and after optimization, such as accuracy, recall, and F1 score, as well as parameter efficiency metrics, such as the number of parameters and the proportion of non-zero parameters. The model's improvements in parameter efficiency and semantic fidelity are demonstrated through charts and textual explanations. Semantic Fidelity Analysis: An in-depth analysis of the model's performance in financial semantic fidelity, combined with the designed evaluation metrics, illustrates the model's ability and accuracy in capturing different financial semantics. Specific case studies can be used to demonstrate the model's understanding of financial text and the consistency between the output results and the actual semantics. Optimization Process Review: A review of the entire iterative optimization process, showing the changes in the rank allocation scheme, the improvement in model performance, and the improvement in semantic fidelity in each iteration. Visualization methods, such as plotting curves of various metrics during the iteration process, provide a more intuitive presentation of the optimization process and its effects.
[0118] Specifically, the report generation process can be performed on the second rank allocation scheme based on the strategy content of the above report generation strategy to generate the corresponding rank allocation analysis report.
[0119] The rank allocation analysis report is format-converted to obtain the corresponding target rank allocation analysis report.
[0120] In this embodiment, the generated rank allocation analysis report is output in document form. Common formats such as PDF and Word can be selected for conversion to obtain the converted target rank allocation analysis report, which is convenient for users to view and understand.
[0121] The target rank allocation analysis report is then processed for output.
[0122] In this embodiment, the target rank allocation analysis report can be generated and output to relevant users via email, message, or interface display to complete the output processing of the target rank allocation analysis report.
[0123] This application obtains the optimal second-rank allocation scheme corresponding to the target large language model; then, based on a preset report generation strategy, it generates a report on the second-rank allocation scheme to obtain a corresponding rank allocation analysis report; subsequently, it performs format conversion on the rank allocation analysis report to obtain the corresponding target rank allocation analysis report; finally, it outputs the target rank allocation analysis report. Based on the above processing flow, this application provides users with a fine-tuned model and a detailed rank allocation analysis report by organizing and summarizing the results of the optimization process of the optimal second-rank allocation scheme, thereby improving the generation efficiency and intelligence of the rank allocation analysis report. Furthermore, users can understand the advantages and limitations of the model, as well as the specific details of the optimization process, through the rank allocation analysis report, providing a reference for subsequent model applications and further optimization.
[0124] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.
[0125] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.
[0126] Furthermore, this application discloses a semantically sensitive adaptive low-rank adaptation scheme for large financial models, belonging to the field of financial artificial intelligence and model optimization technology. The core innovation lies in constructing a dynamic low-rank adaptation framework based on financial semantic sensitivity. Through three major technological breakthroughs—semantic importance assessment, dynamic rank allocation, and sparse regularization constraints—it solves the technical challenge of balancing parameter efficiency and semantic fidelity in the fine-tuning of large financial models.
[0127] Furthermore, through this application, financial institutions can achieve accurate adaptation of large language models in the financial professional field under strict parameter constraints, which not only ensures service quality but also controls deployment costs, providing reliable technical support for the intelligent transformation of finance. The improvements of this application include: (1) Significantly improved parameter efficiency: Under the same parameter budget, the accuracy of financial tasks is significantly improved compared with fixed-rank LoRA; (2) Excellent semantic fidelity: The understanding of financial professional terms is more accurate, and the correctness of business logic is significantly improved; (3) More reasonable resource allocation: Automatic identification and strengthening of the adaptation capacity of key layers significantly improves parameter utilization; (4) Significantly reduced deployment costs: While maintaining performance, the model fine-tuning parameters are significantly reduced, and the deployment resource requirements are greatly reduced.
[0128] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0129] It should be emphasized that, in order to further ensure the privacy and security of the aforementioned target large language model, the target large language model can also be stored in a node of a blockchain.
[0130] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).
[0131] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0132] Further reference Figure 3 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of a task processing device based on a large language model, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.
[0133] like Figure 3 As shown, the task processing device 300 based on a large language model described in this embodiment includes: a first acquisition module 301, a first processing module 302, a first generation module 303, a training module 304, a verification module 305, an adjustment module 306, a second processing module 307, and a third processing module 308. Wherein: The first acquisition module 301 is used to acquire pre-collected financial corpus data; The first processing module 302 is used to process the financial corpus based on a preset large language model, and analyze the semantic importance information of each network layer contained in the large language model based on the obtained model output data. The first generation module 303 is used to perform scheme generation processing on the semantic importance information and the preset parameter budget based on the preset dynamic rank allocation engine to obtain the corresponding specified rank allocation scheme. Training module 304 is used to train the large language model based on a preset training dataset and the specified rank allocation scheme, using a preset sparse regularization constraint module, to obtain a trained first large language model. The verification module 305 is used to fine-tune the first large language model based on a preset test dataset to obtain the corresponding second large language model, and to verify the semantic fidelity of the second large language model. The adjustment module 306 is used to adjust the specified rank allocation scheme based on a preset adjustment strategy to obtain the corresponding target rank allocation scheme if the second language model fails the semantic fidelity verification. The second processing module 307 is used to perform model training, fine-tuning and evaluation on the large language model using the target rank allocation scheme based on a preset iterative processing strategy, until the preset iterative termination condition is met and the corresponding target large language model is obtained. The third processing module 308 is used to process the task data to be processed based on the target large language model.
[0134] In some optional implementations of this embodiment, the first processing module 302 includes: The cleaning submodule is used to clean the financial corpus data to obtain the corresponding target corpus data. The first processing submodule is used to process the target corpus data based on the large language model and record the output features of each network layer in the large language model. The extraction submodule is used to perform information extraction processing based on the output features of each network layer to obtain the semantic information of each network layer. The evaluation submodule is used to perform semantic importance evaluation processing on the semantic information of each network layer to obtain the semantic importance information of each network layer.
[0135] In some optional implementations of this embodiment, the evaluation submodule includes: The calling unit is used to invoke a preset scoring function; An acquisition unit is used to acquire specified semantic information of a specified network layer; wherein, the specified network layer is any one of all network layers contained in the large language model; The calculation unit is used to calculate and process the specified semantic information based on the scoring function to obtain the corresponding scoring data; The normalization unit is used to normalize the scoring data to obtain the corresponding target scoring data; The determining unit is used to use the target score data as the specified semantic importance information of the specified network layer.
[0136] In some optional implementations of this embodiment, the first generation module 303 includes: The first calling submodule is used to call a preset mapping relationship table based on the dynamic rank allocation engine; The second processing submodule is used to query the rank of each network layer corresponding to the semantic importance information of each network layer from the mapping relationship, and construct the corresponding initial rank allocation scheme based on the rank of each network layer. The calculation submodule is used to calculate the parameters based on the rank of each network layer to obtain the corresponding total number of parameters; The analysis submodule is used to obtain the parameter budget of the large language model and analyze the total number of parameters and the parameter budget to obtain the corresponding analysis results; The adjustment submodule is used to adjust the initial rank allocation scheme based on the analysis results to obtain a first rank allocation scheme that meets the budget requirements. The first determining submodule is used to use the first rank allocation scheme as the specified rank allocation scheme.
[0137] In some optional implementations of this embodiment, the first generation module 303 includes: The second calling submodule is used to call a preset dynamic rank allocation model based on the dynamic rank allocation engine; The acquisition submodule is used to acquire the parameter budget corresponding to the large language model; The allocation submodule is used to allocate the parameter budget and the semantic importance information based on the dynamic rank allocation model to obtain the corresponding allocation result; The second determining submodule is used to use the allocation result as the specified rank allocation scheme.
[0138] In some optional implementations of this embodiment, the verification module 305 includes: The first verification submodule is used to perform terminology consistency verification on the second large language model; The second verification submodule is used to perform business logic verification on the second language model if the second language model passes the terminology consistency verification. The third verification submodule is used to perform performance verification on the second language model if the second language model passes the business logic verification. The determination submodule is used to determine whether the second largest language model has failed the semantic fidelity verification if the second largest language model passes the performance verification, and otherwise determine whether the second largest language model has failed the semantic fidelity verification.
[0139] In some optional implementations of this embodiment, the task processing device based on a large language model further includes: The second acquisition module is used to acquire the optimal second-rank allocation scheme corresponding to the target large language model; The second generation module is used to perform report generation processing on the second rank allocation scheme based on a preset report generation strategy to obtain the corresponding rank allocation analysis report. The conversion module is used to perform format conversion processing on the rank allocation analysis report to obtain the corresponding target rank allocation analysis report; The output module is used to process the output of the target rank allocation analysis report.
[0140] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4 This is a basic structural block diagram of the computer device in this embodiment.
[0141] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0142] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.
[0143] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may also include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for task processing methods based on large language models. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.
[0144] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions of the task processing method based on the large language model.
[0145] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.
[0146] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the task processing method based on the large language model described above.
[0147] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0148] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.
Claims
1. A task processing method based on a large language model, characterized in that, Includes the following steps: Acquire pre-collected financial corpus data; The financial corpus is processed based on a pre-defined large language model, and the semantic importance information of each network layer contained in the large language model is analyzed based on the obtained model output data. Based on a preset dynamic rank allocation engine, the semantic importance information and the preset parameter budget are used to generate a scheme to obtain the corresponding specified rank allocation scheme. Based on the preset training dataset and the specified rank allocation scheme, the large language model is trained using the preset sparse regularization constraint module to obtain the trained first large language model. The first language model is fine-tuned based on a pre-set test dataset to obtain the corresponding second language model, and the semantic fidelity of the second language model is verified. If the second largest language model fails the semantic fidelity verification, the specified rank allocation scheme is adjusted based on the preset adjustment strategy to obtain the corresponding target rank allocation scheme. Based on a preset iterative processing strategy, the target rank allocation scheme is used to train, fine-tune, and evaluate the large language model until the preset iterative termination condition is met and the corresponding target large language model is obtained. The target large language model is used to process the task data to be processed.
2. The task processing method based on a large language model according to claim 1, characterized in that, The step of analyzing the semantic importance information of each network layer contained in the large language model based on the obtained model output data specifically includes: The financial corpus data is cleaned to obtain the corresponding target corpus data; The target corpus data is processed based on the large language model, and the output features of each network layer in the large language model are recorded. Information extraction processing is performed based on the output features of each network layer to obtain the semantic information of each network layer. The semantic importance of each network layer is evaluated to obtain the semantic importance information of each network layer.
3. The task processing method based on a large language model according to claim 2, characterized in that, The step of performing semantic importance evaluation on the semantic information of each network layer to obtain the semantic importance information of each network layer specifically includes: Call the preset scoring function; Obtain specified semantic information from a specified network layer; wherein, the specified network layer is any one of all network layers contained in the large language model; The specified semantic information is calculated and processed based on the scoring function to obtain the corresponding scoring data; The scoring data is normalized to obtain the corresponding target scoring data; The target score data is used as the specified semantic importance information for the specified network layer.
4. The task processing method based on a large language model according to claim 1, characterized in that, The step of generating a scheme by processing the semantic importance information and the preset parameter budget based on the preset dynamic rank allocation engine to obtain the corresponding specified rank allocation scheme specifically includes: The dynamic rank allocation engine calls a preset mapping table; The rank of each network layer corresponding to the semantic importance information of each network layer is retrieved from the mapping relationship, and an initial rank allocation scheme is constructed based on the rank of each network layer. The total number of parameters is obtained by calculating the parameters based on the rank of each network layer. Obtain the parameter budget of the large language model, and analyze the total number of parameters and the parameter budget to obtain the corresponding analysis results; Based on the analysis results, the initial rank allocation scheme is adjusted to obtain a first rank allocation scheme that meets the budget requirements; The first rank allocation scheme is used as the designated rank allocation scheme.
5. The task processing method based on a large language model according to claim 1, characterized in that, The step of generating a scheme by processing the semantic importance information and the preset parameter budget based on the preset dynamic rank allocation engine to obtain the corresponding specified rank allocation scheme specifically includes: The dynamic rank allocation engine calls a preset dynamic rank allocation model; Obtain the parameter budget corresponding to the large language model; The parameter budget and semantic importance information are allocated based on the dynamic rank allocation model to obtain the corresponding allocation results; The allocation result is used as the specified rank allocation scheme.
6. The task processing method based on a large language model according to claim 1, characterized in that, The steps for semantic fidelity verification of the second language model specifically include: Term consistency verification was performed on the second largest language model; If the second largest language model passes the terminology consistency verification, then the business logic verification is performed on the second largest language model. If the second largest language model passes the business logic verification, then the performance of the second largest language model will be verified. If the second largest language model passes the performance verification, then the second largest language model is deemed to have failed the semantic fidelity verification; otherwise, the second largest language model is deemed to have failed the semantic fidelity verification.
7. The task processing method based on a large language model according to claim 1, characterized in that, After the step of training, fine-tuning, and evaluating the large language model using the target rank allocation scheme based on the preset iterative processing strategy until the preset iterative termination condition is met and the corresponding target large language model is obtained, the method further includes: Obtain the optimal second-rank allocation scheme corresponding to the target large language model; The second rank allocation scheme is processed to generate a report based on a preset report generation strategy, resulting in a corresponding rank allocation analysis report. The rank allocation analysis report is converted to a new format to obtain the corresponding target rank allocation analysis report. The target rank allocation analysis report is then processed for output.
8. A task processing device based on a large language model, characterized in that, include: The first acquisition module is used to acquire pre-collected financial corpus data; The first processing module is used to process the financial corpus based on a preset large language model, and to analyze the semantic importance information of each network layer contained in the large language model based on the obtained model output data. The first generation module is used to perform scheme generation processing on the semantic importance information and the preset parameter budget based on the preset dynamic rank allocation engine to obtain the corresponding specified rank allocation scheme. The training module is used to train the large language model based on a preset training dataset and the specified rank allocation scheme, using a preset sparse regularization constraint module, to obtain the trained first large language model. The verification module is used to fine-tune the first large language model based on a preset test dataset to obtain the corresponding second large language model, and to verify the semantic fidelity of the second large language model. The adjustment module is used to adjust the specified rank allocation scheme based on a preset adjustment strategy to obtain the corresponding target rank allocation scheme if the second largest language model fails the semantic fidelity verification. The second processing module is used to perform model training, fine-tuning and evaluation of the large language model using the target rank allocation scheme based on a preset iterative processing strategy, until the preset iterative termination condition is met and the corresponding target large language model is obtained. The third processing module is used to process the task data to be processed based on the target large language model.
9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the task processing method based on a large language model as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the task processing method based on a large language model as described in any one of claims 1 to 7.