Artificial intelligence-based task processing method and device, computer equipment and medium

By acquiring historical data and resource status information, and combining a multi-armed gambling machine optimizer and a precision risk control strategy, the optimal fine-tuning configuration scheme is automatically generated, which solves the performance degradation problem caused by improper configuration of large language models in financial tasks and improves the reliability and stability of the model.

CN121920449APending Publication Date: 2026-04-24PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2026-01-08
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Improper fine-tuning of large language models in the financial field leads to performance degradation, affecting their reliability and stability in critical tasks.

Method used

By acquiring historical fine-tuning data of the target task, the initial fine-tuning configuration scheme is determined using a task perceptron. Candidate configuration schemes are generated by combining a multi-armed gambling machine optimizer and a precision risk control strategy. Based on resource status information and cost-benefit data, the final target fine-tuning configuration scheme is obtained, thus realizing the automated fine-tuning of the large language model.

Benefits of technology

This improves the reliability and stability of large language models in financial task processing, ensuring model optimization in terms of resources and performance, and adapting to different task requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121920449A_ABST
    Figure CN121920449A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence, and relates to an artificial intelligence-based task processing method, which comprises the following steps of: analyzing an obtained task type and historical fine tuning data to determine an initial fine tuning configuration scheme; performing fine tuning on the large language model based on the initial fine tuning configuration scheme and collecting performance data; generating a plurality of candidate fine tuning configuration schemes based on the performance data; filtering the candidate fine tuning configuration schemes to obtain a first fine tuning configuration scheme; adjusting the first fine adjustment configuration scheme based on the resource state information to obtain a second fine adjustment configuration scheme; performing scheme screening on the second fine adjustment configuration scheme based on the cost data and the performance income data to obtain a target fine adjustment configuration scheme; and performing task processing on a target large language model obtained by performing fine tuning on the large language model by using the target fine tuning configuration scheme. The method can be applied to task processing scenes in the field of financial science and technology, and reliability and stability of the target large language model on related task processing can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology and can be applied to the financial technology field, particularly to artificial intelligence-based task processing methods, devices, computer equipment, and storage media. Background Technology

[0002] In the application of large language models in traditional finance, the fine-tuning of these models plays a crucial role in task performance. Currently, large language models, with their powerful language understanding and generation capabilities, are increasingly widely used in the financial field, covering multiple key tasks such as financial analysis, risk assessment, and customer service. However, improper fine-tuning can become a bottleneck restricting model performance, leading to performance degradation in key financial tasks and consequently reducing the reliability and stability of large language models in related tasks.

[0003] For example, in claims review scenarios within the financial and insurance sector, traditional methods rely primarily on manual review, which is not only inefficient but also susceptible to subjective biases from reviewers, leading to inconsistent and inaccurate results. Introducing a large language model (MLM) to assist in the review process, through fine-tuning its understanding of key elements such as insurance terms and claims information, can improve efficiency and accuracy. However, improper fine-tuning can cause the model to fail to accurately identify crucial information in claims, resulting in misjudgments or omissions. For instance, for claims involving complex disease clauses and ambiguous expressions, an improperly tuned model might incorrectly determine the validity of the claim, causing unnecessary losses and disputes for both the insurance company and the customer. Therefore, there is an urgent need to provide a scientifically sound and reasonable method for fine-tuning large language models to improve their performance in critical financial tasks and enhance their reliability and stability. Summary of the Invention

[0004] The purpose of this application is to propose a task processing method, apparatus, computer device, and storage medium based on artificial intelligence, so as to solve the technical problem that the reliability and stability of existing large language models are low in related tasks due to improper fine-tuning configuration.

[0005] Firstly, an artificial intelligence-based task processing method is provided, including: Obtain the task type of the target task, and obtain the historical fine-tuning data corresponding to the task type; Based on a preset task sensor, the task type and historical fine-tuning data are analyzed to determine the initial fine-tuning configuration scheme corresponding to the fine-tuning algorithm. Based on the initial fine-tuning configuration scheme, the preset large language model is fine-tuned, and the performance data of the large language model is collected. Based on the performance data, the exploration strategy of the preset multi-armed gambling machine optimizer is adjusted to generate multiple corresponding candidate fine-tuning configuration schemes; The candidate fine-tuning configuration schemes are filtered based on a preset accuracy risk control strategy to obtain the corresponding first fine-tuning configuration scheme. Obtain the current resource status information, and adjust the first fine-tuning configuration scheme based on the resource status information to obtain the corresponding second fine-tuning configuration scheme; Collect cost data and performance benefit data corresponding to the second fine-tuning configuration scheme, and perform scheme screening processing on the second fine-tuning configuration scheme based on the cost data and the performance benefit data to obtain the corresponding target fine-tuning configuration scheme; The target large language model is fine-tuned based on the target fine-tuning configuration scheme to obtain the corresponding target large language model, and the task data to be processed is processed based on the target large language model.

[0006] Secondly, an artificial intelligence-based task processing device is provided, comprising: The acquisition module is used to acquire the task type of the target task and acquire the historical fine-tuning data corresponding to the task type; The analysis module is used to analyze the task type and historical fine-tuning data based on a preset task sensor to determine the initial fine-tuning configuration scheme corresponding to the fine-tuning algorithm. The fine-tuning module is used to fine-tune the preset large language model based on the initial fine-tuning configuration scheme and collect the performance data of the large language model. The generation module is used to adjust the exploration strategy of the preset multi-armed gambling machine optimizer based on the performance data, so as to generate multiple candidate fine-tuning configuration schemes. The filtering module is used to filter the candidate fine-tuning configuration schemes based on a preset accuracy risk control strategy to obtain the corresponding first fine-tuning configuration scheme. The adjustment module is used to obtain the current resource status information and adjust the first fine-tuning configuration scheme based on the resource status information to obtain the corresponding second fine-tuning configuration scheme. The filtering module is used to collect cost data and performance benefit data corresponding to the second fine-tuning configuration scheme, and to perform scheme filtering processing on the second fine-tuning configuration scheme based on the cost data and the performance benefit data to obtain the corresponding target fine-tuning configuration scheme. The processing module is used to fine-tune the large language model based on the target fine-tuning configuration scheme to obtain the corresponding target large language model, and to process the task data to be processed based on the target large language model.

[0007] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the above-described artificial intelligence-based task processing method.

[0008] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the above-described artificial intelligence-based task processing method.

[0009] In the above-mentioned task processing method, apparatus, computer equipment, and storage medium based on artificial intelligence, the following steps are taken: First, the task type of the target task is obtained, and historical fine-tuning data corresponding to the task type is acquired. Then, based on a preset task perceptron, the task type and historical fine-tuning data are analyzed to determine an initial fine-tuning configuration scheme corresponding to the fine-tuning algorithm. Next, based on the initial fine-tuning configuration scheme, a preset large language model is fine-tuned, and the performance data of the large language model is collected. Subsequently, based on the performance data, the exploration strategy of a preset multi-armed gambling machine optimizer is adjusted to generate multiple corresponding candidate fine-tuning configuration schemes. Finally, based on a preset accuracy risk control strategy, the candidate... The first fine-tuning configuration scheme is obtained by filtering selected fine-tuning schemes. The current resource status information is then acquired, and the first fine-tuning configuration scheme is adjusted based on this information to obtain a second fine-tuning configuration scheme. Cost and performance benefit data corresponding to the second fine-tuning configuration scheme are collected, and the second fine-tuning configuration scheme is screened based on this data to obtain a target fine-tuning configuration scheme. Finally, the large language model is fine-tuned based on the target fine-tuning configuration scheme to obtain a target large language model, and the task data to be processed is then processed based on the target large language model. Based on the above automated processing flow, this application combines a multi-stage processing flow that adjusts the exploration strategy to generate candidate fine-tuning configuration schemes, filters schemes based on accuracy risk control strategies, adjusts schemes based on resource status information, and screens schemes based on cost and performance benefit data. By comprehensively considering multiple factors, it can automatically and accurately recommend the optimal target fine-tuning configuration scheme that is most suitable for the current target task, ensuring the accuracy and adaptability of the obtained target fine-tuning configuration scheme. Then, by using the target fine-tuning configuration scheme to fine-tune the large language model to obtain the target large language model, and processing the task data to be processed based on the target large language model, the reliability and stability of the target large language model in related task processing can be effectively improved. Attached Figure Description

[0010] To more clearly illustrate the solutions in this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is an exemplary system architecture diagram to which this application can be applied; Figure 2 This is a flowchart of an embodiment of the AI-based task processing method according to this application; Figure 3 This is a schematic diagram of a structure of an embodiment of the AI-based task processing device according to this application; Figure 4 This is a schematic diagram of the structure of one embodiment of the computer device according to this application. Detailed Implementation

[0012] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains; the terminology used herein in the specification of the application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings of this application, are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings of this application are used to distinguish different objects, not to describe a particular order.

[0013] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0014] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0015] like Figure 1As shown, system architecture 100 may include terminal device 101, network 102, and server 103. Terminal device 101 may be a laptop 1011, tablet 1012, or mobile phone 1013. Network 102 is used as a medium to provide a communication link between terminal device 101 and server 103. Network 102 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0016] Users can use terminal device 101 to interact with server 103 via network 102 to receive or send messages, etc. Various communication client applications can be installed on terminal device 101, such as web browser applications, shopping applications, search applications, instant messaging tools, email clients, social media platform software, etc.

[0017] Terminal device 101 can be various electronic devices with a display screen and support web browsing. In addition to laptops 1011, tablets 1012, or mobile phones 1013, terminal device 101 can also be an e-book reader, an MP3 player (Moving Picture Experts Group Audio Layer III), an MP4 player (Moving Picture Experts Group Audio Layer IV), a laptop computer, and a desktop computer, etc.

[0018] Server 103 can be a server that provides various services, such as a backend server that provides support for the pages displayed on terminal device 101.

[0019] It should be noted that the AI-based task processing method provided in this application embodiment is generally executed by a server / terminal device, and correspondingly, the AI-based task processing device is generally located in the server / terminal device.

[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included.

[0021] Continue to refer to Figure 2This document illustrates a flowchart of an embodiment of the AI-based task processing method according to this application. The order of steps in the flowchart can be changed, and some steps can be omitted, depending on different requirements. The AI-based task processing method provided in this application can be applied to any scenario requiring task processing, and therefore can be applied to products in these scenarios, such as task processing products in the financial insurance field. The AI-based task processing method includes the following steps: Step S201: Obtain the task type of the target task and obtain the historical fine-tuning data corresponding to the task type.

[0022] In this embodiment, the artificial intelligence-based task processing method runs on an electronic device (e.g., Figure 1 The server / terminal device shown can obtain the task type of the target task and the historical fine-tuning data corresponding to the task type through wired or wireless connection. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultrawideband) connections, and other currently known or future-developed wireless connection methods. The executing entity of this application is specifically a task processing system, which can be simply referred to as the system. This application can be applied to task processing scenarios in the fintech field. The aforementioned target task can specifically be a financial task related to a large language model with fine-tuning requirements. By identifying the task type of this financial task, the specific type of the financial task can be clarified. For example, it may include various natural language processing tasks, such as text generation, question answering, translation, summarization, as well as financial news classification, financial sentiment analysis, intelligent customer service, and other task types. Different tasks have different requirements for model performance and data characteristics; accurate task type identification is the foundation for subsequent data processing. Simultaneously, historical fine-tuning data of similar financial tasks corresponding to the aforementioned target task is collected, including hyperparameter configurations, model performance indicators, resource consumption, etc. This historical data will provide a reference for intelligent hyperparameter search, helping to quickly determine the initial configuration range.

[0023] Step S202: Analyze the task type and historical fine-tuning data based on the preset task perception to determine the initial fine-tuning configuration scheme corresponding to the fine-tuning algorithm.

[0024] In this embodiment, the fine-tuning algorithm can be either LoRA (Low-Rank Adaptation) or QLoRA (Quantized Low-Rank Adaptation). A pre-built task perceptron provides an initial LoRA / QLoRA hyperparameter configuration range corresponding to the fine-tuning algorithm based on the task type and historical fine-tuning data. For example, for a financial risk detection task, it might recommend a LoRA rank between 20 and 50, an alpha parameter between 32 and 128, and a quantization bit width between 3 and 6 bits. These ranges provide a general boundary for subsequent searches, avoiding blind searches. Furthermore, a multi-armed gambling machine optimizer analyzes the performance of different hyperparameter configurations in similar tasks based on historical experimental data. Within the range recommended by the financial task perceptron, it selects a relatively optimal initial fine-tuning configuration as the starting point for exploration. For example, based on historical data, it was found that a configuration with LoRA rank between 30 and 40, alpha parameter between 60 and 80, and quantization bit width of 4 bits performed well in risk detection tasks. Therefore, the recommended initial fine-tuning configuration is LoRA rank = 32, alpha = 64, and quantization bit width = 4 bits.

[0025] Step S203: Based on the initial fine-tuning configuration scheme, fine-tune the preset large language model and collect the performance data of the large language model.

[0026] In this embodiment, the aforementioned Large Language Model (LLM) can be a deep learning model with a large number of parameters, pre-trained on massive amounts of text data, corresponding to the task types described above. The LLM is based on a deep learning architecture, such as the Transformer. It learns basic language patterns, grammatical rules, semantic information, and world knowledge through unsupervised learning on a large-scale text corpus (covering various types of text such as books, articles, and web pages). This gives the model powerful language understanding and generation capabilities, enabling it to handle various natural language processing tasks, such as text generation, question answering, translation, and summarization. For example, the GPT series, BERT series, and LLaMA series of large language models can be used. These models absorb rich language knowledge during the pre-training stage and, after fine-tuning or adaptation to specific tasks, can demonstrate excellent performance in different scenarios.

[0027] Large language models are closely related to task type. While the pre-training goal of large language models is usually to learn the general rules of language, they can be adapted to specific tasks through fine-tuning or prompt engineering. For example, in the financial field, large language models can be fine-tuned for tasks such as financial news classification, financial sentiment analysis, and intelligent customer service. In the scenario of collecting model performance data in real time, large language models were used on financial validation sets. By fine-tuning them, they were able to better understand and process financial-related text and data, thereby evaluating the model's performance on financial tasks under different hyperparameter configurations. Furthermore, different tasks have different requirements for models, which affects the selection and use of large language models. For example, for tasks that require generating long texts, a large language model with strong generation capabilities and the ability to handle long sequences might be preferred; for tasks with high real-time requirements, a model with fewer parameters and faster inference speed might be chosen. In financial tasks, if the task involves understanding complex financial terminology and professional knowledge, it may be necessary to select a large language model that has been further pre-trained or fine-tuned with specific financial data to improve the model's performance in the financial domain.

[0028] Furthermore, during the fine-tuning of the aforementioned large language model using the initial fine-tuning configuration, the system collects real-time performance data of the large language model on the financial validation set, including performance metrics such as accuracy, recall, and F1 score, as well as loss function values ​​during training. This performance data reflects the actual performance of the model under the current hyperparameter configuration and serves as an important basis for adjusting the exploration strategy.

[0029] Step S204: Adjust the preset exploration strategy of the multi-armed gambling machine optimizer based on the performance data to generate multiple candidate fine-tuning configuration schemes.

[0030] In this embodiment, the specific implementation process of adjusting the preset multi-armed gambling machine optimizer based on the performance data to generate multiple candidate fine-tuning configuration schemes will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0031] Step S205: Based on a preset accuracy risk control strategy, the candidate fine-tuning configuration schemes are filtered to obtain the corresponding first fine-tuning configuration scheme.

[0032] In this embodiment, the specific implementation process of filtering the candidate fine-tuning configuration schemes based on the preset accuracy risk control strategy to obtain the corresponding first fine-tuning configuration scheme will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0033] Step S206: Obtain the current resource status information, and adjust the first fine-tuning configuration scheme based on the resource status information to obtain the corresponding second fine-tuning configuration scheme.

[0034] In this embodiment, the specific implementation process of obtaining the current resource status information and adjusting the first fine-tuning configuration scheme based on the resource status information to obtain the corresponding second fine-tuning configuration scheme will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0035] Step S207: Collect cost data and performance benefit data corresponding to the second fine-tuning configuration scheme, and perform scheme screening processing on the second fine-tuning configuration scheme based on the cost data and the performance benefit data to obtain the corresponding target fine-tuning configuration scheme.

[0036] In this embodiment, the specific implementation process of collecting cost data and performance benefit data corresponding to the second fine-tuning configuration scheme, and performing scheme screening processing on the second fine-tuning configuration scheme based on the cost data and performance benefit data to obtain the corresponding target fine-tuning configuration scheme will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0037] Step S208: Fine-tune the large language model based on the target fine-tuning configuration scheme to obtain the corresponding target large language model, and process the task data to be processed based on the target large language model.

[0038] In this embodiment, the large language model can be fine-tuned using the obtained target fine-tuning configuration scheme. By adjusting the model's internal parameters and optimizing its structure, the large language model can continuously learn and adapt to relevant business processing (such as financial risk detection), ultimately resulting in a final model that meets requirements in terms of performance (e.g., an expected accuracy of 92%), resource consumption, and cost—the target large language model. Then, the fine-tuned target large language model is used for actual business processing to process task data (such as financial transaction data). For example, the fine-tuned financial risk detection model can be deployed to an insurance company's transaction system. This requires integrating the target large language model into the existing business system to ensure that the model can work collaboratively with other system components. Furthermore, the deployed target large language model is used to perform real-time risk detection on financial transactions, providing support for the insurance company's business decisions. For example, when the target large language model detects a transaction with high credit risk, it promptly issues a warning to the insurance company's risk management department. The risk management department can then further assess and process the transaction based on the warning information from the target large language model, avoiding potential risk losses.

[0039] This application first obtains the task type of the target task and acquires historical fine-tuning data corresponding to the task type; then, it analyzes the task type and historical fine-tuning data based on a preset task perceptron to determine an initial fine-tuning configuration scheme corresponding to the fine-tuning algorithm; subsequently, it fine-tunes a preset large language model based on the initial fine-tuning configuration scheme and collects the performance data of the large language model; next, it adjusts the exploration strategy of a preset multi-armed gambling machine optimizer based on the performance data to generate multiple candidate fine-tuning configuration schemes; and filters the candidate fine-tuning configuration schemes based on a preset accuracy risk control strategy to obtain a corresponding first fine-tuning configuration scheme; further, it acquires current resource status information and adjusts the first fine-tuning configuration scheme based on the resource status information to obtain a corresponding second fine-tuning configuration scheme; then, it collects cost data and performance benefit data corresponding to the second fine-tuning configuration scheme and performs scheme screening based on the cost data and performance benefit data to obtain a corresponding target fine-tuning configuration scheme; finally, it fine-tunes the large language model based on the target fine-tuning configuration scheme to obtain a corresponding target large language model, and processes the task data to be processed based on the target large language model. Based on the above automated processing flow, this application combines a multi-stage processing flow that adjusts the exploration strategy to generate candidate fine-tuning configuration schemes, filters schemes based on accuracy risk control strategies, adjusts schemes based on resource status information, and screens schemes based on cost and performance benefit data. By comprehensively considering multiple factors, it can automatically and accurately recommend the optimal target fine-tuning configuration scheme that is most suitable for the current target task, ensuring the accuracy and adaptability of the obtained target fine-tuning configuration scheme. Then, by using the target fine-tuning configuration scheme to fine-tune the large language model to obtain the target large language model, and processing the task data to be processed based on the target large language model, the reliability and stability of the target large language model in related task processing can be effectively improved.

[0040] In some alternative implementations, step S204 includes the following steps: Based on a preset Bayesian optimization engine, a mapping model between hyperparameters and task performance is constructed according to the performance data.

[0041] In this embodiment, a mapping model between hyperparameters and financial task performance is constructed using a Bayesian optimization engine based on real-time performance data of a large language model during fine-tuning. This mapping model can be understood as a mathematical description of the relationship between hyperparameters and model performance, capable of predicting the potential performance level of the model under different hyperparameter configurations. For example, data analysis reveals that when the LoRA rank increases to 48 and alpha increases to 96, the model's accuracy on risk detection tasks tends to improve further. The Bayesian optimization engine captures this trend and updates its prediction of the optimal configuration.

[0042] Based on a preset iterative strategy, the mapping model is used to control the multi-armed gambling machine optimizer to adjust the exploration strategy corresponding to the hyperparameter configuration, and the hyperparameter search is performed based on the adjusted exploration strategy to obtain multiple corresponding first configuration schemes.

[0043] In this embodiment, the iterative strategy includes: 1) The Bayesian optimization engine feeds back the constructed mapping model and the prediction results for the optimal configuration to the multi-armed optimizer, guiding it to adjust its exploration strategy. Based on this information, the multi-armed optimizer decides whether to continue a fine-grained search around the currently performing hyperparameters or to explore some regions that have not yet been fully explored. For example, if the Bayesian optimization engine predicts that increasing the LoRA rank and alpha parameters may bring better performance, the multi-armed optimizer will adjust its exploration direction in this direction.

[0044] 2) Probabilistic Exploration and Exploitation Balance: The multi-armed gambling machine optimizer employs a probabilistic exploration and exploitation balance strategy. During the exploration phase, it attempts some hyperparameter configurations that haven't been fully explored yet, even if these configurations may not seem optimal at the moment. For example, it might randomly select some LoRA rank, alpha parameter, and quantization bit width combinations that are on the edge of the initial recommendation range or haven't been tried before for testing. During the exploitation phase, it prioritizes those hyperparameter configurations that perform well for further optimization based on existing information. Through this balance strategy, the multi-armed gambling machine optimizer can improve search efficiency by fully utilizing existing experience while exploring new possibilities.

[0045] 3) Dynamically Adjusting the Exploration Direction: Guided by the Bayesian optimization engine, the multi-armed gambling machine optimizer dynamically adjusts the exploration direction of hyperparameter configurations. It continuously evaluates the potential value of different hyperparameter combinations based on real-time data and feedback from the mapping model, determining the next search focus. For example, if a specific range of LoRA rank and alpha parameter combinations shows good performance in multiple trials, the multi-armed gambling machine optimizer will explore this range more deeply, trying different quantization bit width combinations to find the optimal configuration.

[0046] 4) Multi-round Trials and Iterations: Intelligent hyperparameter search is a multi-round iterative process. In each iteration, the multi-armed optimizer selects a new hyperparameter configuration to try based on the current exploration strategy, collects the model's performance data under that configuration, and feeds it back to the Bayesian optimization engine. The Bayesian optimization engine updates the mapping model and optimal configuration prediction based on the new data, and then guides the multi-armed optimizer to the next round of exploration. Through multiple iterations, the system can continuously accumulate experience and gradually approach the optimal hyperparameter configuration. Multiple potential candidate solutions are generated: During the iteration process, the system records the hyperparameter configurations tried in each round and their corresponding model performance. After multiple iterations, multiple configurations showing potential in different aspects are obtained. For example, it may be found that the configuration with rank = 48, alpha = 96, and bit width = 3 bits performs well in accuracy; while the configuration with rank = 45, alpha = 90, and bit width = 4 bits performs well in recall. These solutions all have certain potential and can be used as candidate configurations (i.e., the first configuration) for further evaluation and selection.

[0047] Specifically, based on the strategy content of the above iterative strategy, the above mapping model can be used to control the multi-armed gambling machine optimizer to adjust the exploration strategy corresponding to the hyperparameter configuration, and hyperparameter search can be performed based on the adjusted exploration strategy to obtain multiple corresponding first configuration schemes.

[0048] The first configuration scheme is filtered based on the preset evaluation requirements to obtain the filtered second configuration scheme.

[0049] In this embodiment, the aforementioned evaluation requirements include actual task requirements and evaluation metrics. The system can filter multiple candidate first configuration schemes based on the specific task requirements and evaluation metrics to obtain a filtered second configuration scheme. For example, if the task prioritizes accuracy, a configuration scheme with higher accuracy will be selected first; if both accuracy and recall need to be considered comprehensively, a comprehensive metric such as the F1 score may be used for evaluation and selection. Finally, a set of the most promising candidate configuration schemes is determined for subsequent model training and evaluation.

[0050] The second configuration scheme is used as the candidate fine-tuning configuration scheme.

[0051] This application constructs a mapping model between hyperparameters and task performance based on performance data using a pre-defined Bayesian optimization engine. Then, based on a pre-defined iterative strategy, the mapping model controls a multi-armed optimizer to adjust the exploration strategy corresponding to the hyperparameter configuration. Based on the adjusted exploration strategy, hyperparameter search is performed to obtain multiple first configuration schemes. Subsequently, the first configuration schemes are filtered based on pre-defined evaluation requirements to obtain second configuration schemes. These second configuration schemes are then used as candidate fine-tuning schemes. Based on this processing flow, this application, through the collaborative work of the Bayesian optimization engine and the multi-armed optimizer, fully utilizes historical and real-time data to achieve intelligent hyperparameter search, quickly and accurately generating multiple candidate fine-tuning configuration schemes. This improves the generation efficiency and intelligence of candidate fine-tuning configuration schemes, contributing to improved performance of large language models on related tasks.

[0052] In some optional implementations of this embodiment, step S205 includes the following steps: Collect key indicator data corresponding to the candidate fine-tuning configuration scheme.

[0053] In this embodiment, during the fine-tuning process, the performance of the large language model on key indicators such as financial compliance and risk identification accuracy is monitored in real time and continuously. For the financial risk detection model, its accuracy in identifying different types of risks (such as credit risk, market risk, and operational risk) is monitored in detail, while ensuring that the model strictly complies with financial regulations and regulatory requirements. Furthermore, data on these key indicators of the model at different time points and in different data batches are collected to form a real-time monitoring data stream. This process is like creating a detailed performance profile for each candidate fine-tuning configuration, recording its performance throughout the fine-tuning process, providing a comprehensive and detailed data foundation for subsequent selection.

[0054] A third configuration scheme is selected from the candidate fine-tuning configuration schemes, where the key indicator data is greater than the preset indicator threshold.

[0055] In this embodiment, a series of key indicator thresholds are pre-defined. These thresholds are determined based on the actual needs and risk tolerance of financial operations. For example, for risk identification accuracy, a threshold of 90% might be set; that is, when the accuracy is below 90%, the model performance is considered to have significantly declined. Furthermore, the collected key indicator data is evaluated. Once it is detected that the key indicator of a candidate fine-tuning configuration has decreased beyond the preset threshold—for example, the risk identification accuracy drops from 92% to 88%, exceeding the 2% threshold—the candidate fine-tuning configuration will be eliminated, and a third configuration with key indicator data greater than the threshold will be selected.

[0056] In addition, a rollback mechanism is automatically triggered for configuration schemes exhibiting poor performance. Once triggered, the model configuration of the candidate scheme is rolled back to a more conservative configuration. For example, if the original quantization bit width was 3 bits, it might be rolled back to 4 bits to attempt to restore model performance. This method promptly eliminates candidate configuration schemes whose performance significantly degrades at the current stage and fails to meet business requirements, narrowing the candidate pool. Furthermore, relevant information for each rollback event can be recorded, including rollback time, configuration parameters before and after the rollback, changes in key metrics, etc., and the risk pattern library is updated. These records not only help analyze the reasons for model performance degradation but also provide a reference for subsequent screening, preventing similar problems from recurring.

[0057] The third configuration scheme is filtered based on a preset risk pattern library to obtain a filtered fourth configuration scheme.

[0058] In this embodiment, the aforementioned risk pattern library is a database built based on historical data. This library stores various configuration patterns and risk characteristics that may lead to a decline in model performance. For example, by analyzing historical data, it is found that when the LoRA rank exceeds 50 and the quantization bit width is less than 3 bits, the model is prone to a decrease in accuracy on risk detection tasks. Such configuration schemes will then be marked as potential risk patterns.

[0059] During the fine-tuning process, the system monitors the current candidate configuration schemes in real time. When a similar configuration scheme to the current candidate scheme is detected in the risk pattern library, these potentially risky schemes are filtered out, directly excluding them to further optimize the candidate configuration range and obtain the corresponding fourth configuration scheme. Additionally, timely warnings can be issued to remind engineers to intervene. Engineers can further evaluate and adjust the corresponding candidate configuration schemes based on the warning information. Furthermore, as the fine-tuning process progresses, new rollback events and abnormal data will continuously emerge. The risk warning system will continuously analyze this data, adding new risk patterns to the risk pattern library, thus continuously improving and enriching the library. This allows subsequent screening processes to more accurately and comprehensively identify potential risks, improving the quality and efficiency of the screening process.

[0060] The fourth configuration scheme is used as the first fine-tuning configuration scheme.

[0061] This application collects key indicator data corresponding to candidate fine-tuning configuration schemes; then, it selects a third configuration scheme from the candidate schemes whose key indicator data exceeds a preset threshold; subsequently, it filters the third configuration scheme based on a preset risk model library to obtain a filtered fourth configuration scheme; and finally, it uses the fourth configuration scheme as the first fine-tuning configuration scheme. Based on the above processing flow, this application, through the synergistic effect of multiple stages—collecting key indicator data, screening schemes based on indicator thresholds, and risk assessment based on a risk model library—can gradually filter out fine-tuning configuration schemes that are poorly performing, have existing risks, or have potential risks. The final retained first fine-tuning configuration scheme performs relatively well in multiple aspects and can meet the needs of financial business and risk control requirements, providing a more reliable foundation for subsequently determining the optimal target fine-tuning configuration scheme.

[0062] In some alternative implementations, step S206 includes the following steps: Based on a preset hardware awareness configurator, the current resource status information is monitored in real time.

[0063] In this embodiment, a pre-built hardware-aware configurator is used to collect various resource status information in real time through system interfaces or dedicated hardware monitoring tools. This information covers computing resources (such as GPU utilization and memory usage), storage resources (disk read / write speed and remaining storage space), and network resources (network bandwidth and latency). For example, the current GPU load percentage can be obtained through system monitoring tools, and the remaining memory capacity can be obtained through the memory management module. Furthermore, the collected resource information is analyzed in depth to identify current resource bottlenecks and trends in resource usage. For instance, if GPU utilization is consistently high and trending upwards, while memory usage is also increasing, it may indicate that computing and memory resources will become limiting factors in subsequent training processes.

[0064] Obtain the dynamic adjustment strategy corresponding to the preset elastic training controller.

[0065] In this embodiment, the dynamic adjustment strategy includes: 1. Preliminary adjustment based on resource status. Resource allocation and candidate scheme association: Based on the resource status analysis results, different candidate configuration schemes are matched and associated with available resources. For configuration schemes with high resource requirements, such as those using large-scale neural network structures or high-precision data processing, they will be preferentially allocated to environments with abundant resources; while for schemes with relatively low resource requirements, they will be allocated to regions with relatively scarce resources but which can meet basic needs. Adjusting configuration parameters to adapt to resources: For each candidate configuration scheme, some of its parameters are adjusted according to the current resource status. For example, if computing resources are scarce, for some configuration schemes with high computing performance requirements, the complexity of the model can be appropriately reduced, such as reducing the number of layers or neurons in the neural network; if memory resources are limited, for configuration schemes that require a large amount of cached data, the data loading and caching strategy can be adjusted to reduce the amount of data loaded each time.

[0066] 2. Elastic Training Controller Intervention. Dynamic Learning Rate Adjustment: The elastic training controller dynamically adjusts the learning rate based on resource availability and real-time model training performance. When resources are sufficient and model training is progressing smoothly, the learning rate can be appropriately increased to accelerate model convergence. Conversely, if resources are limited or the model exhibits oscillations or overfitting, the learning rate is reduced to stabilize model training. For example, when low GPU utilization and slow model loss function decline are detected, the learning rate is increased by 10%; when GPU utilization is near its limit and the model loss function fluctuates significantly, the learning rate is reduced by 20%. Optimizer Selection and Parameter Optimization: Based on resource availability and model characteristics, the elastic training controller also selects and adjusts the optimizer's parameters. Different optimizers perform differently under different resource environments and model training stages. For example, when resources are limited, a computationally less computationally intensive optimizer, such as stochastic gradient descent (SGD), is selected; while when resources are abundant and fast convergence is desired, an adaptive optimizer, such as Adam, is chosen. Simultaneously, optimizer parameters, such as momentum coefficients and weight decay coefficients, are adjusted based on training progress to optimize model training effectiveness. Adjusting Batch Size: Batch size is a crucial factor affecting model training efficiency and resource utilization. The elastic training controller dynamically adjusts the batch size based on resource status. When computational resources are sufficient, increasing the batch size improves training efficiency and fully utilizes hardware resources; when resources are scarce, decreasing the batch size avoids training interruptions or performance degradation due to insufficient resources. For example, when GPU memory is sufficient, the batch size can be increased from 32 to 64; when memory is nearing full capacity, the batch size can be decreased to 16.

[0067] 3. Comprehensive Evaluation and Readjustment. Multi-dimensional evaluation of configuration scheme performance: After the above adjustments, the system will evaluate each candidate configuration scheme from multiple dimensions, including model performance metrics (such as accuracy, recall, F1 score, etc.), resource consumption (such as computation time, memory usage, disk read / write volume, etc.), and training stability (such as the trend of loss function changes, overfitting degree, etc.). Through comprehensive evaluation, a full understanding of the performance of each configuration scheme under the current resource conditions is obtained. Readjustment based on evaluation results: Based on the comprehensive evaluation results, the candidate configuration schemes are further adjusted. For configuration schemes with excellent performance, if resources permit, their parameters can be further optimized to pursue better performance; for configuration schemes with poor performance, the reasons are analyzed, and targeted adjustments are made or they are directly eliminated. For example, if the model performance of a certain configuration scheme improves after adjustment but the resource consumption is still high, its resource utilization strategy can be further optimized; if the performance of a certain configuration scheme still cannot meet the requirements after multiple adjustments, it is removed from the candidate list. Continuous monitoring and dynamic looping: Resource adaptive scheduling is a continuous process. The system will continuously monitor the resource status and model training status, and make dynamic adjustments based on real-time feedback. As training progresses, resource conditions may change, and new situations may arise in model performance. Therefore, it is necessary to continuously iterate the above adjustment process to ensure that candidate configuration schemes are always trained and evaluated in the optimal resource environment.

[0068] Based on the resource status information, the first fine-tuning configuration scheme is adjusted using the dynamic adjustment strategy to obtain the adjusted fifth configuration scheme.

[0069] In this embodiment, based on the obtained resource status information, parameter adjustment processing for the first fine-tuning configuration scheme can be performed according to the strategy content of the above dynamic adjustment strategy, and the obtained fifth configuration scheme can be used as the corresponding second fine-tuning configuration scheme.

[0070] The fifth configuration scheme is used as the second fine-tuning configuration scheme.

[0071] This application monitors the current resource status information in real time using a preset hardware-aware configurator; then, it obtains a dynamic adjustment strategy corresponding to a preset elastic training controller; subsequently, based on the resource status information, it uses the dynamic adjustment strategy to adjust the parameters of the first fine-tuning configuration scheme, obtaining an adjusted fifth configuration scheme; and finally, it uses the fifth configuration scheme as the second fine-tuning configuration scheme. Based on this processing flow, by using a hardware-aware configurator to monitor the current resource status information in real time, and then using the dynamic adjustment strategy corresponding to the preset elastic training controller to adjust the parameters of the first fine-tuning configuration scheme, and using the resulting adjusted fifth configuration scheme as the corresponding second fine-tuning configuration scheme, it is possible to intelligently adjust the first fine-tuning configuration scheme to obtain a second fine-tuning configuration scheme that meets the requirements of optimized resource utilization, ensuring the accuracy and adaptability of the obtained second fine-tuning configuration scheme.

[0072] In some alternative implementations, step S207 includes the following steps: Cost data corresponding to the first fine-tuning configuration scheme is collected based on preset monitoring tools.

[0073] In this embodiment, the aforementioned cost data includes resource usage time (i.e., GPU usage time) and power consumption data. Specifically, GPU usage time corresponding to the preset first fine-tuning configuration can be collected using a dedicated GPU monitoring tool. The GPU monitoring tool can obtain the real-time operating status of each GPU, including usage time. For example, during model fine-tuning, the cumulative usage time of each GPU is recorded at regular time intervals (e.g., every minute), thereby obtaining the total GPU usage time under different configurations.

[0074] Alternatively, power can be obtained by installing power monitoring devices on servers or computing equipment. These devices can accurately measure the power consumption of the equipment during operation. For example, a smart meter can be connected to the server's power input to record power consumption data in real time and correlate it with different model configurations for subsequent analysis of power costs under different configuration schemes.

[0075] Based on a preset model evaluation tool, performance gain data corresponding to the first fine-tuning configuration scheme is collected.

[0076] In this embodiment, the aforementioned performance gain data refers to model performance gain data. The performance gain data of the model under different configuration schemes can be evaluated using a representative financial validation set through the model evaluation module. Indicators such as accuracy improvement and recall improvement are calculated. For example, by comparing the accuracy and recall of the model on the validation set with the initial configuration and the current configuration, the performance improvement can be determined.

[0077] The cost data and performance benefit data are evaluated based on a preset evaluation strategy, and a specified fine-tuning configuration scheme that meets the requirements is selected from the first fine-tuning configuration scheme according to the evaluation results.

[0078] In this embodiment, the specific implementation process of evaluating the cost data and performance benefit data based on the preset evaluation strategy, and selecting the specified fine-tuning configuration scheme that meets the requirements from the first fine-tuning configuration scheme according to the obtained evaluation results, will be further described in detail in subsequent specific embodiments of this application, and will not be elaborated on here.

[0079] The specified fine-tuning configuration scheme is used as the target fine-tuning configuration scheme.

[0080] This application collects cost data corresponding to a first fine-tuning configuration scheme using a pre-set monitoring tool; then, it collects performance benefit data corresponding to the first fine-tuning configuration scheme using a pre-set model evaluation tool; subsequently, it evaluates and processes the cost and performance benefit data based on a pre-set evaluation strategy, and selects a specified fine-tuning configuration scheme that meets the requirements from the first fine-tuning configuration schemes based on the evaluation results; finally, it uses the specified fine-tuning configuration scheme as the target fine-tuning configuration scheme. Based on the above processing flow, this application collects cost and performance benefit data corresponding to a first fine-tuning configuration scheme by using monitoring and model evaluation tools, then evaluates and processes the cost and performance benefit data based on the evaluation strategy, and selects a configuration scheme that meets the requirements from the first fine-tuning configuration schemes as the corresponding target fine-tuning configuration scheme based on the evaluation results. This achieves the evaluation of the advantages and disadvantages of different configuration schemes by quantifying costs and benefits, providing a scientific basis for selecting the optimal configuration scheme, and ensuring the accuracy and intelligence of the obtained target fine-tuning configuration scheme.

[0081] In some optional implementations of this embodiment, the cost data includes resource usage time and power consumption data; the step of evaluating the cost data and performance benefit data based on a preset evaluation strategy, and selecting a specified fine-tuning configuration scheme that meets the requirements from the first fine-tuning configuration scheme according to the obtained evaluation results, includes the following steps: Call the preset cost-benefit calculation formula.

[0082] In this embodiment, the cost-benefit calculation formula corresponds to a predefined cost-benefit ratio indicator; it includes: Cost-benefit ratio = Performance improvement / (Computation cost + Power consumption cost). A comprehensive indicator (cost-benefit ratio) is pre-designed to comprehensively compare the performance improvement with computation cost and power consumption.

[0083] Based on the cost-benefit calculation formula, the resource usage time, the power consumption data, and the performance benefit data are calculated and processed to obtain the cost-benefit score corresponding to each of the second fine-tuning configuration schemes.

[0084] In this embodiment, the resource usage time, power consumption data and performance benefit data can be calculated based on the above cost-benefit calculation formula to calculate the cost-benefit score (i.e. cost-benefit ratio) corresponding to each second fine-tuning configuration scheme.

[0085] Specifically, the performance improvement is a weighted average of the accuracy and recall improvements (weights determined based on business needs). The computational cost is calculated based on GPU usage time and the cost per unit time, while the power consumption cost is calculated based on power consumption and the unit electricity price. For each configuration, the cost-benefit ratio is calculated using the above formula. For example, for configuration A, accuracy improves by 3%, recall improves by 2%, GPU usage time is 10 hours, the cost per unit time is 50 yuan, power consumption is 5 kWh, and the unit electricity price is 1 yuan / kWh. Assuming an accuracy weight of 0.6 and a recall weight of 0.4, the performance improvement is 0.6 × 3% + 0.4 × 2% = 2.6%, the computational cost is 10 × 50 = 500 yuan, and the power consumption cost is 5 × 1 = 5 yuan. The cost-benefit ratio is 500 + 52.6% ≈ 0.0051% / yuan.

[0086] The sixth configuration with the highest cost-effectiveness score is selected from all the second fine-tuning configuration options.

[0087] In this embodiment, the highest specified cost-benefit score can be obtained by comparing the calculated cost-benefit scores of each second fine-tuning configuration scheme, and the configuration scheme corresponding to the specified cost-benefit score can be obtained from all the second fine-tuning configuration schemes as the sixth configuration scheme.

[0088] The sixth configuration scheme is used as the designated fine-tuning configuration scheme.

[0089] This application uses a preset cost-benefit calculation formula; then, based on this formula, it calculates and processes resource usage time, power consumption data, and performance benefit data to obtain a cost-benefit score corresponding to each second fine-tuning configuration scheme; subsequently, it selects the sixth configuration scheme with the highest cost-benefit score from all second fine-tuning configuration schemes; and finally, it uses the sixth configuration scheme as the designated fine-tuning configuration scheme. Based on the above processing flow, this application uses a cost-benefit calculation formula to calculate and process resource usage time, power consumption data, and performance benefit data to obtain a cost-benefit score corresponding to each second fine-tuning configuration scheme, and then selects the sixth configuration scheme with the highest cost-benefit score as the corresponding designated fine-tuning configuration scheme. This intelligently and accurately achieves a comprehensive evaluation of different configuration schemes to recommend the optimal configuration scheme, effectively improving the accuracy and intelligence of the recommended designated fine-tuning configuration scheme.

[0090] In some optional implementations of this embodiment, the step of evaluating the cost data and performance gain data based on a preset evaluation strategy, and selecting a specified fine-tuning configuration scheme that meets the requirements from the first fine-tuning configuration scheme according to the obtained evaluation results, includes the following steps: Invoke the preset configuration recommendation engine.

[0091] In this embodiment, the configuration recommendation engine is a pre-built automated tool with configuration scheme recommendation function.

[0092] Based on the configuration recommendation engine, a hierarchical model corresponding to the cost data and the performance benefit data is constructed.

[0093] In this embodiment, the construction process of the hierarchical model includes: 1) Constructing the target layer: clarifying the final evaluation goal, i.e., recommending the optimal configuration scheme. 2) Criterion layer: determining the elements of the criterion layer based on the key factors of configuration recommendation. This includes cost data (including resource usage time and power consumption data), performance benefit data (such as model performance), etc. If the task has special requirements, other criteria can be added, such as scalability, stability, etc. Among them, model performance can be further subdivided into sub-criteria such as accuracy, recall, F1 score, etc. Resource usage time includes sub-criteria such as GPU usage time, and can also include memory usage, CPU utilization. Cost data covers sub-criteria such as power consumption cost / power consumption data, and can also include hardware procurement cost, maintenance cost, etc. 3) Scheme layer: listing all configuration schemes to be evaluated. These configuration schemes can be combined through different model parameter settings, hardware resource configurations, etc. For example, configuration scheme 1 has LoRA rank = 32, alpha = 64, bit width = 4bit, and uses GPU model A; configuration scheme 2 has LoRA rank = 48, alpha = 96, bit width = 8bit, and uses GPU model B, etc.

[0094] Based on the hierarchical model, each of the second fine-tuning configuration schemes is comprehensively evaluated and ranked to obtain the corresponding evaluation and ranking results.

[0095] In this embodiment, the process of comprehensively evaluating and ranking the second fine-tuning configuration scheme includes: 1. Constructing a judgment matrix. 1) Judgment matrix of the criterion layer to the target layer. Organize an expert team or relevant business personnel to compare the importance of each element of the criterion layer relative to the target layer pairwise. Use the 1-9 scale, where 1 indicates that the two elements are equally important, 3 indicates that one element is slightly more important than the other, 5 indicates that it is significantly important, 7 indicates that it is strongly important, 9 indicates that it is extremely important, and 2, 4, 6, and 8 indicate intermediate values ​​between the above adjacent judgments. 2) Judgment matrix of the sub-criterion layer to the criterion layer. For each sub-criterion under each criterion, pairwise comparisons are also performed to construct a judgment matrix. 3) Judgment matrix of the scheme layer to the sub-criterion layer. For each sub-criterion, pairwise comparisons are performed on each configuration scheme in the scheme layer to construct a judgment matrix.

[0096] 2. Hierarchical Single Sorting and Consistency Check. 1) Hierarchical Single Sorting. For each judgment matrix, calculate its maximum eigenvalue and corresponding eigenvector. After normalization, the eigenvector represents the relative importance weight of each element at that level relative to an element at the previous level. Taking the judgment matrix of the criterion layer to the target layer as an example, the eigenvector is calculated using mathematical methods (such as the eigenvalue method), and after normalization, the weights of model performance, resource usage time, and power consumption data are v1, v2, and v3, respectively. 2) Consistency Check Calculation. Consistency index CI = (λmax) / (λmax) n) / n 1. Where λmax is the largest eigenvalue of the judgment matrix, and n is the order of the judgment matrix. Find the average random consistency index RI; RI has a corresponding value for different n. For example, when n=3, RI=0.58. Calculate the consistency ratio CR=RI / CI. If CR<0.1, the judgment matrix is ​​considered to have satisfactory consistency; otherwise, the judgment matrix needs to be readjusted until the consistency requirement is met.

[0097] 3. Overall Hierarchical Ranking and Consistency Check. 1) Overall Hierarchical Ranking. Based on the results of the hierarchical single ranking, calculate the total weight of each element in the scheme layer relative to the target layer. The total weight can be obtained by multiplying the weight of each sub-criterion by the weight of the corresponding criterion, and then summing the weights of all sub-criterions using a weighted sum. For example, suppose the weight of the model performance criterion is w1, the weight of the accuracy sub-criterion under the model performance criterion is w11, the weight of the recall sub-criterion is w12, and the weight of the F1 score sub-criterion is w13; the weight of resource usage time is w2, and the weights of its sub-criterions are w21, w22, and w23; the weight of electricity consumption data is w3, and the weights of its sub-criterions are w31, w32, and w33. For configuration scheme i, its total weight relative to the target layer is Wi = w1(w11wi11+w12wi12+w13wi13)+w2(w21wi21+w22wi22+w23wi23)+w3(w31wi31+w32wi32+w33wi33), where wijk represents the weight of configuration scheme i under sub-criteria j. 2) Consistency check. Calculate the consistency ratio CR_total of the overall hierarchical ranking. If CR_total < 0.1, the overall hierarchical ranking is considered to have satisfactory consistency; otherwise, the judgment matrix needs to be readjusted. The same hierarchical single ranking and consistency check are also performed on the judgment matrix of the sub-criteria layer to the criteria layer and the judgment matrix of the scheme layer to the sub-criteria layer.

[0098] 4. Sort by total weight. Sort the total weights of each configuration scheme in the scheme layer from high to low to obtain the corresponding evaluation ranking results.

[0099] The seventh configuration scheme, which ranks first, is extracted from the evaluation and ranking results.

[0100] In this embodiment, the seventh configuration scheme, which ranks first in the ranking, can be extracted from the obtained evaluation ranking results and used as the final specified fine-tuning configuration scheme.

[0101] The seventh configuration scheme is used as the designated fine-tuning configuration scheme.

[0102] This application utilizes a pre-defined configuration recommendation engine; then, based on the engine, it constructs a hierarchical model corresponding to cost and performance benefit data; subsequently, it comprehensively evaluates and ranks each second fine-tuning configuration scheme based on the hierarchical model, obtaining the corresponding evaluation and ranking results; finally, it extracts the top-ranked seventh configuration scheme from the ranking results and uses this seventh scheme as the designated fine-tuning configuration scheme. Based on this process, this application, by using a configuration recommendation engine and constructing a multi-level hierarchical model to comprehensively evaluate and rank different configuration schemes, effectively improves the accuracy and intelligence of recommending the optimal configuration scheme for the specified fine-tuning configuration.

[0103] In some optional implementations of this embodiment, the system also has visualization and report generation functions, including: 1) Experiment management dashboard display: The experiment management dashboard provides real-time monitoring and result comparison analysis functions for fine-tuning experiments. Through charts, reports, and other forms, it intuitively displays information such as the performance change trend, resource consumption, and cost-effectiveness ratio of large language models under different configuration schemes. For example, it plots a curve showing the change in accuracy with training rounds, demonstrating the convergence speed and final performance of the model under different configurations; it creates a resource consumption bar chart to compare GPU usage time and power consumption under different configurations.

[0104] 2) Generate Configuration Recommendations and Cost-Benefit Reports: Based on the experimental data analysis results, generate configuration recommendations and cost-benefit reports. The reports include optimal configuration recommendations, performance improvement explanations, resource savings, and potential risk warnings. For example, a detailed explanation of the recommended optimal configuration and its expected performance improvement in financial tasks, a comparison of resource consumption and costs under different configurations, an analysis of potential risks, and corresponding recommendations.

[0105] In some alternative implementations, the user information obtained is subject to user consent and complies with relevant laws and policies.

[0106] Furthermore, any software tools or components not belonging to our company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

[0107] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0108] It should be emphasized that, in order to further ensure the privacy and security of the above-mentioned target fine-tuning configuration scheme, the above-mentioned target fine-tuning configuration scheme can also be stored in a node of a blockchain.

[0109] The blockchain referred to in this application is a novel application model of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Essentially, a blockchain is a decentralized database, a chain of data blocks linked together using cryptographic methods. Each data block contains information about a batch of network transactions, used to verify the validity of the information (anti-counterfeiting) and generate the next block. A blockchain can include an underlying blockchain platform, a platform product service layer, and an application service layer.

[0110] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results. Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0111] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0112] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0113] Further reference Figure 3 As a response to the above Figure 2 The implementation of the method shown in this application provides an embodiment of an artificial intelligence-based task processing device, which is similar to... Figure 2 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0114] like Figure 3 As shown, the AI-based task processing device 300 described in this embodiment includes: an acquisition module 301, an analysis module 302, a fine-tuning module 303, a generation module 304, a filtering module 305, an adjustment module 306, a screening module 307, and a processing module 308. Wherein: The acquisition module 301 is used to acquire the task type of the target task and acquire historical fine-tuning data corresponding to the task type; Analysis module 302 is used to analyze the task type and historical fine-tuning data based on a preset task sensor to determine the initial fine-tuning configuration scheme corresponding to the fine-tuning algorithm. The fine-tuning module 303 is used to fine-tune the preset large language model based on the initial fine-tuning configuration scheme and collect the performance data of the large language model. The generation module 304 is used to adjust the exploration strategy of the preset multi-armed gambling machine optimizer based on the performance data, so as to generate multiple candidate fine-tuning configuration schemes. The filtering module 305 is used to filter the candidate fine-tuning configuration schemes based on a preset accuracy risk control strategy to obtain the corresponding first fine-tuning configuration scheme. The adjustment module 306 is used to obtain the current resource status information and adjust the first fine-tuning configuration scheme based on the resource status information to obtain the corresponding second fine-tuning configuration scheme. The filtering module 307 is used to collect cost data and performance benefit data corresponding to the second fine-tuning configuration scheme, and to perform scheme filtering processing on the second fine-tuning configuration scheme based on the cost data and the performance benefit data to obtain the corresponding target fine-tuning configuration scheme. The processing module 308 is used to fine-tune the large language model based on the target fine-tuning configuration scheme to obtain the corresponding target large language model, and to process the task data to be processed based on the target large language model.

[0115] In some optional implementations of this embodiment, the generation module 304 includes: A submodule is constructed to build a mapping model between hyperparameters and task performance based on the performance data, using a preset Bayesian optimization engine. The search submodule is used to control the multi-armed gambling machine optimizer to adjust the exploration strategy corresponding to the hyperparameter configuration based on the preset iteration strategy and the mapping model, and to perform hyperparameter search based on the adjusted exploration strategy to obtain multiple corresponding first configuration schemes. The first filtering submodule is used to filter the first configuration scheme based on preset evaluation requirements to obtain the filtered second configuration scheme. The first determining submodule is used to select the second configuration scheme as the candidate fine-tuning configuration scheme.

[0116] In some optional implementations of this embodiment, the filtering module 305 includes: The first collection submodule is used to collect key indicator data corresponding to the candidate fine-tuning configuration scheme; The second filtering submodule is used to filter out a third configuration scheme from the candidate fine-tuning configuration schemes where the key indicator data is greater than a preset indicator threshold. The filtering submodule is used to filter the third configuration scheme based on a preset risk pattern library to obtain a filtered fourth configuration scheme. The second determining submodule is used to use the fourth configuration scheme as the first fine-tuning configuration scheme.

[0117] In some optional implementations of this embodiment, the adjustment module 306 includes: The monitoring submodule is used to monitor the current resource status information in real time based on the preset hardware sensing configurator; The acquisition submodule is used to acquire the dynamic adjustment strategy corresponding to the preset elastic training controller; The adjustment submodule is used to adjust the parameters of the first fine-tuning configuration scheme based on the resource status information and the dynamic adjustment strategy to obtain the adjusted fifth configuration scheme. The third determining submodule is used to use the fifth configuration scheme as the second fine-tuning configuration scheme.

[0118] In some optional implementations of this embodiment, the filtering module 307 includes: The second collection submodule is used to collect cost data corresponding to the first fine-tuning configuration scheme based on a preset monitoring tool; The third collection submodule is used to collect performance gain data corresponding to the first fine-tuning configuration scheme based on a preset model evaluation tool. The third filtering submodule is used to evaluate the cost data and the performance benefit data based on a preset evaluation strategy, and to filter out the specified fine-tuning configuration scheme that meets the requirements from the first fine-tuning configuration scheme according to the evaluation results. The fourth determination submodule is used to use the specified fine-tuning configuration scheme as the target fine-tuning configuration scheme.

[0119] In some optional implementations of this embodiment, the cost data includes resource usage time and power consumption data; the third filtering submodule includes: The first calling unit is used to call the preset cost-benefit calculation formula; The calculation unit is used to calculate and process the resource usage time, the power consumption data and the performance benefit data based on the cost-benefit calculation formula to obtain the cost-benefit score corresponding to each of the second fine-tuning configuration schemes. A filtering unit is used to select the sixth configuration scheme with the highest cost-effectiveness score from all the second fine-tuning configuration schemes; The first determining unit is used to select the sixth configuration scheme as the designated fine-tuning configuration scheme.

[0120] In some optional implementations of this embodiment, the third filtering submodule includes: A construction unit is used to construct a hierarchical model corresponding to the cost data and the performance benefit data based on the configuration recommendation engine. The processing unit is used to perform comprehensive evaluation and sorting of each of the second fine-tuning configuration schemes based on the hierarchical structure model, and obtain the corresponding evaluation and sorting results. An extraction unit is used to extract the seventh configuration scheme that ranks first from the evaluation and ranking results; The second determining unit is used to select the seventh configuration scheme as the designated fine-tuning configuration scheme.

[0121] To address the aforementioned technical problems, embodiments of this application also provide a computer device. Please refer to [link / reference needed]. Figure 4 , Figure 4This is a basic structural block diagram of the computer device in this embodiment.

[0122] The computer device 4 includes a memory 41, a processor 42, and a network interface 43 that are interconnected via a system bus. It should be noted that only the computer device 4 with components 41-43 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described here is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0123] The computer device can be a desktop computer, laptop, handheld computer, or cloud server, etc. The computer device can interact with the user via a keyboard, mouse, remote control, touchpad, or voice control.

[0124] The memory 41 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the memory 41 may be an internal storage unit of the computer device 4, such as the hard disk or memory of the computer device 4. In other embodiments, the memory 41 may also be an external storage device of the computer device 4, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 4. Of course, the memory 41 may include both the internal storage unit and its external storage device of the computer device 4. In this embodiment, the memory 41 is typically used to store the operating system and various application software installed on the computer device 4, such as computer-readable instructions for task processing methods based on artificial intelligence. In addition, the memory 41 can also be used to temporarily store various types of data that have been output or will be output.

[0125] In some embodiments, the processor 42 may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 42 is typically used to control the overall operation of the computer device 4. In this embodiment, the processor 42 is used to execute computer-readable instructions stored in the memory 41 or to process data, for example, to execute computer-readable instructions of the artificial intelligence-based task processing method.

[0126] The network interface 43 may include a wireless network interface or a wired network interface, which is typically used to establish communication connections between the computer device 4 and other electronic devices.

[0127] This application also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the artificial intelligence-based task processing method described above.

[0128] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0129] Obviously, the embodiments described above are only some embodiments of this application, not all embodiments. The accompanying drawings show preferred embodiments of this application, but do not limit the patent scope of this application. This application can be implemented in many different forms; rather, the purpose of providing these embodiments is to provide a more thorough and comprehensive understanding of the disclosure of this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this application's specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the scope of patent protection of this application.

Claims

1. A task processing method based on artificial intelligence, characterized in that, Includes the following steps: Obtain the task type of the target task, and obtain the historical fine-tuning data corresponding to the task type; Based on a preset task sensor, the task type and historical fine-tuning data are analyzed to determine the initial fine-tuning configuration scheme corresponding to the fine-tuning algorithm. Based on the initial fine-tuning configuration scheme, the preset large language model is fine-tuned, and the performance data of the large language model is collected. Based on the performance data, the exploration strategy of the preset multi-armed gambling machine optimizer is adjusted to generate multiple corresponding candidate fine-tuning configuration schemes; The candidate fine-tuning configuration schemes are filtered based on a preset accuracy risk control strategy to obtain the corresponding first fine-tuning configuration scheme. Obtain the current resource status information, and adjust the first fine-tuning configuration scheme based on the resource status information to obtain the corresponding second fine-tuning configuration scheme; Collect cost data and performance benefit data corresponding to the second fine-tuning configuration scheme, and perform scheme screening processing on the second fine-tuning configuration scheme based on the cost data and the performance benefit data to obtain the corresponding target fine-tuning configuration scheme; The target large language model is fine-tuned based on the target fine-tuning configuration scheme to obtain the corresponding target large language model, and the task data to be processed is processed based on the target large language model.

2. The task processing method based on artificial intelligence according to claim 1, characterized in that, The step of adjusting the preset exploration strategy of the multi-armed gambling machine optimizer based on the performance data to generate multiple candidate fine-tuning configuration schemes specifically includes: Based on a preset Bayesian optimization engine, a mapping model between hyperparameters and task performance is constructed according to the performance data. Based on a preset iterative strategy, the mapping model is used to control the multi-armed gambling machine optimizer to adjust the exploration strategy corresponding to the hyperparameter configuration, and the hyperparameter search is performed based on the adjusted exploration strategy to obtain multiple corresponding first configuration schemes. The first configuration scheme is filtered based on the preset evaluation requirements to obtain the filtered second configuration scheme. The second configuration scheme is used as the candidate fine-tuning configuration scheme.

3. The task processing method based on artificial intelligence according to claim 1, characterized in that, The step of filtering the candidate fine-tuning configuration schemes based on a preset accuracy risk control strategy to obtain the corresponding first fine-tuning configuration scheme specifically includes: Collect key indicator data corresponding to the candidate fine-tuning configuration schemes; A third configuration scheme is selected from the candidate fine-tuning configuration schemes, where the key indicator data is greater than the preset indicator threshold. The third configuration scheme is filtered based on a preset risk pattern library to obtain a filtered fourth configuration scheme. The fourth configuration scheme is used as the first fine-tuning configuration scheme.

4. The task processing method based on artificial intelligence according to claim 1, characterized in that, The step of obtaining the current resource status information and adjusting the first fine-tuning configuration scheme based on the resource status information to obtain the corresponding second fine-tuning configuration scheme specifically includes: Based on a preset hardware awareness configurator, the current resource status information is monitored in real time. Obtain the dynamic adjustment strategy corresponding to the preset elastic training controller; Based on the resource status information, the first fine-tuning configuration scheme is adjusted using the dynamic adjustment strategy to obtain the adjusted fifth configuration scheme. The fifth configuration scheme is used as the second fine-tuning configuration scheme.

5. The task processing method based on artificial intelligence according to claim 1, characterized in that, The step of collecting cost data and performance benefit data corresponding to the second fine-tuning configuration scheme, and performing scheme screening processing on the second fine-tuning configuration scheme based on the cost data and the performance benefit data to obtain the corresponding target fine-tuning configuration scheme, specifically includes: Cost data corresponding to the first fine-tuning configuration scheme is collected based on preset monitoring tools; Based on a preset model evaluation tool, performance gain data corresponding to the first fine-tuning configuration scheme is collected; The cost data and performance benefit data are evaluated based on a preset evaluation strategy, and a specified fine-tuning configuration scheme that meets the requirements is selected from the first fine-tuning configuration scheme based on the evaluation results. The specified fine-tuning configuration scheme is used as the target fine-tuning configuration scheme.

6. The task processing method based on artificial intelligence according to claim 5, characterized in that, The cost data includes resource usage time and power consumption data; the step of evaluating the cost data and performance benefit data based on a preset evaluation strategy, and selecting a specified fine-tuning configuration scheme that meets the requirements from the first fine-tuning configuration scheme based on the obtained evaluation results, specifically includes: Call the preset cost-benefit calculation formula; Based on the cost-benefit calculation formula, the resource usage time, the power consumption data, and the performance benefit data are calculated and processed to obtain the cost-benefit score corresponding to each of the second fine-tuning configuration schemes. Select the sixth configuration option with the highest cost-effectiveness score from all the second fine-tuning configuration options; The sixth configuration scheme is used as the designated fine-tuning configuration scheme.

7. The task processing method based on artificial intelligence according to claim 5, characterized in that, The step of evaluating the cost data and performance benefit data based on a preset evaluation strategy, and selecting a specified fine-tuning configuration scheme that meets the requirements from the first fine-tuning configuration schemes based on the obtained evaluation results, specifically includes: Call the preset configuration recommendation engine; Based on the configuration recommendation engine, a hierarchical model corresponding to the cost data and the performance benefit data is constructed. Based on the hierarchical structure model, each of the second fine-tuning configuration schemes is comprehensively evaluated and ranked to obtain the corresponding evaluation and ranking results. Extract the seventh configuration scheme that ranks first from the evaluation and ranking results; The seventh configuration scheme is used as the designated fine-tuning configuration scheme.

8. A task processing device based on artificial intelligence, characterized in that, include: The acquisition module is used to acquire the task type of the target task and acquire the historical fine-tuning data corresponding to the task type; The analysis module is used to analyze the task type and historical fine-tuning data based on a preset task sensor to determine the initial fine-tuning configuration scheme corresponding to the fine-tuning algorithm. The fine-tuning module is used to fine-tune the preset large language model based on the initial fine-tuning configuration scheme and collect the performance data of the large language model. The generation module is used to adjust the exploration strategy of the preset multi-armed gambling machine optimizer based on the performance data, so as to generate multiple candidate fine-tuning configuration schemes. The filtering module is used to filter the candidate fine-tuning configuration schemes based on a preset accuracy risk control strategy to obtain the corresponding first fine-tuning configuration scheme. The adjustment module is used to obtain the current resource status information and adjust the first fine-tuning configuration scheme based on the resource status information to obtain the corresponding second fine-tuning configuration scheme. The filtering module is used to collect cost data and performance benefit data corresponding to the second fine-tuning configuration scheme, and to perform scheme filtering processing on the second fine-tuning configuration scheme based on the cost data and the performance benefit data to obtain the corresponding target fine-tuning configuration scheme. The processing module is used to fine-tune the large language model based on the target fine-tuning configuration scheme to obtain the corresponding target large language model, and to process the task data to be processed based on the target large language model.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the task processing method based on artificial intelligence as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the artificial intelligence-based task processing method as described in any one of claims 1 to 7.