Multi-scene multi-base large model engine system

Through the multi-task instruction fine-tuning platform, large-model automation evaluation platform and large-model management platform of the multi-scene multi-base large-model engine system, the shortcomings of traditional single large-model technology in dealing with different types of tasks and quickly switching scenarios are solved, and the dynamic adaptation, efficient evaluation and precise deployment of the model are achieved to meet the complex needs of the financial industry.

CN120011187AActive Publication Date: 2025-05-16BEIJING ZHONGKE JINCAI TECH

Patent Information

Application Number
CN202411904885.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-16
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Traditional single large-model technology is difficult to handle different types of tasks, especially when switching quickly between different business scenarios. It lacks dynamic fine-tuning and evolution capabilities, and it is difficult to meet the financial industry's demand for intelligent and precise services, and it is difficult to efficiently evaluate and deploy the most suitable models.

Method used

A multi-scene multi-base large-model engine system is proposed, including the front-end interaction layer, business function layer, core technology layer and underlying support layer. The core technology layer includes a multi-task instruction fine-tuning platform, a large-model automation evaluation platform and a large-model management platform. Through these platforms, the dynamic adaptation, efficient evaluation and precise deployment of models are achieved.

Benefits of technology

The dynamic adaptation, efficient evaluation and precise deployment of the model are realized, which can meet the complex and diversified business needs of the financial industry, improve the optimization effect and deployment efficiency of the model in actual applications, and enhance the competitiveness of the banking industry in intelligent and precise services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011187A_ABST
    Figure CN120011187A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of artificial intelligence model management, and discloses a multi-scene multi-base large model engine system. Comprising a front-end interaction layer, a service function layer, a core technology layer and a bottom supporting layer, the front-end interaction layer is used for providing an access mode, an interface protocol, authority management and log recording; the business function layer is used for meeting core business requirements of multiple scenes in the financial industry; the core technology layer comprises a multi-task instruction fine tuning platform, a large model automatic evaluation platform and a large model management platform; according to the method, the optimal model is dynamically matched, the task is efficiently executed, and through the multi-task instruction fine adjustment platform and the automatic evaluation platform, the model performance can be continuously optimized, the efficient utilization of model resources is realized, and meanwhile, the accuracy and response efficiency of service processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence model management technology, and more specifically, to a multi-scene and multi-base large model engine system. Background Art

[0002] Currently, the financial industry, such as banking and insurance, has achieved remarkable results in the application of large model technology. However, with the continuous expansion of business scale and the increasing complexity of scenarios, traditional single large model technology has exposed many shortcomings in meeting the diversified needs of the financial industry.

[0003] These deficiencies are mainly manifested in the following aspects:

[0004] First, it is difficult to handle different types of tasks, especially when switching quickly between different business scenarios;

[0005] Second, it lacks the ability to dynamically fine-tune and evolve, making it difficult to meet the banking industry's growing demand for intelligent and precise services and to adapt to rapidly changing market demands;

[0006] Third, it is difficult to efficiently evaluate and quickly find the most suitable model from multiple models. This is mainly reflected in two aspects: first, the evaluation coverage is limited and cannot fully reflect the actual performance of the model on complex tasks; second, the lack of dynamic adaptation capabilities to different task characteristics makes it difficult to accurately evaluate the comprehensive performance of multi-task models, which restricts the optimization effect and deployment efficiency of the model in practical applications.

[0007] In view of this, the present invention proposes a multi-scene multi-base large model engine system to solve the above problems. Summary of the invention

[0008] In order to overcome the above defects of the prior art and to achieve the above objectives, the present invention provides the following technical solutions: a multi-scenario multi-base large model engine system, including a front-end interaction layer, a business function layer, a core technology layer and an underlying support layer;

[0009] The front-end interaction layer is used to provide access methods, interface protocols, authority management and log records;

[0010] The business function layer is used to meet the core business needs of multiple scenarios in the financial industry;

[0011] The core technology layer includes a multi-task instruction fine-tuning platform, a large model automatic evaluation platform, and a large model management platform;

[0012] The underlying support layer includes data and storage, security, computing and reasoning support, and DevOps and operation and maintenance.

[0013] Furthermore, the access method supports Web terminals, mobile terminals, third-party application APIs, multi-dimensional terminal interactions, and business system access;

[0014] The interface protocol interacts via HTTP, WebSocket, REST, JSON format, and HTML protocol;

[0015] The authority management and log recording are used to perform multi-level authority management on access users and systems, and record access logs at the same time.

[0016] Furthermore, the business function layer includes an intelligent customer service module, an intelligent knowledge base question and answer module, an intelligent report generation module, an intelligent contract review module, a financial product investment education module, an intelligent question and answer module, and an intelligent marketing module.

[0017] Furthermore, the multi-task instruction fine-tuning platform is used to identify real-time task requirements and dynamically load the corresponding LoRA fine-tuning module:

[0018] Among them, the method of identifying real-time task requirements is:

[0019] After receiving the task input, the system parses the input content through the front-end interface and extracts the input data;

[0020] Extract task features from input data through a natural language processing model and map the extracted task features into high-dimensional feature vectors;

[0021] Assign tasks to specific scenario tags based on task routing rules and analyze the fine-grained characteristics of task requirements;

[0022] Parse the multimodal data contained in the input to determine the type of data involved in the task;

[0023] Combined with the task scenario database in the system, the task characteristics are matched with the predefined scenario characteristics to obtain the real-time task requirements;

[0024] In the LoRA fine-tuning module, multiple LoRA fine-tuning modules are pre-trained and registered, each LoRA fine-tuning module focuses on several scenario requirements;

[0025] Register task feature labels for each LoRA fine-tuning module, and match real-time task requirements with module registration features through feature extraction;

[0026] The dynamic loading includes: when the task matches a single LoRA fine-tuning module, dynamically loading the LoRA fine-tuning module;

[0027] When the task characteristics cover multiple LoRA fine-tuning modules, the module combination mechanism is adopted, and multiple LoRA fine-tuning modules are loaded to process the task.

[0028] Furthermore, when the multi-task instruction fine-tuning platform is trained, multiple task modules are integrated into a unified model base through LoRA technology;

[0029] By adopting a multi-expert model framework, the corresponding LoRA fine-tuning modules are dynamically activated according to the input task characteristics; through the selection mechanism, only the LoRA fine-tuning modules required for the current task are loaded.

[0030] Furthermore, when the multi-task instruction fine-tuning platform is training, it also includes an early warning mechanism, and the early warning mechanism performs an early warning in the following manner:

[0031] Collect multiple key data related to resources in real time, including computing resources, memory and storage, network and DNS health checks;

[0032] After collecting key data, the resource bottleneck is identified through anomaly detection algorithms, and the identified resource bottleneck is notified to the user in real time;

[0033] Among them, if an abnormal node appears during the training process, the abnormal node will be automatically isolated and detected;

[0034] The nodes are identified through anomaly detection algorithms, and nodes with abnormal operation are recorded as abnormal nodes, and vice versa, they are recorded as healthy nodes;

[0035] After automatically isolating the abnormal node, continue to monitor and take measures to restore the node;

[0036] Record abnormal logs and isolation processes in the database, and use historical data to optimize abnormal detection algorithms and automatic isolation;

[0037] The automatic isolation includes dynamic task migration, communication isolation, partial task restart and resource recovery.

[0038] Furthermore, the large model automated evaluation platform includes multi-dimensional performance evaluation, automated testing and quantitative feedback, and model version comparison and error analysis;

[0039] Among them, multi-dimensional performance evaluation is used to evaluate the performance of the model in multiple dimensions, including accuracy, response time, and resource consumption;

[0040] Automated testing and quantitative feedback are used for scenario-based automated testing functions;

[0041] Model version comparison and error analysis is used to perform horizontal comparative analysis of different model versions.

[0042] Furthermore, the large model management platform includes intelligent routing and model adaptation, AI-Agent development and operation platform, and full life cycle management;

[0043] The intelligent routing and model adaptation automatically selects and integrates the optimal model by combining the results of the large model evaluation platform;

[0044] The AI-Agent development and operation platform is based on an enhanced retrieval algorithm for knowledge texts and is used to combine contextual information with a real-time updated knowledge base;

[0045] The lifecycle management is used to cover the model training, evaluation, selection, deployment and iterative optimization process to form a closed-loop management.

[0046] Furthermore, the large-model automated evaluation platform also includes task weighting and dynamic scoring, wherein the task weighting and dynamic scoring set up a task weighting system according to scenario requirements, and the task weighting system includes static weight allocation and dynamic weight allocation.

[0047] Furthermore, the large-model automated evaluation platform also includes scenario-based and multi-task evaluations, which provide feedback and verification on the model by simulating real business environments and multi-task adaptability evaluations.

[0048] The technical effects and advantages of the multi-scene multi-base large model engine system of the present invention are as follows:

[0049] The present invention realizes dynamic adaptation, efficient evaluation and precise deployment of models through the collaborative work of a multi-task instruction fine-tuning platform, a large model automatic evaluation platform and a large model management platform, and can meet the complex and diverse business needs of the financial industry. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a structural schematic diagram of the multi-scene multi-base large model engine system of the present invention;

[0051] Figure 2 It is a schematic diagram of the architecture of the present invention. DETAILED DESCRIPTION

[0052] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0053] Example 1

[0054] See also Figure 1 and Figure 2 As shown, the multi-scenario multi-base large model engine system described in this embodiment includes a front-end interaction layer, a business function layer, a core technology layer and an underlying support layer;

[0055] The front-end interaction layer is used to provide access methods, interface protocols, authority management and log records;

[0056] The business function layer is used to meet the core business needs of multiple scenarios in the financial industry;

[0057] The core technology layer includes a multi-task instruction fine-tuning platform, a large model automatic evaluation platform, and a large model management platform;

[0058] The underlying support layer includes data and storage, security, computing and reasoning support, and DevOps and operation and maintenance.

[0059] Furthermore, the access method supports Web terminals, mobile terminals, third-party application APIs, multi-dimensional terminal interactions, and business system access;

[0060] The interface protocol interacts via HTTP, WebSocket, REST (such as RESTful architecture), JSON format, and HTML protocol;

[0061] The authority management and log recording are used to perform multi-level authority management on access users and systems, and record access logs;

[0062] Specifically, multiple access methods can meet the multi-scenario needs of the financial industry; multiple interface protocols can ensure the high compatibility and low latency performance of the system; permission management and logging can ensure the security and compliance of the system.

[0063] Furthermore, the business function layer includes an intelligent customer service module, an intelligent knowledge base question and answer module, an intelligent report generation module, an intelligent contract review module, a financial product investment education module, an intelligent question and answer module, and an intelligent marketing module;

[0064] Specifically, the intelligent customer service module supports multiple rounds of dialogue and context understanding, and is suitable for scenarios such as account inquiries and loan applications, providing customers with a continuous and smooth interactive experience and improving customer service quality; the intelligent knowledge base question-and-answer module extracts information from documents such as financial reports and customer feedback, provides accurate knowledge support for bank employees, helps quickly answer compliance questions, and improves the efficiency and accuracy of information queries; the intelligent report generation module supports the generation of due diligence, credit reports and research reports, with a professional banking style and high accuracy, provides in-depth support for decision-making, and meets the reporting needs of banks; the intelligent contract review module conducts intelligent review of contracts in banking business, automatically identifies key terms and potential risks, ensures the compliance and effectiveness of contracts, and improves review efficiency; the wealth management product investment education module: provides education and guidance functions for wealth management products, helps customers understand complex wealth management product information, and improves investors' risk awareness and decision-making ability; the intelligent question module (such as ChatBI) supports data interaction and intelligent chart generation, provides trend forecasting and data analysis, helps bank personnel quickly gain insight into business data, and supports business decision-making; the intelligent marketing module provides personalized marketing copy recommendations for banks through customer behavior data analysis, improves marketing effectiveness, and helps banks achieve accurate customer reach.

[0065] It should be noted that this system can accurately match the diverse needs of financial business scenarios and improve response efficiency and adaptability through multi-model collaboration. For example, the system not only supports efficient fine-tuning of specific scenarios, but also can evaluate and select the optimal model in real time to ensure that the model performance is highly consistent with business needs. The application of this system has significantly enhanced the competitiveness of the banking industry in intelligent and precise services, and provided customized intelligent support for it to cope with the rapidly changing market environment.

[0066] Furthermore, the multi-task instruction fine-tuning platform is used to identify real-time task requirements and dynamically load the corresponding LoRA fine-tuning module:

[0067] Among them, the method of identifying real-time task requirements is:

[0068] After the system receives the task input (such as natural language questions, structured queries, text analysis requests, etc.), it parses the input content through the front-end interface and extracts the input data. For example, in the intelligent customer service scenario, the input may be the user's natural language questions; in the financial calculation scenario, the input may be a digital data table or calculation formula;

[0069] The task features of the input data are extracted through the natural language processing (NLP) model; the main extracted features include: task type (such as question answering, analysis, generation); scenario category (such as financial report generation, risk control, intelligent customer service, etc.); complexity features (such as whether the task requires multiple rounds of interaction or high-precision calculation); the extracted task features are mapped to high-dimensional feature vectors. For example, the input sentence can be converted into a vector representation through a pre-trained model (such as BERT, GPT), with additional task scenario labels and keywords;

[0070] Assign tasks to specific scenario labels based on classification algorithms (e.g., logistic regression, multi-layer perceptron) or task routing rules (e.g., “How is the loan application process?” → intelligent customer service scenario; “Calculate the risk indicators of the current investment portfolio,” → financial analysis scenario);

[0071] Analyze the fine-grained characteristics of task requirements through rule-based engines or deep learning models (e.g., does it require real-time processing? Does it involve multi-turn dialogue? Does it require specific knowledge base retrieval support?)

[0072] Parse the multimodal data (text, tables, images, etc.) that may be contained in the input to determine the type of data involved in the task;

[0073] Combined with the task scenario database in the system, the task characteristics are matched with the predefined scenario characteristics to determine the most suitable task requirements;

[0074] The system generates dynamic feedback based on the feature matching results (for example, if a task is identified as requiring knowledge base support, the feedback is "knowledge base question and answer module needs to be loaded");

[0075] LoRA fine-tuning module includes:

[0076] During the system development phase, multiple LoRA fine-tuning modules are pre-trained and registered for different business scenarios or task types, each focusing on one or more scenario requirements (e.g., intelligent customer service module: for multi-round dialogue and context understanding; financial analysis module: for high-precision calculation and table data processing; risk review module: for compliance and risk control scenarios);

[0077] Register task feature labels for each LoRA fine-tuning module (e.g., intelligent customer service module: task type = question and answer, scenario category = customer service, multi-round dialogue = True; financial analysis module: task type = analysis, scenario category = finance, high calculation accuracy = True);

[0078] The real-time task requirements are matched with the module registration features through feature extraction. (The matching methods include: rule matching based on task labels and module labels. For example, when the task requirements contain the keywords "risk" and "audit", the risk review module is matched first; the cosine similarity or other similarity indicators are calculated by vectorizing the task features and the module feature vectors, and the LoRA fine-tuning module with the highest similarity is selected;)

[0079] Dynamic loading includes: when a task matches a single LoRA fine-tuning module, the system dynamically loads the module (for example, the user asks "How do I calculate the return on investment?" The system loads the financial analysis module);

[0080] When the task features cover multiple modules (such as the need to semantically understand user input and combine it with knowledge base extraction), the system adopts a module combination mechanism, loads multiple LoRA fine-tuning modules, and processes tasks sequentially or in parallel (for example, the intelligent customer service module + knowledge base question and answer module collaboratively processes tasks);

[0081] Moreover, after the task is executed, the system monitors the module’s output quality, response time and other related indicators, and records the compatibility between the task and the module;

[0082] Improve the matching accuracy between tasks and modules based on historical data and model performance optimization rules (for example, if a task is associated with a specific module multiple times, the system can automatically enhance the priority of the module);

[0083] The combination mechanism includes:

[0084] Extract task features based on task type (e.g., question answering, generation, calculation) and scenario requirements (e.g., financial report generation, risk analysis);

[0085] Task characteristics include: input data type: text, table, image; target output type: classification, summary, answer, etc.; task complexity: whether it involves multiple rounds of dialogue, high-precision calculation;

[0086] Decompose complex tasks into multiple subtasks (e.g., financial report generation task, subtask 1: extract key data from documents (information extraction);

[0087] Classify open source models according to their characteristics (e.g., question-answering tasks: ChatGLM, LLaMA; generation tasks: GPT-4, Bloom, etc.);

[0088] Select the most suitable model for each subtask (e.g., information extraction: T5, because it performs better on document processing and extraction tasks);

[0089] Dynamically select the most suitable model based on task characteristics, using a rule engine or deep learning classifier based on feature matching;

[0090] Set up task pipelines for multiple models to ensure that tasks are completed in a logical order (e.g., information extraction → data analysis → text generation);

[0091] For independent subtasks, parallel execution is adopted;

[0092] Unify and optimize the outputs of multiple models (e.g., fuse answers from multiple models, such as weighted voting, confidence ranking);

[0093] Set evaluation indicators and adjust the model combination method based on the evaluation results (for example, replace poorly performing models and redefine the division of tasks);

[0094] Specifically, for the seven core application scenarios in the banking industry (i.e., intelligent customer service, intelligent knowledge base Q&A, intelligent report generation, intelligent contract review, intelligent investment education for wealth management products, intelligent questions and answers, and intelligent marketing), the platform efficiently fine-tunes through the LoRA fine-tuning module to achieve high-precision support for multiple tasks and multiple scenarios; furthermore, this modular design enhances the flexibility and scalability of the system in complex tasks, enabling it to quickly respond to changes in business needs.

[0095] Furthermore, the large model automated evaluation platform includes multi-dimensional performance evaluation, automated testing and quantitative feedback, and model version comparison and error analysis;

[0096] Among them, multi-dimensional performance evaluation is used to evaluate the performance of the model in multiple dimensions such as accuracy, response time, and resource consumption;

[0097] Automated testing and quantitative feedback are used for scenario-based automated testing functions to quickly evaluate the performance of the model in different business scenarios, generate quantitative feedback for each fine-tuning module, and guide model optimization;

[0098] Model version comparison and error analysis are used to conduct horizontal comparative analysis of different model versions. Through automatic scoring, error classification and detailed reporting functions, the advantages and disadvantages of the model can be quickly identified, providing a scientific basis for continuous iterative optimization.

[0099] The performance of the rapid assessment model in different business scenarios includes:

[0100] Extract a small representative dataset from banking scenarios (usually through sampling techniques such as random sampling or stratified sampling) (for example, for intelligent customer service tasks, the sample includes common user questions and high-frequency questions);

[0101] Load predefined test tasks and run the model for prediction, using preset metrics (such as accuracy, response time) to instantly calculate model performance;

[0102] For each task, the system only tests key scenarios and indicators (for example, the intelligent customer service task only tests the fluency and semantic consistency of multi-round conversations, and the report generation task only tests the accuracy of key data extraction);

[0103] Generate a summary report based on the evaluation results, pointing out the strengths and weaknesses of the model;

[0104] The model version comparison and error analysis (horizontal comparison and analysis of different model versions) includes:

[0105] Quantify the results of each dimension (accuracy, response time, etc.) to generate a score (for example, accuracy: 95%; response time: 450 milliseconds (better than the standard 500 milliseconds);

[0106] Set weights based on task priorities and calculate the total score of the model;

[0107] An evaluation report is then automatically generated, which includes detailed data, model strengths and weaknesses, and improvement suggestions;

[0108] The guiding model optimization comprises:

[0109] The module with the lowest score in the evaluation report is considered the main bottleneck (for example, the response time of a module is higher than the standard, and the semantic consistency score in the multi-turn dialogue task is low);

[0110] Use model pruning or quantization to optimize the module (for example, for high response time issues, find the root cause of the problem through quantitative analysis, such as excessive model parameters, which can be optimized by model pruning or quantization technology; for multi-round dialogue issues, increase task-related training data or improve fine-tuning modules);

[0111] Specifically, in the intelligent customer service scenario, the quantitative feedback showed a "low semantic consistency score" (80%); based on the feedback, the system suggested improvements to the context understanding capability: increasing data fine-tuning for multiple rounds of conversations and optimizing the weight distribution of the LoRA fine-tuning module;

[0112] The automatic scoring includes:

[0113] The system records the input and output results of the model in real time when the task is executed; calculates the score (such as accuracy, response time) according to the corresponding indicator formula; and synthesizes the total score through the task weight mechanism to obtain the automatic scoring result;

[0114] The error categories include:

[0115] The system classifies the model output and marks correct and incorrect results. It also performs cluster analysis on incorrect results and attributes them to specific question types (such as semantic understanding errors and knowledge base call failures). For example, in the intelligent customer service task, the system did not correctly answer the question "loan application process" and the error was classified as "knowledge loss".

[0116] The detailed report is used to automatically generate a visual report, which includes evaluation indicator scores, error classification statistics and optimization suggestions;

[0117] Specifically, the large model automated evaluation platform can fully reflect the applicability and stability of the model in banking scenarios through multi-dimensional performance evaluation; automated testing and quantitative feedback can quickly evaluate the performance of the model in different business scenarios through scenario-based automated testing functions, generate quantitative feedback for each fine-tuning module, and guide model optimization; model version comparison and error analysis support horizontal comparative analysis of different model versions, and quickly identify the advantages and disadvantages of the model through automatic scoring, error classification and detailed reporting functions, providing a scientific basis for continuous iterative optimization;

[0118] It should be noted that the rapid evaluation model has the following advantages: through sample simplification and parallel calculation, the evaluation time is usually a few minutes, which is highly time-efficient; no human intervention is required, and the system automatically completes task allocation, execution and result analysis; the scenario coverage is limited, and the rapid evaluation only covers some key scenarios to screen models or find obvious problems.

[0119] Furthermore, the large model management platform includes intelligent routing and model adaptation, AI-Agent development and operation platform, and full life cycle management;

[0120] The intelligent routing and model adaptation automatically selects and integrates the optimal model by combining the results of the large model evaluation platform;

[0121] The AI-Agent development and operation platform is based on an enhanced retrieval algorithm for knowledge texts and is used to combine contextual information with a real-time updated knowledge base;

[0122] The lifecycle management is used to cover the model training, evaluation, selection, deployment and iterative optimization process to form a closed-loop management;

[0123] Among them, the AI-Agent development and operation platform introduces the contextual understanding capabilities of Transformer-based sequence modeling technology (such as BERT and GPT);

[0124] Support content updates of dynamic knowledge bases through incremental update technologies (such as real-time index updates or dynamic vector storage);

[0125] Convert multimodal information such as text, tables, and images into a unified vector representation and perform retrieval based on a unified semantic space (for example, in financial analysis, relevant content can be extracted from both text and charts at the same time);

[0126] Combine user behavior data and task priority, and use reinforcement learning (such as RLHF, Reinforcement Learning with Human Feedback) to optimize the ranking of search results;

[0127] Lifecycle management includes: support for the complete training process from zero-based training to fine-tuning, including PromptTuning, LoRA fine-tuning and other technologies;

[0128] Analyze model performance through multi-dimensional evaluation (such as accuracy, response time, resource consumption, etc.);

[0129] Intelligently select the optimal model based on the evaluation results, support cloud, edge and local deployment, meet the financial industry's high requirements for data security, and continuously optimize the model based on feedback during operation;

[0130] It should be noted that the multi-scenario and multi-base large model engine system has achieved all-round technological innovation from dynamic task fine-tuning, efficient evaluation to intelligent management. It not only improves the adaptability and flexibility of the model in the financial industry, but also significantly improves business processing efficiency and overall system performance.

[0131] Furthermore, in the underlying support layer, the data and storage are used to support database systems such as TDSQL and Oceanbase, are compatible with a variety of unstructured data formats (such as PDF, DOCX, JPEG, etc.), and support distributed file storage (DFS, Minio, Elasticsearch);

[0132] The security assurance system includes internal control security audit, system security, network security and privacy protection to ensure the data compliance and security of the system in financial scenarios;

[0133] The calculation and reasoning support certain AI processors, GPUs (such as A100, V100) and other high-performance computing devices to improve model training and reasoning speed;

[0134] The DevOps and operation and maintenance provide automated testing, CI / CD processes and a bug feedback platform to achieve efficient development and operation integration.

[0135] Furthermore, when the multi-task instruction fine-tuning platform is trained, multiple task modules are integrated into a unified model base through LoRA technology;

[0136] Adopting the multi-expert model (MoE) framework, the corresponding LoRA fine-tuning module is dynamically activated according to the input task characteristics. Through the weighting or selection mechanism, only the modules required for the current task are loaded;

[0137] It supports dynamic adaptability of multi-tasking, and ensures that the system can quickly respond to the diverse needs of the banking industry by activating different LoRA fine-tuning modules to handle tasks such as question answering, sentiment analysis, and financial calculations.

[0138] Furthermore, when the multi-task instruction fine-tuning platform is training, it also includes an early warning mechanism, and the early warning mechanism performs an early warning in the following manner:

[0139] When the system is running, multiple key data related to resources are collected in real time, including computing resources, memory and storage, network, DNS health check (specifically, computing resources include GPU / NPU utilization, monitoring memory utilization and computing core (ComputeCoreUtilization) occupancy; its key indicators are: GPU core utilization (such as reaching more than 90%), memory utilization (such as reaching more than 95% or memory overflow), temperature (such as exceeding the recommended temperature, usually 70℃-80℃), CPU utilization: monitoring the host CPU utilization, especially in the data preprocessing stage. Key indicators: single-core / multi-core occupancy exceeds 85%, data I / O delay. Memory and storage Storage includes RAM, detects whether there is memory leak or memory overflow, key indicators: memory usage exceeds 90%, and frequently triggered memory paging (PageSwapping). Storage (Disk): Insufficient disk capacity or read and write speed bottleneck, key indicators: available disk space is less than 10%, and IOPS (input / output operations per second) is abnormally low. Network includes network bandwidth, detects the transmission bandwidth of training data from storage to nodes, key indicators: bandwidth utilization exceeds 90%, network jitter (Jitter) and high latency (Latency). DNS health check: detects DNS resolution delay and whether there is resolution failure, key indicators: DNS resolution time is greater than 200ms, and resolution failure rate exceeds 2%);

[0140] After collecting key data, resource bottlenecks are identified through anomaly detection algorithms (the anomaly detection algorithms include any one of threshold trigger detection, dynamic baseline detection, time series analysis, multi-dimensional indicator association analysis, or pattern matching and AI detection, wherein pattern matching and AI detection can detect unknown anomalies by identifying known problem patterns and using deep learning models);

[0141] Notify users of identified resource bottlenecks in real time through dashboard display, log records, email, SMS, and instant messaging tools (such as Slack and DingTalk). The notification content includes bottleneck description, possible problems, and recommended solutions.

[0142] Among them, if abnormal nodes appear during the training process, the system will automatically isolate and detect the abnormal nodes;

[0143] The nodes are identified through anomaly detection algorithms, and nodes with abnormal operation are recorded as abnormal nodes, and vice versa, they are recorded as healthy nodes;

[0144] After automatically isolating abnormal nodes, continue to monitor the system and take measures to restore nodes or optimize system performance;

[0145] The system records the exception log and isolation process in the database;

[0146] And use historical data to optimize anomaly detection algorithms and automatic isolation methods;

[0147] The automatic isolation includes dynamic task migration, communication isolation, partial task restart and resource recovery;

[0148] The dynamic task migration includes migrating the training tasks on the abnormal nodes to other healthy nodes, such as using a distributed training framework (such as Horovod, Ray) to support task migration (for example, if the GPU of a node fails, the tasks of the node are suspended and the task slices are allocated to the backup nodes);

[0149] The communication isolation includes stopping the communication between the abnormal node and other nodes to prevent the abnormal state from spreading, and updating the topology structure between nodes (such as the Ring or Tree structure of distributed training) (for example, if a node has a network instability problem, the system will disconnect its communication with the cluster and reconfigure the communication topology);

[0150] The partial task restart includes partially restarting the tasks on the abnormal node without affecting the global training process (for example, if the data file read by a node is damaged, the data shard of the node is reloaded and the task is restarted);

[0151] The resource recovery includes releasing the resources occupied by abnormal nodes to prevent resource waste (terminating the process that is out of control due to crash, and releasing the video memory and internal memory).

[0152] The method of restoring the node includes: detecting and restoring hardware anomalies caused by temporary problems (such as re-enabling the node after the temperature drops) and service restart (automatically restarting the crashed service process).

[0153] The optimizing system performance includes adjusting system configuration according to abnormal patterns, dynamically adjusting resource allocation strategies, and optimizing communication topology to improve fault tolerance.

[0154] Specifically, for example, when a GPU overheating problem occurs, the phenomenon is manifested as the GPU temperature of a certain node reaching 90°C; the system stops the tasks on the GPU and migrates the tasks to the backup node; reduces the GPU frequency, waits for the temperature to recover and then re-enables it, updates the cooling strategy to prevent similar problems from recurring; for example, when a network delay problem occurs, the phenomenon is manifested as a sudden increase in communication delay between nodes; detection and isolation; the system disconnects the abnormal node and reconfigures the communication topology; then checks the network equipment and adjusts the data transmission strategy.

[0155] Furthermore, the large model automated evaluation platform also includes task weighting and dynamic scoring, wherein task weighting and dynamic scoring set up a task weighting system according to scenario requirements. For example, financial calculation tasks focus on calculation accuracy, while intelligent customer service tasks focus more on language fluency and adaptability. A comprehensive score is dynamically generated to ensure the pertinence and practicality of the evaluation results. The task weighting system includes static weight allocation and dynamic weight allocation.

[0156] For static weight allocation, the weight of each evaluation indicator is defined based on the priority of the business scenario;

[0157]

[0158] Among them, Weighted Score represents the weighted score, n is the number of evaluation indicators, W i is the weight of the i-th evaluation indicator, s i is the score corresponding to the i evaluation indicators, where is equal to 1;

[0159] For dynamic weight allocation, it dynamically adjusts the weights based on real-time task characteristics;

[0160] Through a rule engine (such as RuleEngine): dynamically match weight templates based on task tags (such as "finance" and "customer service");

[0161] Through reinforcement learning: optimize the weight allocation strategy through historical task data;

[0162] Design multiple expert modules (Experts), and dynamically select the expert that best suits the task type corresponding to the current task through the gating network (GatingNetwork);

[0163] Among them, the gating network uses task features as input and outputs task weights when selecting experts;

[0164] The output of the MoE is used as the basis for scoring and combined with a weighting system to generate a comprehensive score;

[0165] The weighted formula becomes:

[0166]

[0167] Among them, Gating(w i ) is the weight dynamically adjusted by the gating network, which adjusts the original weight W of each indicator according to the input context information (such as the characteristics of the current task, historical data, environmental changes, etc.) i Make dynamic adjustments;

[0168] Extract features from task input (such as text, speech, or structured data) to obtain task type (question-answering, calculation, generation), data complexity (text length, semantic depth), and business scenario (intelligent customer service, financial calculation, risk review);

[0169] According to the extracted features, appropriate weights are selected from predefined weight templates;

[0170] Evaluate the tasks and calculate the individual scores of each indicator;

[0171] Specifically, for example, task type: user request to generate due diligence report; task characteristics: task label: report generation; data complexity: high (need to process tables and multiple paragraphs of text). Weight distribution: accuracy = 0.6, response time = 0.3, resources = 0.1; score calculation: accuracy score = 90%, response time score = 85%, resource score = 80%; the comprehensive score is substituted into the above formula (the changed weighted formula) to get 87.5;

[0172] It should be noted that by utilizing the expert allocation and gated network mechanism of MoE, more refined dynamic weighting can be achieved; MoE is suitable for scenarios with diverse task types and complex features; when the task complexity is high, combining the MoE framework can further improve the dynamic adaptation capability and scoring accuracy, especially for optimization scenarios with diversified task requirements in the financial field.

[0173] Furthermore, the large model automated evaluation platform also includes scenario-based and multi-task evaluation, which provides refined feedback for the model by simulating real business environments (such as smart investment advisors, compliance reviews, etc.); the multi-task adaptability evaluation module verifies the stability and output quality of the model during task switching;

[0174] Wherein, the method of simulating the real business environment is:

[0175] Collect and analyze common tasks and user needs in the target business, and define possible inputs and outputs in the scenario (for example, in a smart investment advisory scenario, users input investment goals and risk preferences, and output recommended investment portfolios; input contract or policy texts, and output compliance analysis results);

[0176] Define specific task types and their goals for each business scenario (e.g., information extraction task: extract key terms from a document, question answering task: answer user questions based on a knowledge base);

[0177] Collect historical data from real business and perform data preprocessing (such as customer questions, contract texts, market data, etc., data preprocessing includes data cleaning and data labeling);

[0178] Use generative models (such as GPT) to create simulated input data to simulate diverse user interactions;

[0179] Generate test tasks randomly or according to rules based on scenario task definitions (for example, simulate users asking questions such as “How to configure a low-risk investment portfolio?” or inputting contract text such as “Do the terms stipulated by Party A in the contract comply with tax policies?”);

[0180] Input simulation tasks into business-related subsystems to ensure that tasks are consistent with actual processes (for example, in the smart investment advisory scenario, call real-time market data, and in the contract review scenario, use a dynamic knowledge base);

[0181] Call the target model to complete the task and record the output results;

[0182] Define evaluation indicators specific to each task (for example, in the case of smart investment advisors: the risk-return ratio of the recommended portfolio and the timeliness of the recommendation; in the case of compliance review: the accuracy and recall rate of clause detection);

[0183] Compare the task output with the expected results and calculate various indicators;

[0184] Types of errors in the output of labeling and classification tasks (e.g., logical errors, knowledge gaps);

[0185] The task switching method is as follows:

[0186] The system parses the input content and extracts the task features;

[0187] Based on task characteristics, classify input into different task types;

[0188] The system encapsulates the processing logic of each task into an independent module (such as the LoRA fine-tuning module);

[0189] Dynamically load the corresponding modules according to the task classification results;

[0190] Specifically, the task description is: the user continuously asks different types of questions: (Question 1: "What is the current loan interest rate?" → data query task; Question 2: "How to calculate the monthly loan payment?" → financial calculation task; Question 3: "Are there any low-risk investment products recommended?" → smart investment advisory task);

[0191] The task switching process is as follows: Input parsing: (Question 1: Identify the keyword "loan interest rate", classified as a data query task; Question 2: Identify the keyword "monthly payment", classified as a financial calculation task; Question 3: Identify the keyword "low-risk investment", classified as a smart investment advisory task);

[0192] The modules are loaded as follows: (data query task: load the knowledge base question and answer module to extract answers from real-time interest rate data; financial calculation task: load the financial calculation module to calculate monthly payments using formulas; smart investment advisory task: load the recommendation module to call the investment analysis model to generate product recommendations);

[0193] The output results are: (the answer to question 1: "The current loan interest rate is 4.5%;" the answer to question 2: "Based on the input amount and term, the monthly payment is ¥5000;" the answer to question 3: "Based on your risk preference, it is recommended to invest in fixed-income products).

[0194] In this embodiment, by matching and loading the LoRA fine-tuning module, the appropriate LoRA module can be accurately selected according to the characteristics of the task, thereby improving the task processing efficiency. Moreover, if the task characteristics cover multiple modules, a module combination mechanism can be adopted to improve the adaptability and flexibility of the system. Through a unified model base, the overhead of model switching is reduced, and the efficiency and accuracy of task processing are improved by activating the corresponding expert module. Only the required LoRA module is loaded according to the task requirements, thereby avoiding the loading of irrelevant modules and optimizing the use of computing resources. Through real-time resource monitoring and anomaly detection, problems can be discovered in advance and measures can be taken to prevent system crashes or performance degradation. By automatically isolating abnormal nodes and recycling resources, the need for manual intervention can be greatly reduced, and the autonomy and reliability of the system can be improved. Multi-dimensional performance evaluation can ensure that the model can maintain high performance in different scenarios, avoiding overfitting or underfitting. Multi-task evaluation can ensure that the system can perform well in a variety of task scenarios and improve task adaptability. Through intelligent routing and model adaptation, resources can be efficiently configured, invalid calculations can be reduced, and life cycle management can ensure that the model can be effectively managed at each stage of the life cycle, thereby improving the continuous value of the model.

[0195] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0196] In the several embodiments provided by the present invention, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only one, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0197] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

[0198] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. Multi-scene multi-base large model engine system, characterized by: Including front-end interaction layer, business function layer, core technology layer and underlying support layer; The front-end interaction layer is used to provide access methods, interface protocols, authority management and log records; The business function layer is used to meet the core business needs of multiple scenarios in the financial industry; The core technology layer includes a multi-task instruction fine-tuning platform, a large model automatic evaluation platform, and a large model management platform; The underlying support layer includes data and storage, security, computing and reasoning support, and DevOps and operation and maintenance.

2. The multi-scene multi-base large model engine system according to claim 1 is characterized in that: The access method supports Web, mobile, third-party application API, multi-dimensional terminal interaction and business system access; The interface protocol interacts via HTTP, WebSocket, REST, JSON format, and HTML protocol; The authority management and log recording are used to perform multi-level authority management on access users and systems, and record access logs at the same time.

3. The multi-scene multi-base large model engine system according to claim 1 is characterized in that: The business function layer includes an intelligent customer service module, an intelligent knowledge base question and answer module, an intelligent report generation module, an intelligent contract review module, a financial product investment education module, an intelligent question and answer module, and an intelligent marketing module.

4. The multi-scene multi-base large model engine system according to claim 1, characterized in that: The multi-task instruction fine-tuning platform is used to identify real-time task requirements and dynamically load the corresponding LoRA fine-tuning module: Among them, the method of identifying real-time task requirements is: After receiving the task input, the system parses the input content through the front-end interface and extracts the input data; Extract task features from input data through a natural language processing model and map the extracted task features into high-dimensional feature vectors; Assign tasks to specific scenario tags based on task routing rules and analyze the fine-grained characteristics of task requirements; Parse the multimodal data contained in the input to determine the type of data involved in the task; Combined with the task scenario database in the system, the task characteristics are matched with the predefined scenario characteristics to obtain the real-time task requirements; In the LoRA fine-tuning module, multiple LoRA fine-tuning modules are pre-trained and registered, each LoRA fine-tuning module focuses on several scenario requirements; Register task feature labels for each LoRA fine-tuning module, and match real-time task requirements with module registration features through feature extraction; The dynamic loading includes: when the task matches a single LoRA fine-tuning module, dynamically loading the LoRA fine-tuning module; When the task characteristics cover multiple LoRA fine-tuning modules, the module combination mechanism is adopted, and multiple LoRA fine-tuning modules are loaded to process the task.

5. The multi-scene multi-base large model engine system according to claim 1 is characterized in that: When the multi-task instruction fine-tuning platform is trained, multiple task modules are integrated into a unified model base through LoRA technology; By adopting a multi-expert model framework, the corresponding LoRA fine-tuning modules are dynamically activated according to the input task characteristics; through the selection mechanism, only the LoRA fine-tuning modules required for the current task are loaded.

6. The multi-scene multi-base large model engine system according to claim 1, characterized in that: When the multi-task instruction fine-tuning platform is training, it also includes an early warning mechanism, and the early warning mechanism performs an early warning in the following manner: Collect multiple key data related to resources in real time, including computing resources, memory and storage, network and DNS health checks; After collecting key data, the resource bottleneck is identified through anomaly detection algorithms, and the identified resource bottleneck is notified to the user in real time; Among them, if an abnormal node appears during the training process, the abnormal node will be automatically isolated and detected; The nodes are identified through anomaly detection algorithms, and nodes with abnormal operation are recorded as abnormal nodes, and vice versa, they are recorded as healthy nodes; After automatically isolating the abnormal node, continue to monitor and take measures to restore the node; Record abnormal logs and isolation processes in the database, and use historical data to optimize abnormal detection algorithms and automatic isolation; The automatic isolation includes dynamic task migration, communication isolation, partial task restart and resource recovery.

7. The multi-scene multi-base large model engine system according to claim 1 is characterized in that: The large model automated evaluation platform includes multi-dimensional performance evaluation, automated testing and quantitative feedback, and model version comparison and error analysis; Among them, multi-dimensional performance evaluation is used to evaluate the performance of the model in multiple dimensions, including accuracy, response time, and resource consumption; Automated testing and quantitative feedback are used for scenario-based automated testing functions; Model version comparison and error analysis is used to perform horizontal comparative analysis of different model versions.

8. The multi-scene multi-base large model engine system according to claim 1, characterized in that: The large model management platform includes intelligent routing and model adaptation, AI-Agent development and operation platform, and full life cycle management; The intelligent routing and model adaptation automatically selects and integrates the optimal model by combining the results of the large model evaluation platform; The AI-Agent development and operation platform is based on an enhanced retrieval algorithm for knowledge texts and is used to combine contextual information with a real-time updated knowledge base; The lifecycle management is used to cover the model training, evaluation, selection, deployment and iterative optimization process to form a closed-loop management.

9. The multi-scene multi-base large model engine system according to claim 1, characterized in that: The large-model automated evaluation platform also includes task weighting and dynamic scoring, wherein the task weighting and dynamic scoring set up a task weighting system according to scenario requirements, and the task weighting system includes static weight allocation and dynamic weight allocation.

10. The multi-scene multi-base large model engine system according to claim 1, characterized in that: The large-model automated evaluation platform also includes scenario-based and multi-task evaluations, which provide feedback and verification on the model by simulating real business environments and multi-task adaptability evaluations.

Citation Information

Patent Citations

  • Sensitive data management system and management method based on block chain

    CN115665145A

  • Automatic financial report question and answer method and device based on large model

    CN117235233A

  • Platform for building intelligent operation system based on enterprise management information

    CN119046415A

  • Intelligent question number large-screen display system and method based on multi-modal large model

    CN119088820A

  • Selection and optimization of prediction algorithms for 3GPP network data analytics function (NWDAF)

    US20240040402A1

Cited By

  • Data asset secure storage method based on block chain technology

    CN120257332A

  • Well logging data interpretation method, device and equipment based on large model fine tuning and storage medium

    CN120430421A

  • Logging data interpretation method, device and equipment based on large model fine-tuning and storage medium

    CN120430421B

  • Power market large language model evaluation method and device based on dynamic scene perception

    CN120910511A

  • Large model evaluation method and device

    CN121029623A