Full-automatic data processing and large model fine tuning system based on AI-Agent
Through the fully automated data processing system based on AI-Agent, the problem of inefficient identification of complex concepts and data structure in financial data processing is solved, and the accurate identification and efficient conversion of financial data is realized, the accuracy and efficiency of data processing are improved, and the needs of multilingual and global financial business are adapted to the needs of multilingual and global financial services.
Patent Information
- Application Number
- CN202510620374.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-08-19
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing technology is difficult to accurately identify complex financial concepts in financial data processing, and the traditional data structured methods are inefficient, unable to fully utilize the value of the financial knowledge base, and cannot meet the needs of financial institutions for data processing accuracy and efficiency.
A fully automated data processing system based on AI-Agent is adopted, including user interaction modules, data processing technology modules, functional support modules, data storage modules and interface modules. Multimodal recognition technology, automatic cleaning algorithms and large language models are used, and financial recognition Agents are combined, structured output Agents and Q&A generation Agents are achieved through knowledge embedding algorithms and dynamic programming algorithms.
It improves the accuracy and efficiency of financial data processing, ensures data security, enhances the robustness of the system, supports multilingual processing, adapts to global financial business scenarios, and provides powerful data processing and intelligent service support.
Smart Images

Figure CN120508636A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of financial data processing, and in particular to a fully automated data processing and large-model fine-tuning system based on AI-Agent. Background Art
[0002] In the financial sector, data processing and analysis are crucial to the decision-making and operations of financial institutions. Financial data is characterized by large amounts of information, diverse formats, and high levels of technical expertise. The core task of financial data processing is to quickly and accurately identify key information from this massive amount of financial data and transform it into structured data.
[0003] In existing financial data processing, financial concept identification often relies on traditional keyword matching and simple pattern recognition methods. Data structured processing is typically achieved through manually formulated rules or simple template matching, lacking effective processing of complex data structures and semantic information. Financial knowledge bases are often utilized through simple text retrieval, making it difficult to deeply explore relationships between knowledge.
[0004] However, existing technologies have numerous shortcomings. Traditional financial concept recognition methods are prone to misjudgments and omissions, and are unable to accurately process semantically complex financial concepts. Data structuring methods that manually formulate rules and template matching are inefficient and difficult to adapt to changes in data formats and requirements. Simple text retrieval methods fail to fully leverage the value of financial knowledge bases, fail to provide comprehensive and accurate support for financial analysis and decision-making, and struggle to meet the growing demands of financial institutions for data processing accuracy and efficiency. To address this, we propose a fully automated data processing and large-scale model fine-tuning system based on AI agents. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a fully automated data processing and large-model fine-tuning system based on AI-Agent, which solves the problems existing in traditional financial concept recognition methods.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a fully automated data processing and large model fine-tuning system based on AI-Agent, including the following modules:
[0007] User interaction module: with registration and login, file upload, scene selection and chat dialogue functions;
[0008] Data processing technology module: It can perform user registration and login verification, use multimodal recognition technology to identify and extract data, use fusion automatic cleaning algorithms to clean data, and carry out complex data processing;
[0009] Functional support module: includes account management, multimodal recognition technology integration, pre-made cleaning rule setting, and agent construction based on locally deployed large language models;
[0010] Data storage module: uses database to store structured data in system operation;
[0011] Logging module: records user operations, data processing process and system operation status information;
[0012] The interface module includes data identification and extraction interface, data cleaning interface, block interface and Agent interface.
[0013] Preferably, the multimodal recognition technology integrates OCR, Surya, Florence-2, and Whisper Small algorithms to process document, voice, and image multimodal data.
[0014] Preferably, the fusion automatic cleaning algorithm for data cleaning in the data processing technology module includes removal of sensitive information, text normalization and quality filtering operations, and supports multi-language mixed processing. The removal of sensitive information is combined with regular expressions, keyword matching and semantic understanding technology; text normalization utilizes deep learning text generation models, quality filtering combined with heuristic rules and machine learning classifiers.
[0015] Preferably, a closed-loop ecology of "data input - intelligent processing - scenario adaptation - knowledge retrieval - accurate response" is constructed in the data processing technology module. By establishing a regression model to analyze the relationship between the change in noise rate before and after data cleaning and the change in the accuracy of the Agent system, the effectiveness analysis results of the closed-loop optimization of the data processing process are obtained.
[0016] Preferably, the Agent in the functional support module is constructed based on the locally deployed open source large language model DeepSeek, and the Dify platform is used to build a specific function Agent system.
[0017] Preferably, the specific function agents include financial identification agents, structured output agents, and question-answer generation agents;
[0018] The Financial Identification Agent uses the RAG financial knowledge base, combined with natural language processing and machine learning algorithms, to identify and classify key information in financial data;
[0019] The structured output agent uses data mapping and template matching technology to convert unstructured information into structured data;
[0020] The question-answer generation agent uses answer-driven automatic question-answer generation technology and combines anomaly detection and correction mechanisms to generate answers.
[0021] Preferably, the financial identification agent uses a knowledge embedding algorithm to convert financial concepts in the RAG financial knowledge base into low-dimensional vector representations. When identifying key information of financial data, it calculates the similarity between the input data and the knowledge base vector to determine the matching financial concepts; the structured output agent uses a dynamic programming algorithm to optimize the data mapping and template matching process in the process of converting unstructured information into structured data, and adjusts the matching strategy in real time according to data characteristics and target structure.
[0022] Preferably, the data storage module constructs a RAG financial knowledge base by collecting financial knowledge from multiple channels and integrates it with the Agent system. In the knowledge retrieval link, a vector similarity matching algorithm is used to retrieve relevant information. Through experimental comparison, a data confidence mechanism is established when using and not using the RAG knowledge base to calculate the accuracy of the question-answering process.
[0023] Preferably, the data identification and extraction interface in the interface module traverses the specified input directory files, calls the corresponding API parsing interface according to the file extension, and stores the parsed text content as a JSON file.
[0024] Preferably, the data cleaning interface processes JSON files in batches, calls three types of cleaning algorithms to clean the data, and stores the processed data in a designated directory.
[0025] This invention provides a fully automated data processing and large-model fine-tuning system based on AI-Agent. It has the following beneficial effects:
[0026] 1. This invention uses a financial identification agent and a structured output agent to work together. It first uses a knowledge embedding algorithm to convert RAG financial knowledge base concepts into vectors. It then preprocesses and vectorizes the input financial data. It then uses a Euclidean distance algorithm to match financial concepts. Finally, it uses a dynamic programming algorithm to convert unstructured data into structured data, which is then output. This series of algorithms accurately identifies financial concepts and efficiently transforms data structures, comprehensively improving the accuracy and efficiency of financial data processing and assisting financial institutions in conducting in-depth analysis and decision-making.
[0027] 2. This invention uses regular expressions, keyword matching, and semantic understanding technology to accurately remove sensitive information and ensure data security. It uses deep learning models to achieve text normalization and combines heuristic rules and machine learning classifiers to complete quality filtering to improve data quality and standardization. It supports multi-language mixed processing and adapts to global financial business scenarios. The integration of multiple technologies reduces data errors and noise, enhances system robustness, and lays a solid foundation for financial data processing, analysis, and model training.
[0028] 3. By integrating multimodal recognition, data cleaning, agent construction and other technologies, the present invention improves data processing efficiency and quality, ensures data security and privacy, enhances the application effectiveness in the financial field, optimizes the fine-tuning effect of large models, and improves user experience with a simple and easy-to-use interface and efficient interactive functions, providing financial institutions and users with powerful data processing and intelligent service support. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] Figure 1 This is the module diagram of the fully automated data processing and large model fine-tuning system based on AI-Agent;
[0030] Figure 2 This is a flow chart of the structured financial data output of the present invention;
[0031] Figure 3 Schematic diagram of the data processing technology flow of the present invention;
[0032] Figure 4 This is a project architecture diagram of the present invention;
[0033] Figure 5 This is a diagram showing the division of the core components of the present invention;
[0034] Figure 6 This is the data cleaning classification diagram of the present invention. DETAILED DESCRIPTION
[0035] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the specification of the present invention. Obviously, the embodiments described are only some embodiments of the present invention, not all embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0036] Example:
[0037] Please see the attached Figure 1 - Attachment Figure 6 , an embodiment of the present invention provides a fully automated data processing and large model fine-tuning system based on AI-Agent, including the following modules:
[0038] User interaction module: It has registration and login, file upload, scenario selection and chat dialogue functions. Users upload documents, audio and video and other multi-format data, and the information is extracted through multimodal recognition technologies such as OCR, converted into structured JSON storage and executed according to preset cleaning rules to form a standardized data base. After the user selects the three core application scenarios, the system starts the Agent thinking module, analyzes the questions through Embedding vectorization technology, and conducts related searches with the search engine and the knowledge base that can be uploaded independently to obtain TopK matching results. Finally, relying on the deep search capability of the knowledge base, it integrates and outputs structured professional content, completes cause analysis and background interpretation, and builds a closed-loop ecosystem of "data input - intelligent processing - scenario adaptation - knowledge retrieval - accurate response", providing efficient and accurate intelligent business chat support for scenarios such as financial analysis and enterprise data management, enabling value mining and decision-making upgrades in business chat scenarios, and realizing a full-link business chat system with intelligent data processing and scenario-based interaction.
[0039] Overall technical architecture
[0040] Stage 1: Data Upload
[0041] Focusing on multi-format file input compatibility, it supports more than 20 file formats including PDF, DOCX, MP3 audio, video files, etc., providing multiple data sources for subsequent processing and ensuring that all types of user documents can be connected to the system.
[0042] Stage 2: Data Identification and Extraction
[0043] It uses multimodal recognition technology, integrated OCR word recognition, specializing in text recognition and extraction of images and traces, Surya can be used for image or video content analysis, Florence-2 multimodal model is good at processing mixed text and image content, WhisperSmal l voice recognition model, and accurate audio-to-text conversion. Four algorithms are used to uniformly convert document content of different modalities into JSON format, realizing cross-modal and accurate extraction of text, voice, image and other information, and opening up the key link from unstructured data to structured expression.
[0044] Stage 3: Data filtering and cleaning
[0045] Build a data purification system through the "prefabricated cleaning rules" module, and perform three core operations with the help of built-in and customizable sensitive word libraries.
[0046] First, sensitive information such as user privacy data is removed to avoid data security risks at the source; second, text normalization is performed to unify format standards; finally, quality filtering is performed to screen out duplicate, erroneous or low-value data, and a loop optimizer mechanism is used to ensure that the output data meets security and regulatory requirements, laying a solid data foundation for advanced processing.
[0047] Stage 4: Advanced Data Processing
[0048] By integrating cleaned data, search engine capabilities, and the user-uploaded RAGflow knowledge base, the system conducts in-depth processing using an agent system. Ultimately, the results are applied to diverse financial scenarios: powering account management robots to intelligently understand and respond to household inquiries; assisting fraud detection models in extracting abnormal patterns from transaction data; and supporting compliance monitoring systems in summarizing and analyzing regulatory documents. This approach closely aligns with the LLM fine-tuning requirements for financial applications in the competition, achieving a closed-loop process from data processing to value creation in these scenarios.
[0049] Data processing technology module: It can perform user registration and login verification, use multimodal recognition technology to identify and extract data, use fusion automatic cleaning algorithms to clean data, and carry out complex data processing;
[0050] Functional support module: includes account management, multimodal recognition technology integration, pre-made cleaning rule setting, and agent construction based on locally deployed large language models;
[0051] Data storage module: uses database to store structured data in system operation;
[0052] Logging module: records user operations, data processing process and system operation status information;
[0053] Interface module, including data identification and extraction interface, data cleaning interface, block interface and Agent interface;
[0054] Data identification and extraction interface
[0055] Omniparse_run is an interface function for batch parsing files of different formats. It traverses the files in the specified input directory, calls the corresponding API parsing interface according to the file extension, and stores the parsed text content as a JSON file.
[0056] 1) Parameter explanation
[0057] Omniparse_run parameter explanation
[0058]
[0059] 2) Return results
[0060] Omniparse_run returns the result
[0061] Type Description
[0062] Successfully returned json format data: {"text": extracted text}
[0063] Failed Parsing failed and returned an error message
[0064] 3) Processing flow
[0065] ① Make sure the output_dir directory exists to store the parsing results.
[0066] ②Traverse the input_dir directory and obtain all processable files.
[0067] ③Select the API endpoint based on the file extension and send the file as multipart / form-data.
[0068] ④Parse the JSON returned by the API and extract the text segment.
[0069] ⑤Save the result as a JSON file with the naming format: {original file name}.json.
[0070] ⑥ Output log
[0071] 4.2 Data Cleansing Interface
[0072] Datajuicer_run is used to batch process JSON files, call three types of cleaning algorithms for data cleaning, and store the processed data in the specified directory.
[0073] 1) Parameter explanation
[0074] Datajuicer_run parameter explanation
[0075] Parameter type description
[0076] dataset_folder str dataset directory
[0077] export_path str Cleaning result output directory
[0078] 2) Return results
[0079] Generate the processed file {original file name}.json and file attribute file {original file name}_stats.json in the export_path directory
[0080] 3) Processing flow
[0081] ①Traverse all .json files in the dataset_folder directory.
[0082] ② Use the thread pool to concurrently perform the following steps to clean the data: a. Read the config_all.yaml configuration file.
[0083] b. Generate a temporary YAML configuration file to store cleaning parameters.
[0084] c. Execute process_data.py to clean the data.
[0085] d. After cleaning is complete, delete the temporary YAML file.
[0086] ③After cleaning is completed, delete the monitor directory under the export_path directory
[0087] 4.3 Block Interface
[0088] The chunk_run function processes the JSON parsed files in the datajuicer_output directory, cleans the text, and generates text files according to the specified pattern.
[0089] 1) Parameter explanation
[0090] Chunk_run Parameter Explanation
[0091] Parameter type description
[0092] mode str application scenario, passed in by the front end
[0093] chunk_size int chunk size, determined by the large model capability
[0094] 2) Return results
[0095] Chunkrun returns results
[0096] type
[0097] illustrate
[0098] A .txt file is generated in the chunk_output directory successfully, with each line containing JSON records.
[0099] Failed to generate the file and returned an error message
[0100] 3) Processing flow
[0101] ① Initialize the output directory: Make sure the chunk_output directory exists.
[0102] ② Traverse JSON files: Only process .json files, ignoring files ending with _status.json. Read the JSON data and ensure it is not empty and correctly formatted.
[0103] ③Text processing: Extract the content of the “text” segment and split it into paragraphs according to \n\n.
[0104] ④ Text cleaning: Remove LaTeX formulas and special symbols. Merge short paragraphs (length < 30).
[0105] ⑤Data storage can be based on the selected application scenario):
[0106] Chatbot mode: Generate text_id for each paragraph and store it in a .txt file.
[0107] Other modes: Each chunk_size segment forms a chunk, and each chunk is saved as a separate .txt file containing multiple text_id-text records.
[0108] ⑥ Save results: All processed .txt files are stored in the chunk_output directory 4.4Agent interface
[0109] Dify_run is used as an Agent interface to batch process text files. The core process includes:
[0110] ① File upload: Upload the .txt file in the chunk_output directory to DifyAPI.
[0111] ②API interaction: Call DifyChatAPI to process text and extract the returned JSON data.
[0112] ③Data cleaning: Extract the “answer” segment from the JSON result returned by the API.
[0113] ④Data storage: Save the extracted content to the dify_output directory.
[0114] ⑤ File merging: Integrate the files in the dify_output directory and merge them into the final output directory files.
[0115] 1) Parameter explanation
[0116] Dify_run parameter explanation
[0117]
[0118] 2) Request API explanation
[0119] ①Upload files
[0120] Responsible for uploading .txt files to the Dify platform and returning file_id for subsequent chat-messages requests
[0121] Upload file method
[0122] Http request method Url path
[0123] POST https: / / api.dify.ai / v1 / files / upload
[0124] Upload file parameter explanation
[0125]
[0126] Upload file return result
[0127] Type Description
[0128] Successfully returns the file ID in the chat room
[0129] Failure returns error information
[0130] ②Send a secret request
[0131] Responsible for uploading .txt files to the Dify platform and returning file_id for subsequent chat-messages requests
[0132] How to send a secret request
[0133] Http request method Url path
[0134] POST https: / / api.dify.ai / v1 / chat-messages
[0135] Send Tianji request parameter explanation
[0136] Parameter type description
[0137] Headers dict request headers
[0138] JSON-body object file information, workflow parameters, query content, etc.
[0139] Send Tianji request and return result
[0140] Type Description
[0141] A txt file containing the output content of each node in the workflow is successfully returned
[0142] Failure returns error information
[0143] 3) Return results
[0144] Generate the processed txt file in the headquarters folder
[0145] 4) Processing flow
[0146] ① Read the .txt file in the chunk_output directory
[0147] ②Concurrently upload these files to DifyAPI
[0148] ③Get the JSON data returned by the API
[0149] ④ Extract "answer" related content
[0150] ⑤Store the processed results in the dify_output directory
[0151] ⑥Integrate the data and merge the block files into the output directory.
[0152] The multimodal recognition technology integrates OCR, Surya, Florence-2, and WhisperSmal l algorithms to process multimodal data such as documents, voice, and images.
[0153] The integrated automatic cleaning algorithm for data cleaning in the data processing technology module includes sensitive information removal, text normalization, and quality filtering operations, while supporting multi-language mixed processing. Sensitive information removal is combined with regular expressions, keyword matching, and semantic understanding technologies; text normalization utilizes a deep learning text generation model, quality filtering combined with heuristic rules and machine learning classifiers, and establishes the following operation steps:
[0154] Step 1: Remove sensitive information
[0155] Sub-step 1.1: Regular Expression Matching
[0156] Define a series of regular expression patterns to match common sensitive information, such as ID card numbers, bank card numbers, and phone numbers. Traverse the text and use regular expressions to match. If a match is successful, remove or replace the matched content.
[0157] Sub-step 1.2: Keyword matching
[0158] Prepare a list of sensitive keywords, traverse the text, and check whether the text contains the keywords in the list. If so, replace the keywords with an empty string or other specified characters;
[0159] Sub-step 1.3: Semantic Understanding Technology
[0160] Use a pre-trained language model to encode the text, and then use a classifier to determine whether the text contains sensitive information. If so, the sensitive part is processed;
[0161] Corresponding formula:
[0162] ScaledDot-ProductAttention:
[0163]
[0164] Where Q is the query matrix, K is the key matrix, V is the value matrix, and d k is the dimension of the key
[0165] Multi-HeadAttention:
[0166] MultiHead(Q,K,V)=Concat(head1,…,head h )W O
[0167] in, and W O is the learnable weight matrix, h is the number of heads
[0168] Step 2: Text Normalization
[0169] Sub-step 2.1: Using a Deep Learning Text Generation Model
[0170] Input the text into the pre-trained text generation model, and the model generates normalized text based on the input;
[0171] Corresponding formula:
[0172] Linear combination: z = w0 + w1x1 + w2x2 + ... + w n x n =w T x+b
[0173] Sigmoid function:
[0174]
[0175] Prediction probability: P(y=1|x)=σ(w T x+b)
[0176] In text generation, the model predicts the next word x t+1 The probability distribution of is:
[0177] P(x t+1 |x1,x2,…,x t )=softmax(W vocab h t +b vocab )
[0178] Among them, W vocab is the vocabulary projection matrix, h t is the hidden state of the decoder at time step t, b vocab is the bias term;
[0179] Step 3: Quality Filtering
[0180] Sub-step 3.1: Heuristic rule filtering
[0181] Filter text based on preset heuristic rules, such as text length, repeated character ratio, etc. For example, filter out text that is too short or too long;
[0182] Sub-step 3.2: Machine Learning Classifier Filtering
[0183] Use a trained machine learning classifier (such as logistic regression) to classify the text and determine whether the text is low-quality text. If so, filter out the text
[0184] Corresponding formula:
[0185] Linear combination: z = w0 + w1x1 + w2x2 + ... + w n x n =w T x+b
[0186] Sigmoid function:
[0187] Prediction probability: P(y=1|x)=σ(w T x+b)
[0188] Where w is the weight vector, x is the input feature vector, b is the bias term, and y is the classification label (0 or 1);
[0189] For multilingual mixed data, we use language detection models to distinguish text languages and apply independent filtering using sensitive word libraries in Chinese and English. Secondly, for text normalization, we use natural language processing techniques such as word segmentation, part-of-speech tagging, and syntactic analysis to convert text into a unified format, correct spelling errors, standardize professional terminology, and ensure text consistency.
[0190] Secondly, in the quality filtering phase, we combine heuristic rule-based machine learning classification to quickly filter out obviously low-quality data, such as duplicate text and content that is too short or too long. We also use human classifiers to accurately identify potentially low-quality data, such as data containing noise or incomplete information, to improve overall data quality. Through the integration of these technologies, we achieve automated cleansing of financial data throughout the entire process, effectively reducing the risk of PII leakage and meeting the needs of diverse financial scenarios.
[0191] The data processing technology module constructs a closed-loop ecosystem of "data input - intelligent processing - scenario adaptation - knowledge retrieval - accurate response". By establishing a regression model to analyze the relationship between the change in noise rate before and after data cleaning and the change in the accuracy of the Agent system, the effectiveness analysis results of the closed-loop optimization of the data processing process are obtained. The following algorithm is established here:
[0192] 1. Data Input: Users upload documents, audio, and video data in various formats through the system's user interaction module. The system utilizes multimodal recognition technology, integrating algorithms such as OCR, Surya, Florence-2, and WhisperSmal l, to uniformly convert document content from various modalities into JSON format. OCR technology extracts text from images and scans; Surya supports multiple languages, handles various document types, and analyzes layouts; Florence-2 understands and generates text associated with images; and WhisperSmal l specializes in speech recognition.
[0193] 2. Intelligent processing
[0194] 2.1 Data Cleaning: Utilizes a fusion automatic cleaning algorithm. To remove sensitive information, a combination of regular expressions and keyword matching is used. Regular expressions are used to match sensitive information formats such as ID card numbers and bank card numbers, and keyword matching is then used for further screening to reduce the risk of sensitive information leakage. Text normalization utilizes natural language processing techniques, such as word segmentation, part-of-speech tagging, and syntactic analysis, to convert text into a unified format, correct spelling errors, and standardize professional terminology. Quality filtering combines heuristic rules with machine learning classifiers. Heuristic rules filter out duplicate text and content that is too short or too long, while classifiers accurately identify potentially low-quality data.
[0195] 2.2 Data Segmentation and Feature Extraction: Cleaned data is segmented according to specific patterns to prepare for subsequent processing. For example, segmentation can be based on text length, semantics, and other factors. Feature extraction algorithms are selected based on different data types and application scenarios. For example, for text data, a bag-of-words model or TF-IDF algorithm may be used to extract text features and convert the text into numerical vectors for easier processing.
[0196] 3. Scenario Adaptation: Based on the application scenario selected by the user, such as question and answer generation, fraud detection, and compliance monitoring, the system uses classification algorithms to map data to the corresponding scenario. The decision tree algorithm can classify data based on its characteristics and gradually divide the data into different scenario branches by judging different features in the data. The support vector machine algorithm finds an optimal hyperplane to separate the data points of different scenarios, thus achieving adaptation of data and scenarios.
[0197] 4. Knowledge retrieval: User questions are parsed into vector form through embedding vectorization technology, and then the search engine is coordinated with the knowledge base that can be uploaded independently to conduct related retrieval. Vector similarity matching algorithms, such as the cosine similarity algorithm, are used to calculate the similarity between the user question vector and the text vector in the knowledge base. The formula is A is the user question vector, and B is the text vector in the knowledge base. The higher the similarity, the stronger the relevance between the text and the question, thereby obtaining the Top-K matching results.
[0198] 5. Accurate response
[0199] 5.1 Answer Generation: Relying on the deep search capabilities of the knowledge base, the Aqent system in the system integrates and processes the retrieved information. The question and answer generation agent uses answer-driven automatic question and answer generation technology, combining the knowledge base content and natural language processing algorithms to generate answer content;
[0200] 5.5 Answer Optimization and Output: Utilize anomaly detection and correction mechanisms to determine the rationality of answers by setting rules or using machine learning models. If an answer is found to be abnormal or inaccurate, re-search the knowledge base or adjust the generation strategy, optimize the answer, and output it to the user.
[0201] The Dify platform was selected as the functional support module to build an AI agent system for structured processing of cleaned data. Dify uses visual process orchestration, such as dialogue flow nodes and tool call links, to transform the complex multi-round dialogue logic in financial scenarios into a graphical configuration. The system consists of three core agents: a financial identification agent, a structured output agent, and a question-and-answer generation agent.
[0202] The financial identification agent leverages the RAG financial knowledge base to accurately identify key information within data. To build the RAG financial knowledge base and enhance the agent system's ability to process financial data, this project employed the following technical approaches. Financial knowledge was collected from multiple sources, including authoritative financial websites, professional financial databases, and industry reports, covering financial terminology, theories, models, and regulations and policies. User-provided financial data, such as company financial statements, financial transaction records, and portfolio information, was also collected. After parsing and structuring, this data was incorporated into the knowledge base. Dify, a low-code platform for building agents, comes with its own knowledge base. However, the knowledge base has limited file size and is not ideal for embedding and segmenting unstructured data. Therefore, this project chose RAGFlow to build an external knowledge base and integrate it with Dify. Technically, the project leveraged Dify's external knowledge base API to integrate RAGflow with Dify. By configuring the API Endpoint and API Key, Dify connects to the RAGflow knowledge base. Within RAGflow, the project organizes collected financial knowledge and user data to build a structured financial knowledge base, regularly updating the content to ensure its currency and accuracy. When a user raises a financial question, the Agent system retrieves relevant information from the knowledge base through RAGflow and, combined with the user's data, generates accurate and detailed Q&A content.
[0203] For example, when a user inquires about a company's financial status, the system can retrieve the company's financial statement data from the knowledge base, analyze its assets and liabilities and profitability, and generate a corresponding response. This integration not only enhances the Agent system's data processing capabilities in the financial field, but also enables efficient retrieval of user information and generation of questions and answers, providing users with more professional and valuable financial information services. The combination of Dify and RAGFlow has significantly improved the functionality and development efficiency of intelligent applications, providing strong support for knowledge-driven AI application development and promoting the widespread application of intelligent technologies in real-world scenarios.
[0204] The structured output agent accurately translates the identified information into structured data. Through standardized processing procedures and algorithms, it translates unstructured financial data into a structured format that is easy to analyze and process, improving the usability and manageability of the data.
[0205] The Q&A Generation Agent utilizes answer-driven automatic Q&A generation technology, combined with the RAG financial knowledge base and hallucination detection and correction system to ensure accurate and reliable answers. Based on the user's question, the agent extracts relevant information from the knowledge base and generates clear and accurate responses, overcoming the hallucination issues that can occur in traditional Q&A systems and improving the efficiency and quality of information users acquire.
[0206] The specific function agents include financial identification agents, structured output agents, and question-answer generation agents;
[0207] The Financial Identification Agent uses the RAG financial knowledge base, combined with natural language processing and machine learning algorithms, to identify and classify key information in financial data;
[0208] The structured output agent uses data mapping and template matching technology to convert unstructured information into structured data;
[0209] The question-answer generation agent uses answer-driven automatic question-answer generation technology and combines anomaly detection and correction mechanisms to generate answers.
[0210] The financial identification agent uses a knowledge embedding algorithm to convert financial concepts in the RAG financial knowledge base into low-dimensional vector representations. When identifying key information in financial data, it calculates the similarity between the input data and the knowledge base vectors to determine the matching financial concepts. The structured output agent uses a dynamic programming algorithm to optimize the data mapping and template matching process in the process of converting unstructured information into structured data, and adjusts the matching strategy in real time based on data characteristics and target structure. The algorithm established here is as follows:
[0211] 1. Collect and organize RAG financial knowledge base and financial data to be processed
[0212] Collect various financial documents, reports, news and other data to build the RAG financial knowledge base, and obtain new financial data that needs to be processed, such as financial reports and transaction records;
[0213] 2. Convert financial concepts in the RAG financial knowledge base into low-dimensional vector representations
[0214] The knowledge embedding algorithm is used to process the financial concepts and their relationships in the knowledge base. First, the entities (financial concepts) and relationships in the knowledge base are represented as vectors in a vector space. Then, the vector representation is adjusted by training based on the loss function of the TransE algorithm.
[0215] TransE algorithm:
[0216] Loss function:
[0217]
[0218] Distance metrics:
[0219]
[0220] Among them, S is the set of positive sample triplets, S′ (h,r,t) is the set of negative sample triplets generated for the positive sample (h, r, t), γ is the interval size, x and y are vectors, and n is the dimension of the vector;
[0221] 3. Preprocess the financial data to be processed and convert it into vector representation
[0222] The input financial data is preprocessed by tokenization, stop word removal, and stemming. The processed data is then converted into low-dimensional vectors using the same feature extraction and vectorization methods used for knowledge base concept embedding.
[0223] 4. Calculate the similarity between the input data vector and the knowledge base vector to determine the matching financial concepts
[0224] Use the vector similarity calculation algorithm to calculate the similarity between the input data vector and the financial concept vector in the knowledge base. Set a similarity threshold and use financial concepts with scores higher than the threshold as matching results.
[0225] Euclidean distance algorithm:
[0226] in, is the vector representation of the input data, is the vector representation of a financial concept in the knowledge base, and n is the dimension of the vector;
[0227] 5. Determine the target structured data template
[0228] Define the output structured data format based on business requirements and data characteristics, such as designing table structure and field names;
[0229] 6. Initialize the state and parameters of the dynamic programming algorithm
[0230] Define the dynamic programming state dp[i][jj] to represent the optimal matching score when processing the i-th element of the unstructured data and the j-th element of the template. Initialize the dp array, dp[0][0] = 0;
[0231] For i = 1 to the length of the unstructured data:
[0232] dp[i][0]=dp[i-1][0]+penalty delete
[0233] For j = 1 to template length:
[0234] dp[0][j]=dp[0][j-1]+penalty insert
[0235] 7. Calculate the value of the dp array according to the state transfer equation
[0236] According to the state transition equation of dynamic programming
[0237]
[0238] Calculate the values of the dp array in sequence. Where match(i,j) represents the score when the i-th element of the unstructured data matches the j-th element of the template, and penalty delete Indicates the penalty score for deleting an element in unstructured data, penalty insert represents the penalty score for inserting an element in unstructured data;
[0239] 8. Start backtracking from the last element of the dp array to generate structured data
[0240] Start backtracking from dp[unstructured data length][template length], and determine whether to match, delete, or insert based on the maximum value when the state transition occurs, gradually building up structured data.
[0241] 9. Output structured financial data
[0242] Output the generated structured data in a suitable format (such as table, JSON, etc.) for subsequent analysis and use.
[0243] The data storage module builds a RAG financial knowledge base by collecting financial knowledge from multiple channels and integrates it with the Agent system. In the knowledge retrieval link, it uses a vector similarity matching algorithm to retrieve relevant information. Through experimental comparison, a data confidence mechanism is established when using and not using the RAG knowledge base to calculate the accuracy of the question-answering process.
[0244] The data identification and extraction interface in the interface module traverses the specified input directory files, calls the corresponding API parsing interface according to the file extension, and stores the parsed text content as a JSON file.
[0245] The data cleaning interface processes JSON files in batches, calls three types of cleaning algorithms to clean the data, and stores the processed data in a specified directory.
[0246] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. AI-Agent-based fully automated data processing and large model fine-tuning system, characterized by: Includes the following modules: User interaction module: with registration and login, file upload, scene selection and chat dialogue functions; Data processing technology module: can perform user registration and login verification, use multimodal recognition technology to identify and extract data, use fusion automatic cleaning algorithms to clean data, and carry out complex data processing; Functional support module: includes account management, multimodal recognition technology integration, pre-made cleaning rule setting, and agent construction based on locally deployed large language models; Data storage module: uses database to store structured data in system operation; Logging module: records user operations, data processing process and system operation status information; The interface module includes data identification and extraction interface, data cleaning interface, block interface and Agent interface.
2. The AI-Agent-based fully automated data processing and large model fine-tuning system according to claim 1 is characterized in that: The multimodal recognition technology integrates OCR, Surya, Florence-2, and Whisper Small algorithms to process multimodal data such as documents, voice, and images.
3. The AI-Agent-based fully automated data processing and large model fine-tuning system according to claim 1 is characterized in that: The integrated automatic cleaning algorithm for data cleaning in the data processing technology module includes removal of sensitive information, text normalization and quality filtering operations, and supports multi-language mixed processing. The removal of sensitive information combines regular expressions, keyword matching and semantic understanding technology; text normalization utilizes deep learning text generation models, quality filtering combined with heuristic rules and machine learning classifiers.
4. The AI-Agent-based fully automated data processing and large model fine-tuning system according to claim 1 is characterized in that: The data processing technology module constructs a closed-loop ecosystem of "data input - intelligent processing - scenario adaptation - knowledge retrieval - accurate response". By establishing a regression model to analyze the relationship between the change in noise rate before and after data cleaning and the change in the accuracy of the Agent system, the effectiveness analysis results of the closed-loop optimization of the data processing process are obtained.
5. The AI-Agent-based fully automated data processing and large model fine-tuning system according to claim 1 is characterized in that: The Agent in the functional support module is built based on the locally deployed open source large language model DeepSeek, and uses the Dify platform to build a specific function Agent system.
6. The AI-Agent-based fully automated data processing and large model fine-tuning system according to claim 5 is characterized in that: The specific function agents include financial identification agents, structured output agents, and question-answer generation agents; The Financial Identification Agent uses the RAG financial knowledge base, combined with natural language processing and machine learning algorithms, to identify and classify key information in financial data; The structured output agent uses data mapping and template matching technology to convert unstructured information into structured data; The question-answer generation agent uses answer-driven automatic question-answer generation technology and combines anomaly detection and correction mechanisms to generate answers.
7. The AI-Agent-based fully automated data processing and large model fine-tuning system according to claim 6 is characterized in that: The financial identification agent uses a knowledge embedding algorithm to convert financial concepts in the RAG financial knowledge base into low-dimensional vector representations. When identifying key information in financial data, it calculates the similarity between the input data and the knowledge base vectors to determine the matching financial concepts. The structured output agent uses a dynamic programming algorithm to optimize the data mapping and template matching process in the process of converting unstructured information into structured data, and adjusts the matching strategy in real time according to data characteristics and target structure.
8. The AI-Agent-based fully automated data processing and large model fine-tuning system according to claim 1 is characterized in that: The data storage module builds a RAG financial knowledge base by collecting financial knowledge from multiple channels and integrates it with the Agent system. In the knowledge retrieval link, it uses a vector similarity matching algorithm to retrieve relevant information. Through experimental comparison, a data confidence mechanism is established when using and not using the RAG knowledge base to calculate the accuracy of the question-answering process.
9. The AI-Agent-based fully automated data processing and large model fine-tuning system according to claim 1 is characterized in that: The data identification and extraction interface in the interface module traverses the specified input directory files, calls the corresponding API parsing interface according to the file extension, and stores the parsed text content as a JSON file.
10. The AI-Agent-based fully automated data processing and large model fine-tuning system according to claim 1 is characterized in that: The data cleaning interface processes JSON files in batches, calls three types of cleaning algorithms to clean the data, and stores the processed data in a specified directory.