Financial data processing method and device, computer device and storage medium

By using object detection algorithms and cross-modal models to automate the processing of financial images, high-quality multimodal datasets are generated. This solves the quality and applicability issues in multimodal data acquisition and VQA data generation in the financial field, and achieves efficient and accurate data processing and model application.

CN121861640BActive Publication Date: 2026-07-24SHANGHAI UNIVERSITY OF FINANCE AND ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHANGHAI UNIVERSITY OF FINANCE AND ECONOMICS
Filing Date
2026-03-16
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing technologies in the financial field suffer from difficulties in acquiring multimodal data and generating VQA data, resulting in inconsistent quality and an inability to meet professional analysis needs, leading to poor model training and application results.

Method used

We employ object detection algorithms and cross-modal models to automatically identify and classify financial images. By combining multi-model collaborative processing and multi-layer filtering mechanisms, we generate high-quality multimodal datasets, including steps such as image annotation, cross-modal feature matching, and environmental perturbation testing, to ensure the accuracy and applicability of the data.

Benefits of technology

It enables efficient and automated collection and processing of financial image data, significantly reducing labor costs, improving data quality and the accuracy of model applications in the financial field, and generating question-and-answer pairs closely aligned with financial scenarios to support in-depth financial reasoning and analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121861640B_ABST
    Figure CN121861640B_ABST
Patent Text Reader

Abstract

The application discloses a financial data processing method and device, computer equipment and a storage medium. The method comprises: acquiring a financial image; detecting the financial image according to a target detection algorithm, identifying a financial image region, labeling key elements in the financial image region, and generating an image labeling file; based on a preset financial image classification label, using a cross-modal model to perform cross-modal feature matching on the image labeling file to obtain a classification result and a classification confidence; when the classification confidence is lower than a preset threshold, associating and fusing the key elements and the classification result through a data interface to correct the classification result and the classification confidence. The embodiment of the application realizes automatic and high-precision collection of massive image data in the financial field, reduces the cost of manual screening, and provides high-quality basic image materials for subsequent VQA data synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to a financial data processing method, apparatus, computer equipment, and storage medium, belonging to the field of multimodal big data processing and artificial intelligence technology. Background Technology

[0002] The training of large-scale language models in the financial field currently faces a critical challenge in acquiring multimodal data. On the one hand, financial images and their corresponding visual question-answering (VQA) data suffer from significant difficulties in acquisition and inconsistent quality. Traditional image collection methods rely on single recognition algorithms, resulting in low recognition accuracy and poor screening efficiency; while image quality screening relies too heavily on manual judgment, making it difficult to meet the processing needs of massive amounts of data. These problems severely restrict the effective training and application of multimodal models in the financial field.

[0003] Existing technologies have significant shortcomings in the data generation stage. VQA data generation is prone to disconnect from real-world financial scenarios and insufficient difficulty gradient design, resulting in training data that fails to meet professional analytical needs. Traditional methods lack a deep understanding of financial business scenarios and cannot accurately recreate complex financial analysis processes, causing the generated question-and-answer data to remain superficial and unable to support in-depth financial reasoning and analysis.

[0004] To address the aforementioned pain points, there is an urgent need to build an efficient, accurate, and deeply tailored multimodal big data processing solution for financial scenarios. Summary of the Invention

[0005] In view of this, this application provides a financial data processing method, apparatus, computer equipment, and storage medium. The embodiments of this application realize the automated and high-precision collection of massive image data in the financial field, reduce the cost of manual screening, and provide high-quality basic image materials for subsequent VQA data synthesis.

[0006] The first aspect of this application discloses a financial data processing method, the method comprising: acquiring a financial image; detecting the financial image according to a target detection algorithm, identifying financial image regions, and annotating key elements within the financial image regions to generate an image annotation file; performing cross-modal feature matching on the image annotation file based on preset financial image classification labels using a cross-modal model to obtain classification results and classification confidence; when the classification confidence is lower than a preset threshold, associating and fusing the key elements with the classification results through a data interface to correct the classification results and classification confidence.

[0007] Furthermore, acquiring the financial image includes: recording detailed source information of the financial image, which includes at least one of the following: the issuing institution, report name, and publication date of the financial research report; the disclosing entity, year, and disclosure platform link of the corporate annual report; the name, version number, and open-source license type of the open-source dataset; the domain name, data crawling time, and corresponding financial product code of the financial website; and using differentiated permission verification methods to verify the detailed source information from different sources. If the detailed source information is compliant, the financial image is acquired and a copyright registration permission certificate ledger is generated. The permission certificate ledger includes a unique data identifier, source details, permission certificate document, registration time, and registrant.

[0008] Furthermore, the financial images include at least one of the following: line chart, bar chart, pie chart, financial relationship diagram, financial statement, supporting data table, official seal image, and candlestick chart.

[0009] Furthermore, it also includes: based on the business objectives of each sub-scenario in the front-end, middle-end, and back-end, and according to the principle of difficulty stratification, pre-configuring scenario-specific prompt word templates. The scenario-specific prompt word templates include three elements: data type, task objective, and business constraints, used to limit the content generated by the large model to fit financial logic. Financial images and corresponding text data are input with the scenario-specific prompt word templates into a visual-text collaborative reasoning model to generate question-answer pairs, while ensuring data scale through quantity control. The question-answer pairs are then classified and quality-screened using a financial semantic classification model to form a multimodal dataset covering multiple sub-scenarios. The quality screening includes: questions must contain financial professional terms corresponding to the sub-scenario, and answers must have unique and verifiable derivation basis.

[0010] Furthermore, the step of inputting financial images and corresponding text data with the scene-specific prompt word template into the visual-text collaborative reasoning model to generate question-answer pairs, while ensuring data scale through quantity control, includes: matching the corresponding text data to the corresponding sub-scene template, filling the prompt words with specific data information according to the corresponding sub-scene template, converting the financial images into a model-compatible input format to form a visual data-structured prompt word input package; inputting the input package into the visual-text collaborative reasoning model to obtain question-answer pairs; determining whether the number of question-answer pairs for each type of sub-scene meets the preset value, and if not, regenerating question-answer pairs by supplementing the same type of input package or adjusting the prompt word template until the requirements are met.

[0011] Furthermore, it also includes: using the visual-text collaborative reasoning model to automatically score the multimodal dataset in multiple dimensions, retaining question-answer pairs that simultaneously meet the threshold requirements of five dimensions: image information density, QA semantic validity, data diversity, objectivity, and computational complexity; verifying the filtered question-answer pairs and corresponding financial images according to manual annotation and a preset first data review standard; and making a decision on the verified question-answer pairs according to an expert consensus voting mechanism. When all experts unanimously determine that the verified question-answer pairs meet the preset second data review standard, the verified question-answer pairs are included in the final dataset.

[0012] Furthermore, the acquisition of financial images also includes: generating four types of environmental disturbance data based on the financial images to test the robustness of the multimodal large language model under non-ideal visual conditions; the four types of environmental disturbance data include: key information occlusion: performing partial occlusion or blurring on the core decision information area in the financial image, with the range controlled within a first threshold range of the core decision information area, simulating scenarios such as document folding, stain coverage, or incomplete scanning; redundant image interference: superimposing visually identical but business-irrelevant image elements on the financial image, using a second threshold transparency coverage, simulating scenarios such as multiple reports on the same page or misinsertion of charts; missing relevant information: deleting some content in the financial image that is related to the question-and-answer task; and superimposing irrelevant information: adding text or graphic elements unrelated to the task to the edges of the financial image, occupying a third threshold visual space, simulating scenarios such as handwritten annotations, advertising watermarks, or irrelevant text annotations.

[0013] A second aspect of this application discloses a financial data processing apparatus, comprising: an acquisition module for acquiring a financial image; a generation module for detecting the financial image using a target detection algorithm, identifying a financial image region, and labeling key elements within the financial image region to generate an image labeling file; a matching module for performing cross-modal feature matching on the image labeling file based on preset financial image classification labels using a cross-modal model to obtain a classification result and a classification confidence level; and a correction module for correcting the classification result and classification confidence level by associating and fusing the key elements with the classification result through a data interface when the classification confidence level is lower than a preset threshold.

[0014] A third aspect of this application discloses a computer-readable storage medium comprising a stored program, wherein the program, when running, controls the execution of the financial data processing method of the above embodiments in the processor of the device.

[0015] A fourth aspect of this application discloses a computer device, the computer device including a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed by the above-described financial data processing method.

[0016] Compared with the prior art, the embodiments of this application have the following beneficial effects: 1) Cost reduction and efficiency improvement: By automating the entire process from data crawling to filtering, replacing the traditional manual methods, data processing costs can be significantly reduced and processing cycles can be greatly shortened.

[0017] 2) Quality Improvement: By adopting multi-model collaborative processing and multi-layer filtering mechanism, the output financial image classification is ensured to have high accuracy, while the error rate of VQA data is controlled at an extremely low level, and the overall data quality is significantly improved compared with the existing solution.

[0018] 3) Scenario Adaptation: The generated data is closely focused on the professional analysis needs in the financial field, such as research report interpretation and investment decision-making. It can effectively improve the accuracy of the model in related tasks and has practical value that can be directly applied to various intelligent financial products. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on the structures shown in these drawings without creative effort.

[0020] Figure 1 A flowchart illustrating a financial data processing method provided in this application embodiment.

[0021] Figure 2 This is a structural diagram of a financial data processing system provided in an embodiment of this application.

[0022] Figure 3 This is a structural diagram of a financial data processing device provided in an embodiment of this application. Detailed Implementation

[0023] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0024] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0025] Example 1: Figure 1 This is a flowchart illustrating a financial data processing method provided in an embodiment of this application. Figure 1 As shown, the method includes: S101. Obtain financial images.

[0026] In this embodiment, visual data acquisition follows the core principle of precise matching between scenario, source, and tool. Differentiated acquisition paths and tools are designed for eight types of key visual data in financial scenarios to ensure a high degree of data compatibility with financial business needs. The eight types of key visual data in financial scenarios (also known as financial images) include line charts, bar charts, pie charts, financial relationship diagrams, financial statements, supporting data tables, official seal images, and candlestick charts.

[0027] Regarding the collection of line charts, bar charts, pie charts, and financial relationship diagrams, the target data must meet the following requirements: ≥2 data dimensions, distinguishable trend characteristics, clear coordinate axes and legends, and the ability to support tasks such as market trend interpretation, industry comparison, and entity correlation analysis. In terms of sources, research reports published by financial institutions are selected, with priority given to reports marked "publicly available for academic research," ensuring the data's relevance and accessibility. For collection tools, a customized PDF parser is used to filter reports containing sections on "data trend analysis," "industry comparison," and "entity correlation logic," such as the "2024 Banking Industry Profit Trend Analysis" section.

[0028] Regarding the collection of financial statements and supporting data tables, the financial statements must include a complete balance sheet, income statement, and cash flow statement structure. The supporting data tables must contain multi-step calculation logic and financial professional indicators to support both the routine financial analysis tasks in the "financial analysis and decision support" scenario of the middle platform and the complex data reasoning tasks in the "financial risk control and asset optimization" scenario of the back-end. The primary data source is publicly available annual reports from enterprises. Annual report PDFs are obtained through official interfaces of enterprise information disclosure platforms (such as the disclosure platform designated by the China Securities Regulatory Commission), and financial statements are extracted using PDF table extraction tools to ensure data field completeness. The secondary source is professional examination question banks such as the Chinese Certified Public Accountant (CPA) exam and actuary exam. Materials containing table-based questions are filtered through compliant exam question bank interfaces, and supporting data tables requiring multi-step calculations and involving the interpretation of professional terminology are extracted to ensure the data is professional and challenging.

[0029] Regarding the acquisition of official seal images, the target data must be typical official seals in the financial sector (such as corporate financial seals, bank seals, and seals of financial regulatory agencies). The text clarity of the seal in the image must be ≥90%, with no obvious occlusion or blurring. This data can be used in the front-end "financial seal recognition" sub-scenario to test the model's ability to recognize and verify compliance markers on financial documents (such as seals on contracts and bills). The data source is exclusively the TrOCR-Seal-Recognition dataset from the open-source dataset Gmgge, obtained directly through the dataset's official interface. During acquisition, images of seals with clear annotations and no copyright restrictions are selected, while images with blurry text or those from non-financial sectors (such as seals of administrative organs) are removed. Images are also uniformly cropped to a specification where the seal area occupies ≥80% of the image to ensure the model focuses on the core recognition object.

[0030] For candlestick chart data collection, the target data needs to cover mainstream financial products such as stocks and futures, with time periods including daily and weekly candlesticks, and be correlated with technical indicators such as MACD (Moving Average Convergence Divergence) and RSI (Relative Strength Index). This data should support the front-end "candlestick chart analysis" sub-scenario, evaluating the model's ability to interpret technical trading signals (such as long lower shadows, MACD golden / death crosses, and RSI overbought / oversold signals). Data sources are targeted through web scraping tools from publicly available financial websites (such as stock exchange websites and compliant financial data service provider platforms). The web scraping parameters are preset as follows: "Financial product type: Shanghai and Shenzhen A-shares, domestic commodity futures; Time range: the past 5 years; Technical indicators: default loading of MACD (parameters 12, 26, 9) and RSI (parameters 6, 12, 24); Image format: including complete candlestick period annotations, indicator curves, and values." After scraping, the images undergo consistency processing (such as standardizing axis scales and indicator color annotations) to avoid format differences affecting subsequent QA processes for synthesis and model evaluation.

[0031] To ensure that all collected data is free of copyright disputes and complies with the requirements for academic research and public use, a three-step compliance verification process of "source tracing - permission confirmation - copyright registration" is designed to form a complete copyright control ledger. Specifically, the acquisition of financial images includes: S1011. Record detailed source information of the financial images, including at least one of the following: the issuing institution, report name and publication date of the financial research report; the disclosing entity, year and disclosure platform link of the corporate annual report; the name, version number and open source license type of the open source dataset; the domain name, data crawling time and corresponding financial product code of the financial website.

[0032] In this embodiment, the original data acquisition path can be traced back through the source information to ensure data traceability.

[0033] S1012. For the detailed source information from different sources, a differentiated permission verification method is used for verification. If the detailed source information is compliant, the financial image is acquired and a copyright registration permission certificate ledger is generated. The permission certificate ledger includes a unique data identifier, source details, permission certificate document, registration time, and registrant.

[0034] Financial research reports, corporate annual reports, and publicly available financial website data should have their data labeling verified through official channels. Only content labeled "publicly available for academic research" or "free for non-commercial use" should be selected. For data requiring authorization, usage authorization documents should be obtained through academic cooperation channels. For open-source datasets, the terms of their open-source licenses (such as the MIT or Apache licenses) must be verified to confirm that the license allows for secondary use, modification, and distribution without additional commercial licensing restrictions. For professional examination question bank data, data usage agreements should be signed with the examination organizers or compliant question bank service providers, clearly stating that the data is only used for dataset construction in academic research and may not be used for commercial profit purposes.

[0035] It is worth noting that, regarding copyright registration, a corresponding ledger of "data-source-proof of authorization" is established for all data that has been verified through authorization. The ledger includes the data's unique identifier (such as an MD5 value), source details, proof of authorization documents (such as a scanned copy of the license agreement or a link to an open-source license), registration time, and registrant. The ledger is stored in encrypted form to ensure that copyright information is searchable and verifiable, preventing copyright disputes in subsequent data use.

[0036] S102. Detect the financial image according to the target detection algorithm, identify the financial image region, and annotate the key elements in the financial image region to generate an image annotation file.

[0037] Continuing with examples of line charts, bar charts, pie charts, and financial relationship diagrams, the YOLO algorithm is used to perform page-by-page image detection on the crawled PDF documents, identify financial image regions (such as line charts, bar charts, pie charts, etc.) in the documents, and accurately label key elements (axis ticks, data labels, chart titles) within the images, generating image annotation files containing element location information.

[0038] S103. Based on the preset financial image classification labels, a cross-modal model is used to perform cross-modal feature matching on the image annotation file to obtain the classification result and classification confidence.

[0039] Continuing with examples of line charts, bar charts, pie charts, and financial relationship diagrams, we apply pre-defined financial image classification labels (such as "revenue trend chart," "balance sheet chart," and "stock candlestick chart") based on the CLIP model. We then perform cross-modal feature matching on the images detected by YOLO, completing image segmentation (removing document background and text interference areas) and classification. The output is a clean dataset of classified financial images, thus distinguishing between line charts, bar charts, pie charts, and financial relationship diagrams while excluding irrelevant images. Furthermore, the extracted images are uniformly converted to PNG format to preserve the original data clarity and ensure that subsequent models can accurately identify chart elements.

[0040] S104. When the classification confidence level is lower than a preset threshold, the key elements are associated and fused with the classification result through a data interface to correct the classification result and the classification confidence level.

[0041] The YOLO element annotation results are linked and fused with the CLIP classification results through a data interface. When the CLIP classification confidence is lower than a preset threshold (such as 85%), the key element information of YOLO annotation is called for secondary verification to correct the classification results.

[0042] For example, the system obtains the list of key elements of the image output by the YOLO model (e.g., ["Axis Labels: Quarterly (Q1-Q4)", "Data Labels: Revenue (RMB 100 Million)", "Chart Type: Bar Chart"]) and the candidate categories and their confidence scores output by the CLIP model (e.g., [("Revenue Trend Chart", 0.82), ("Balance Sheet Chart", 0.15)]). Then, it enters the logical matching and adjudication stage: it iterates through the CLIP candidate categories and queries a pre-defined financial chart knowledge base to obtain the required and excluded element sets for each category. By calculating the matching degree between YOLO elements and the knowledge base rules—that is, by counting the supporting evidence (the number of elements matching the required elements) and opposing evidence (the number of elements appearing in the excluded element set)—a final adjudication is made based on the matching results. If the YOLO evidence strongly supports the CLIP preferred classification (e.g., matching all required elements and having no excluded elements), the confidence of that classification is significantly increased and adopted. If the YOLO evidence contradicts the preferred classification but strongly supports other candidate classifications, the CLIP result is overturned and the classification supported by YOLO is adopted. If the evidence is ambiguous and cannot clearly support any candidate classification, the image is marked as "requiring manual review," thus completing the closed-loop processing flow from evidence extraction to fusion decision.

[0043] This embodiment employs automated web crawling technology to batch crawl research reports, financial reports, and annual reports in PDF format from publicly available financial information websites (such as stock exchange announcement platforms and official websites of financial data service providers). It introduces the YOLO object detection algorithm and the CLIP cross-modal model to construct a dual-model collaborative processing architecture, thereby achieving high-precision automatic classification and processing of financial images, significantly improving the accuracy of image classification and overall processing efficiency.

[0044] In one embodiment, the method further includes: text data association with the goal of strongly binding "visual data-text information" to extract text content directly related to visual data, providing financial business background and logical support for subsequent scenario-based QA synthesis.

[0045] In this embodiment, for line charts, bar charts, and other images in financial research reports, the text descriptions and data interpretation conclusions of the chapters containing the images are extracted to obtain the corresponding business background and data logic. For financial statements in corporate annual reports, the interpretations of the report data in the notes to the financial statements and the management discussion and analysis chapters are extracted to clarify the calculation methods and business meanings of the financial data. For the accompanying data tables in professional exam question banks, the question stems and answer explanations are extracted to obtain the data calculation logic and explanations of professional terms. For official seal images and candlestick charts, the corresponding scene description text is extracted to clarify the business scenarios and related factors of the images. Finally, the extracted text data and corresponding visual data are bound together with a unique identifier and stored in a "visual data + text data" package format (e.g., each data contains a PNG image and a TXT text). This ensures that during the subsequent QA pair synthesis process, the model can simultaneously obtain visual information and text background to generate QA content that conforms to the financial business logic and avoids QA pairs being detached from actual business scenarios.

[0046] In some embodiments, the method further includes: Step 1. Based on the business objectives of each sub-scenario in the front-end, middle-end, and back-end, and according to the principle of difficulty stratification, pre-configure scenario-specific prompt word templates. The scenario-specific prompt word templates include three elements: data type, task objective, and business constraints, which are used to limit the content generated by the large model to fit the financial logic.

[0047] In this embodiment, the prompt word design is centered on "aligning with financial business logic and constraining the direction of generated content". It is led by experts with more than 10 years of experience in the financial industry. It combines the business objectives and data characteristics of multiple sub-scenarios of the front-end, middle-end and back-end to build scenario-based prompt word templates to ensure that the generated QA pairs are highly adapted to actual financial business.

[0048] It should be noted that the template design follows the principles of scenario uniqueness, element completeness, and difficulty stratification: Scenario uniqueness requires each sub-scenario to have a corresponding exclusive prompt word template to avoid cross-scenario confusion. For example, the front-end "Financial Seal Recognition" sub-scenario template focuses on "official seal entity recognition and compliance judgment," while the back-end "Financial Strategy Optimization" sub-scenario template focuses on "strategy adjustment logic and risk-return balance." Element completeness requires all templates to clearly include the data type of the specified associated visual data (financial images), clearly define the QA's task objectives for the business purpose, and define the business constraints that limit the boundaries of financial business logic, avoiding the generation of content detached from actual business operations. Difficulty stratification requires that the same sub-scenario cover three levels of difficulty: basic, medium, and complex, to match the capability requirements of different business stages.

[0049] Example (1): In the "K-line chart analysis" sub-scenario on the front end, the basic difficulty template requires "judging the pattern of a single K-line", and the complex difficulty template requires "judging the future trend by combining the MACD golden cross / death cross". The template example for this sub-scenario is as follows: "Data type: daily K-line chart with MACD indicator; Task objective: answer questions related to stock price trends based on K-line patterns and MACD indicator; Business constraints: the correlation logic between the length of the K-line body, the characteristics of the shadow line and the position relationship of the MACD indicator (such as the DIFF line and DEA line) must be clearly mentioned, and the answer must be directly verifiable through the data in the chart. Please generate 3 single-choice questions, each with 4 options, and the correct answer is unique."

[0050] Example (2), the template for the "Investment Analysis" sub-scenario of the middle platform is: "Data type: corporate financial statements (balance sheet + profit statement); Task objective: evaluate the investment value of enterprises based on the report data and answer investment decision-related questions; Business constraints: it needs to involve the calculation of core financial indicators such as ROE (return on equity) and debt-to-equity ratio, and needs to be compared and analyzed with the industry average level. The questions need to include decision-oriented content such as "whether to invest" and "investment risk points". Please generate 2 open-ended questions and 1 multiple-choice question. The answers need to be verifiable by deduction through report data and industry common sense."

[0051] Example (3), the template for the "Financial Risk and Policy Analysis" sub-scenario in the background is: "Data type: line graph containing interest rate fluctuations + policy text fragments; Task objective: analyze the risk impact of interest rate policies on the financial market and answer questions related to risk assessment; Business constraints: the causal relationship between policies (such as central bank interest rate cuts) and interest rate fluctuations must be clearly defined, professional terms such as "liquidity risk" and "market pricing risk" must be mentioned, and the questions must include counterfactual assumptions (such as "what would be the trend of risk changes if interest rates were not lowered"). Please generate 2 true / false questions and 1 multiple-choice question, and the answers must conform to the logic of financial policy transmission."

[0052] Step 2. Input the financial images and corresponding text data, along with the scene-specific prompt word templates, into the visual-text collaborative reasoning model to generate question-answer pairs, while ensuring the data scale through quantity control.

[0053] Using visual data, associated text data, and corresponding contextualized prompts as input, the initial QA pairs are generated based on the QwenVL-Plus-latest model, which has been verified by InstructBLIP to have reliable visual-text collaborative reasoning capabilities. At the same time, the data scale is ensured through quantity control.

[0054] Step 21. Match the corresponding text data to the corresponding sub-scene template, and fill in the specific data information for the prompt words according to the corresponding sub-scene template. Convert the financial image into a model-compatible input format to form a visual data-structured prompt word input package.

[0055] During the input data preprocessing stage, the eight types of visual data collected are assigned to corresponding sub-scene templates according to the one-to-one correspondence principle of "visual data type - scene template". For example, candlestick charts are matched with the front-end "candlestick chart analysis" template, financial statements are matched with the middle-end "investment analysis" and back-end "financial data reasoning and interpretation" templates, and official seal images are matched with the front-end "financial seal recognition" template. The visual data is also converted into a model-compatible input format (such as PNG format, 300dpi resolution), and specific data information is filled in for prompt words according to the corresponding scene template (such as "data type: profit statement of a new energy company in 2024"), forming a unified input package of "visual data + structured prompt words".

[0056] Step 22. Input the input packet into the visual-text collaborative reasoning model to obtain question-answer pairs.

[0057] During model generation, to ensure the professionalism and accuracy of the generated content, the inference parameters of the QwenVL-Plus-latest model are set as follows: Temperature = 0.3 (to reduce randomness and ensure answer stability), Maximum Generation Length = 512 tokens (to meet the QA requirement for complete expression), and top_p = 0.9 (to focus on high-probability reasonable output). The model first parses key information in the visual data (such as the opening / closing price of the K-line, the core fields of the financial statements, and the text content of the official seal), and then combines the task objectives and business constraints in the prompts to generate QA pairs that meet the needs of the scenario. For example, if the input is "K-line chart with MACD indicator + front-end K-line chart analysis template", the model will generate the question "Based on the K-line and MACD indicator in the chart, did the stock show a MACD golden cross on December 6, 2024?" and the answer "Yes, on that day the DIFF line crossed the DEA line from bottom to top, and the K-line closed with a long lower shadow, verifying the validity of the golden cross".

[0058] Step 23. Determine whether the number of question-answer pairs for each sub-scene meets the preset value. If not, regenerate question-answer pairs by supplementing the same type of input package or adjusting the prompt word template until the requirements are met.

[0059] Regarding quantity control, for the eight types of visual data, the number of QA pairs generated for each type of data must be ≥2000 to ensure that a single type of data can support the evaluation needs of the corresponding sub-scenario. For example, for the front-end sub-scenario "K-line chart analysis" for candlestick charts, ≥2000 QA pairs covering different candlestick patterns (long lower shadow, doji, double top) and MACD / RSI indicator combinations must be generated. For financial statements, ≥1000 QA pairs must be generated for each of the middle-end sub-scenario "investment analysis" and the back-end sub-scenario "financial data reasoning and interpretation", totaling ≥2000. If the number of generated QA pairs for a single type of data is insufficient, it can be regenerated by supplementing the same type of visual data and adjusting the prompt word template (such as adding "constraints to generate questions of different difficulty") until the quantity requirement is met.

[0060] Step 3. Use a financial semantic classification model to classify and screen the question-answer pairs for different scenarios, forming a multimodal dataset covering multiple sub-scenarios. The quality screening includes: the question must contain financial terminology for the corresponding sub-scenario, and the answer must have a unique and verifiable reasoning.

[0061] The Qwen-max model, which has strong text classification and financial semantic understanding capabilities, is used to classify the initial QA pairs into scenarios and select content that meets the quality requirements, thus initially forming a multimodal dataset.

[0062] For example, in terms of scenario classification logic, the Qwen-max model first provides multiple sub-scenarios with "scenario definitions + keyword libraries" as classification criteria: the front-end sub-scenario "financial seal recognition" is defined as "the task of identifying the official seal entity in financial documents and verifying compliance," and the keyword library includes "official seal," "financial seal," "compliance," and "authenticity." The back-end sub-scenario "asset allocation analysis" is defined as "the task of optimizing multi-asset class allocation and balancing risk and return," and the keyword library includes "asset allocation," "portfolio," "risk-return ratio," and "investment portfolio."

[0063] After the initial QA pairs are input into the Qwen-max model, the model uses semantic matching (comparing the terms in the QA pairs with the keyword libraries of each scenario) and business logic judgment (analyzing the fit between the QA pair's task objective and the scenario definition) to assign each QA pair to a unique sub-scenario: QA pairs containing "official seal ownership" and "verify authenticity" are categorized into the front-end "financial seal recognition," while QA pairs containing "asset allocation ratio" and "investment portfolio returns" are categorized into the back-end "asset allocation analysis." Finally, 10% of the classification results are randomly selected for review by financial experts to assess their accuracy. If the accuracy rate is <95%, the scenario keyword library and classification prompts are optimized (e.g., supplementing the "sub-scenario - irrelevant terms" exclusion list), and the classification process is re-executed until the accuracy rate is ≥95%.

[0064] Furthermore, the quality screening follows two main criteria: the inclusion of financial terminology and the unique verifiable answer. The inclusion of financial terminology requires QA pairs to contain core professional terms relevant to the sub-scenario, excluding basic questions lacking professional content. For example, QA pairs in the front-end "Financial Data Statistics" sub-scenario must contain terms such as "year-on-year growth rate" and "month-on-month comparison," while those in the middle-end "Industry Analysis and Inference" sub-scenario must contain terms such as "industry chain" and "competitive landscape." QA pairs without corresponding terms (e.g., "Does the chart contain a line graph?") will be eliminated. The unique verifiable answer requires the answer to have clear verification evidence (directly readable from visual data, calculated using fixed formulas, or derived from financial business rules), excluding QA pairs with ambiguous or multiple solutions. For example, the question "According to the financial statements, what was the company's operating revenue in 2023?" (the answer can be directly read from the statements) meets the requirement, while the question "What is the company's future revenue growth potential?" (the answer is subjective and cannot be verified by existing data) will be eliminated. After scenario classification and quality screening, QA pairs that were incorrectly classified, lacked professional terminology, or had unverifiable answers were removed. Content that met the requirements was retained, initially forming a multimodal dataset covering multiple sub-scenarios: approximately 8,700 QA pairs were retained in the front-end "Financial Knowledge and Data Analysis" scenario, approximately 4,650 in the middle-end "Financial Analysis and Decision Support" scenario, and approximately 2,498 in the back-end "Financial Risk Control and Asset Optimization" scenario, laying the foundation for in-depth review by the subsequent quality control layer.

[0065] In some embodiments, the method further includes: Step 4. Utilize the visual-text collaborative reasoning model to perform automated multidimensional scoring on the multimodal dataset, retaining question-answer pairs that simultaneously meet the threshold requirements of five dimensions: image information density, QA semantic validity, data diversity, objectivity, and computational complexity.

[0066] In this embodiment, the first-level control uses the Qwen-VL-Plus-latest model as the core tool. It automatically filters the initially formed multimodal QA pairs through a preset multi-dimensional scoring system, eliminating erroneous, ambiguous and low-quality data from the source, and ensuring that the data entering the subsequent review stage has basic usability.

[0067] This phase focuses on five core evaluation dimensions: image information density, QA semantic validity, data diversity, objectivity, and computational complexity. Each dimension has clearly defined scoring criteria and thresholds (maximum score 10 points, threshold ≥ 6 points). The image information density dimension assesses the richness of key information in the visual data (such as indicator annotations in candlestick charts or the completeness of fields in financial statements), preventing images lacking information (such as line charts containing only a single data point) from entering the subsequent process. The QA semantic validity dimension judges the logical consistency between the question and the answer. For example, if the question asks "What is the ROE of a certain company in 2023?", the answer must be calculated or interpreted around that indicator. If the answer is irrelevant... Questions that only describe company revenue are deemed invalid; the data diversity dimension requires that QA pairs generated from similar visual data cover different business scenarios or levels of difficulty. For example, candlestick chart QA pairs should include different candlestick patterns (long lower shadow, doji) and indicator combinations (MACD, RSI) to avoid repeating a single type of question; the objectivity dimension requires excluding subjective and speculative content, such as QA pairs like "Will this stock rise in the future?" which lack clear verification basis; the computational complexity dimension requires that QA pairs involving numerical calculations include multi-step reasoning (such as calculating the gross profit margin from the profit and loss statement data and then comparing it with the industry average), excluding simple questions that only require direct data reading.

[0068] The Qwen-VL-Plus-latest model scores each QA pair according to the above criteria, retaining only content that meets the threshold requirements in all five dimensions, and removing data that fails to meet the standard in a single dimension or whose total score in multiple dimensions is less than 30. At the same time, it verifies the rationality of the filtering logic through exponential decay similarity analysis to ensure that the similarity between the model and human evaluation at this stage is above 0.61, thus reducing the pressure on subsequent manual review.

[0069] Step 5. Verify the filtered question-and-answer pairs and corresponding financial images according to the manual annotation method and the preset first data review standard.

[0070] The secondary control is carried out by a team of undergraduate finance majors who have undergone rigorous screening and training. It focuses on the accuracy of QA details and the adaptability to the scenario, and achieves in-depth verification of the data after automated filtering through manual annotation.

[0071] Before the review, the selection process for annotators requires them to pass an assessment that includes financial expertise (such as financial indicator calculation and understanding of financial terminology) and data review standards (such as QA logic consistency judgment). Only those who score ≥80 points can participate in the annotation work. Furthermore, they must complete practical training with 100 sample data points before formal annotation to ensure that the annotation standards are consistent.

[0072] During the review process, the annotators must verify each QA (Question and Answer) across three core dimensions: First, regarding accuracy, the matching degree between the answer and the visual data must be verified. For example, in the financial statement QA, the "debt-to-equity ratio calculation result" needs to be recalculated against the report data; in the candlestick chart QA, the "MACD golden cross judgment" needs to be confirmed by considering the positional relationship between the DIFF and DEA lines in the chart. Second, regarding scenario matching, it is necessary to check whether the QA conforms to the business logic of its sub-scenario. For example, in the front-end "financial seal recognition" sub-scenario, the QA focuses on the identification of the seal entity. If it contains content outside the scenario scope, such as "analyzing the financial status of the company to which the seal belongs," it is considered a mismatch. Third, regarding the completeness of visual elements, it is necessary to confirm whether the QA fully utilizes the key elements in the visual data. For example, the financial statement QA should cover the linked analysis of the balance sheet and profit and loss statement. If it only involves a single table field and is not associated with other visual information, it is considered incomplete.

[0073] Furthermore, during the annotation process, annotators need to mark questionable data as "pending verification" and record specific issues (such as "error in answer calculation logic" or "misclassification of scene"). After all data annotations are completed, the team needs to conduct cross-validation and review the data marked "pending verification" a second time to ensure the consistency of the review results. In the end, only QA pairs that have passed the review and are undisputed are retained, while content with errors, scene mismatches, or insufficient use of visual elements is removed.

[0074] Step 6. Make a decision on the verified question-and-answer pairs according to the expert consensus voting mechanism. When all experts unanimously determine that the verified question-and-answer pairs meet the preset second data review criteria, the verified question-and-answer pairs will be included in the final dataset.

[0075] The third-level control, as the final stage of quality control, consists of an audit team of three experts with more than 10 years of experience in the financial industry. The team focuses on the rigor of the business logic, policy compliance, and accuracy of professional terminology of the QA, and decides whether the data is included in the final dataset through a consensus voting mechanism.

[0076] Among the core dimensions of the audit, logical rigor requires assessing whether the reasoning chain of the QA (Question and Answer) matches the actual financial business process. For example, the QA for the "Financial Strategy Optimization" sub-scenario in the backend should include a complete logic of "strategy adjustment - risk assessment - return forecast." If there are logical gaps (such as only mentioning strategy adjustment without analyzing risks), it is judged as not rigorous. Policy compliance requires verifying whether the content complies with current financial regulatory policies. For example, QA involving "private equity fund sales" should comply with the "Interim Measures for the Supervision and Administration of Private Investment Funds," avoiding illegal statements (such as "private equity funds can be sold to the general public"). Terminology accuracy requires ensuring the standardized use of financial professional terminology. For example, the calculation method (weighted average or diluted) for "ROE (Return on Equity)" must be clearly stated to avoid terminology confusion (such as mistakenly writing "net profit margin" instead of "gross profit margin").

[0077] The review process employs a "one-vote veto system," requiring three experts to independently evaluate each QA pair. Only when all three agree that the data meets the standards can it be included in the final dataset. If there is disagreement (e.g., two approve and one disapprove), an expert meeting must be held to discuss and reach a consensus based on actual financial business cases (e.g., compliance handling methods for similar scenarios). If a consensus cannot be reached, the data will be removed. If a single QA pair is disapproved by two or more experts, it will be directly excluded.

[0078] In addition, the expert team also needs to summarize common problems found during the review process (such as frequent misuse of certain terms) and feed them back to the data synthesis layer to optimize the prompt word template (such as supplementing terminology specification instructions), forming a closed loop of quality control to ensure the professional credibility of the final dataset.

[0079] Steps S101-S104 can be further applied to the image information density evaluation of primary control, improving the model's automatic recognition accuracy of key elements in visual data, thereby enhancing the efficiency and reliability of overall quality control.

[0080] In some embodiments, acquiring the financial image further includes: generating four types of environmental perturbation data based on the financial image to test the robustness of the multimodal large language model under non-ideal visual conditions.

[0081] In this embodiment, the environmental disturbance data synthesis is an enhancement scheme designed by the VisFinEval platform to improve the challenge and realism of the dataset. By simulating common visual disturbance problems in real financial scenarios, four types of disturbance data are generated to test the robustness of multimodal large language models (MLLMs) under non-ideal visual conditions, ensuring that the model evaluation results are more consistent with actual business application scenarios. The synthesis logic, implementation method, and corresponding test objectives of each type of disturbance data are deeply linked to the pain points of financial business, as detailed below: 1) Key information occlusion: Perform partial occlusion or blurring on the core decision information area in the financial image, with the range controlled within the first threshold range of the core decision information area, simulating scenarios such as document folding, stain coverage, or incomplete scanning.

[0082] The key information occlusion focuses on core decision-making information in financial visual data. By partially occluding or blurring the data, it simulates problems such as "document folding, stains covering, and incomplete scanning" in real-world scenarios, and tests the model's ability to reason and complete missing information.

[0083] During the synthesis process, key areas in the visual data are first located: for candlestick charts, elements that affect trend judgment, such as "opening / closing price labels and MACD indicator curves," are obscured; for financial statements, key statistical information such as "core field data (e.g., asset-liability ratio, total operating revenue) and header field names" are obscured; for official seal images, elements that identify the company, such as "official seal text (e.g., company name, seal type) and the five-pointed star logo," are obscured; and for line charts / bar charts, trend analysis data such as "axis scales and data peak point labels" are obscured.

[0084] The occlusion method employs "random region blurring + local pixel coverage" to ensure that the occlusion range is controlled within 30%-50% of the key areas (neither completely hiding information nor increasing the difficulty of recognition), and the occlusion pattern simulates real-world scenarios (such as irregular stains or linear folding marks). For example, in financial statement perturbations, the value of the "Total Current Assets" field in the balance sheet is obscured, retaining only the field name and some surrounding data; in candlestick chart perturbations, the "intersection area of ​​the DIFF line and DEA line in the MACD indicator" is blurred to test whether the model can use other unobscured candlestick patterns (such as long lower shadows) to assist in judging trends. This type of perturbation data is mainly used to evaluate the model's logical reasoning ability in "incomplete information" scenarios, avoiding the model relying solely on complete data to achieve high accuracy and ignoring the common problem of data corruption in actual business.

[0085] 2) Redundant image interference: Visually similar but business-irrelevant image elements are superimposed on financial images, and the second threshold transparency is used to cover them, simulating scenarios where multiple reports are on the same page or charts are mistakenly inserted.

[0086] Redundant image interference simulates real-world problems such as "multiple reports on the same page, misplaced charts, and other document fragments mixed in during scanning" by superimposing unrelated but visually similar image elements onto the original financial visual data, thus testing the model's ability to distinguish between target data and redundant information.

[0087] When compositing, redundant elements are selected based on the principle of "visual similarity + business irrelevance": For financial statements, similar report segments from other companies are overlaid (e.g., a segment of Company A's profit statement is overlaid on Company B's balance sheet, even though the table formats are the same, the data is different); for candlestick charts, partial candlestick charts from other stocks are overlaid (e.g., a segment of Stock C's daily candlestick chart is overlaid on Stock D's weekly candlestick chart, even though the time periods are similar, the price trends are different); for line charts, trend curves from unrelated industries are overlaid (e.g., a line chart of sales volume in the home appliance industry is overlaid on a line chart of net profit in the banking industry, even though the curve styles are similar, the meanings of the coordinate axes are different).

[0088] The overlay method employs a "semi-transparent layer overlay," controlling the transparency of redundant elements to 40%-60% (ensuring no complete obscuring of the original data, but without creating visual interference). The overlay positions are randomly distributed in non-edge areas of the original image (such as the center of a report or a densely trending area in a candlestick chart). For example, in a line chart of net profit in the banking industry, a semi-transparent overlay of a line chart of sales in the home appliance industry is used. The two curves are similar in color and both are labeled with "year-on-year growth rate." The test aims to determine whether the model can distinguish the target data from redundant information using the axis labels ("Net Profit in the Banking Industry" vs. "Home Appliance Sales"). This type of perturbation data is used to verify the model's "visual focusing ability," preventing the model from misinterpreting data due to irrelevant information and ensuring accurate location of target information even in real-world business scenarios with multiple documents.

[0089] 3) Missing relevant information: Remove the parts of the financial image that are related to the question-and-answer task.

[0090] The missing information test simulates real-world scenarios such as "documents not being scanned across multiple pages, data truncation due to incorrect table formatting, and missing key attachments" by deleting parts of the original visual data that are related to the problem. This tests the model's ability to identify and judge "missing information" and prevents the model from forcibly generating answers when there is insufficient information.

[0091] During the synthesis process, based on the task objectives of the QA pairs, "relevant information necessary to answer the questions" is removed from the visual data: For financial QA pairs involving multiple tables (such as "calculating ROE through the balance sheet and income statement"), only one table is retained, and the other key table is deleted; for candlestick chart QA pairs involving time series comparisons (such as "comparing the candlestick trends of January and February 2024"), the candlestick data for one month is deleted; for line chart QA pairs involving multi-dimensional analysis (such as "analyzing the correlation trend between 'fixed asset investment' and 'industrial added value'"), one line of line data is deleted. For example, in the QA pair "calculating a company's ROE in 2023", only the income statement data is retained, and the "net assets" field data in the balance sheet is deleted. In this case, the model needs to recognize that "net asset data is missing, and ROE cannot be calculated," rather than forcibly generating an incorrect result. The core testing objective of this type of perturbation data is to assess the model's "information integrity judgment ability," ensuring that the model can accurately identify constraints when encountering incomplete data in actual business, rather than generating misleading conclusions. This is crucial for scenarios requiring rigorous data support, such as financial risk control and investment decisions.

[0092] 4) Overlay of irrelevant information: Add text or graphic elements unrelated to the task to the edge of the financial image, occupying the third threshold visual space, to simulate handwritten annotations, advertising watermarks or irrelevant text annotation scenarios.

[0093] Irrelevant information overlay involves adding text or graphic elements unrelated to the task to the original visual data to simulate real-world problems such as "handwritten annotations on documents, watermarks for advertisements, and irrelevant text annotations," testing the model's ability to "select and separate visual information from semantic information."

[0094] When compositing, select elements that are relevant to the financial scenario but irrelevant to the current task: For financial statements, add "handwritten annotations (such as 'This data is pending verification'), company logo watermarks, and irrelevant field comments (such as adding 'Number of employees: 1000' next to the 'Operating Revenue' field)"; for official seal images, add "document number, approver's signature, and irrelevant institution name watermarks"; for candlestick charts, add "stock commentary text annotations (such as 'Buy Recommendation' and 'Risk Warning'), and irrelevant technical indicators (such as adding a randomly generated 'custom indicator' curve next to the MACD indicator)".

[0095] The method of adding text / graphics is to separate them from the original data but present them in the same frame. This ensures that irrelevant information does not directly cover key data, but occupies 10%-20% of the visual space of the image (creating semantic interference). For example, in a candlestick chart, the text annotation "Analyst's recommended holding position" is added to the edge of the image to test whether the model can ignore this irrelevant advice and judge the trend solely based on the candlestick and MACD indicators. This type of perturbation data is used to evaluate the model's "semantic focusing ability," avoiding the model being misled by irrelevant text or graphics, and ensuring that it can still extract core decision data from information-heavy financial documents, which is consistent with the application scenario in actual business where "documents contain multiple types of redundant information."

[0096] Combination Figure 2 As shown, a financial data processing system may include at least: I. Data Collection Module: Utilizes automated web crawling technology to batch crawl research reports, financial reports, and annual reports in PDF format from publicly available financial information websites (such as stock exchange announcement platforms and official websites of financial data service providers); introduces the YOLO object detection algorithm and the CLIP cross-modal model to construct a dual-model collaborative processing architecture.

[0097] II. Data Processing Module; 2.1 The automated image quality screening submodule is used to remove low-quality images (such as blurry or missing information images), reducing invalid calculations in subsequent VQA data synthesis and improving data processing efficiency. The Qwen2.5-72B-VL multimodal large model is used as the core of quality assessment, constructing a two-dimensional evaluation index system: Sharpness assessment: Inputting an image to Qwen2.5-72B-VL, the model outputs an image sharpness score (0-100 points). The score is calculated based on image edge gradient values, pixel uniformity, and noise density. When the score is <60 points, it is judged as a low-sharp image and discarded. Information content assessment: A preset template of key information for financial images (e.g., "containing time dimension + numerical dimension + indicator name") is used. Qwen2.5-72B-VL uses visual understanding and text recognition to determine whether the image meets the template requirements. If key information is missing (e.g., no coordinate axis labels, blurry or unrecognizable data values), it is judged as a low-information image and discarded.

[0098] 2.2 The automated VQA data synthesis submodule is used to generate high-quality VQA pairs tailored to financial scenarios, providing professional training data for the large language model training. Using the Qwen-Plus-VL large model as the core of generation, a VQA generation prompt template library for the financial field is constructed. The templates include scenario-specific terms (such as "based on financial chart analysis" and "interpretation combined with financial indicators"), question type terms (such as "calculation-based," "trend analysis-based," and "risk assessment-based"), and answer format requirements (such as "numerical value + unit + analysis conclusion"). The screened financial images and corresponding templates are input into Qwen-Plus-VL. Based on the image's visual features and financial domain knowledge, the model generates VQA data pairs containing "image ID-question-answer." For example, for a "2023 quarterly revenue bar chart of a certain company," the generated question is "What percentage increase was the company's Q3 2023 revenue compared to Q2? Please analyze the rationality of this increase in conjunction with industry quarterly trends." The answer includes the specific growth rate value, the calculation process, and industry trend matching analysis.

[0099] 2.3 Two-layer quality filtration submodule; 2.3.1 An automated question quality filtering unit is implemented to ensure that generated questions are relevant to financial scenarios and possess a certain level of difficulty, avoiding simplistic and generic issues and enhancing the training value of VQA data. A two-dimensional filtering mechanism is constructed at the technical level: Scenario Relevance Filtering: A financial terminology lexicon (containing 500+ professional terms such as "revenue growth rate," "debt-to-equity ratio," and "price-to-earnings ratio") is built. Token segmentation is performed on the generated questions, and the frequency of term occurrences and their relevance to image content are statistically analyzed. When the percentage of term occurrences is <15% or the relevance is <70% (based on word vector similarity), the question is considered a scenario deviation and is filtered out; for example, non-financial scenario questions such as "What color is this chart?" are removed. Difficulty Gradient Filtering: A difficulty assessment model is built based on token complexity to statistically analyze the logic contained in the question. The number of relational terms (such as "year-on-year / month-on-month", "growth / decline", "reason / impact") and the length of calculation-related tokens (such as "average revenue in the past five years" containing 5 calculation-related tokens) are set with difficulty thresholds: simple questions (number of logical terms < 2 and length of calculation-related tokens < 3) account for ≤ 10%, medium-difficulty questions (number of logical terms 2-3 and length of calculation-related tokens 3-5) account for 40%-50%, and high-difficulty questions (number of logical terms ≥ 4 and length of calculation-related tokens ≥ 6) account for ≥ 40%. By controlling these thresholds, it is ensured that the difficulty of the questions meets the training requirements of financial models.

[0100] 2.3.2 The manual sampling verification unit is used to supplement the omissions of automated filtering, further improve the quality of VQA data, and ensure data reliability. A stratified sampling mechanism is adopted, randomly selecting 5%-10% of the samples from the automatically filtered VQA data. A review team composed of financial analysts (≥3 years of experience) and AI data annotation experts (≥2 years of experience) conducts manual review based on review standards (consistency between question and image, accuracy of answer, and financial professionalism). VQA pairs with problems (such as incorrect answer calculations or questions irrelevant to the image) are marked and fed back to the Qwen-Plus-VL model for prompt optimization, forming a closed loop of "generation-filtering-review-optimization".

[0101] III. Data Evaluation Module; A multi-dimensional evaluation index system was constructed to quantitatively evaluate the VQA dataset: Relevance index: Calculate the semantic similarity between the question and the image (based on the BERT model to calculate the similarity of token embedding vectors), requiring a similarity ≥ 85%; Accuracy index: Randomly select 1000 VQA pairs, have 3 financial experts independently judge the accuracy of the answers, and calculate the percentage of unanimous agreement that the answers are accurate, requiring an accuracy rate ≥ 95%; Completeness index: Check whether the answers contain all the information required by the question (e.g., calculation questions require numerical values, units, and calculation processes), with a completeness compliance rate ≥ 98%; Model adaptability index: Use the VQA data to train a large financial model (e.g., Llama3-70B-Fin), and compare the accuracy improvement of the model before and after training on the financial question-answering test set (containing 500 professional questions), requiring an accuracy improvement ≥ 20%; Evaluation feedback mechanism: If a certain index fails to meet the standard (e.g., accuracy < 95%), trace back to the data processing stage (e.g., Qwen-Plus-VL generation parameters, quality filtering threshold) for adjustment, regenerate the VQA data, and evaluate again until all indexes meet the standard.

[0102] This solution is designed for visual-language financial tasks and is compatible with various mainstream multimodal models, including both closed-source and open-source models. For closed-source models, the system supports the integration of cutting-edge models such as OpenAI's GPT-4V and Anthropic's Claude 3.5 (Sonnet), as well as related versions of the Gemini series (e.g., Google's Gemini 1.5). These closed-source models are typically integrated into the system via API interfaces, leveraging their powerful pre-training capabilities to achieve complex understanding and analysis of financial scenarios. For open-source models, the system also supports direct deployment and fine-tuning of model weights, such as the Qwen-VL series (Qwen-VL, Qwen-VL-Chat, etc.), InternVL, MiniGPT-4, and LLaVA visual-language models. These open-source models can be customized and optimized within the system to adapt to specific needs in the financial field. For example, Alibaba Cloud's open-source Qwen-VL-max achieved an overall accuracy of 76.3% on financial multimodal benchmarks, outperforming non-human experts but still lagging behind financial professionals. In addition, the industry has also seen the emergence of multimodal models customized for the financial sector, such as the FinTrAL series, which achieves financial multimodal capabilities approaching those of GPT-4. This system ensures support for these various models through a flexible architecture, enabling their widespread application in multimodal data processing and analysis within financial business scenarios. In financial applications, different types of multimodal models each have their strengths. Closed-source large models (such as GPT-4V) often excel in complex reasoning and general visual understanding, suitable for parsing diverse content such as financial reports and invoices; open-source models (such as Qwen-VL and LLaVA) can be fine-tuned for specific domains to achieve high accuracy in specific tasks, such as reading invoices, recognizing official seals, or analyzing professional charts. The system leverages these advantages, processing multimodal data (text, tables, graphs, etc.) generated in financial business workflows using the corresponding models, thereby covering the task requirements of each stage from the front-end to the back-end, achieving a comprehensive evaluation and improvement of the models' visual and language understanding capabilities. By being compatible with both closed-source and open-source models, this system has built a unified multimodal model training and evaluation platform that can utilize the latest proprietary model capabilities while also allowing for in-depth customization of open-source models to meet the professional requirements of practical applications in the financial industry.

[0103] The model training module supports efficient training and fine-tuning of the visual-language model, employing an image-text alignment training mechanism to enhance the model's ability to understand joint image and text inputs. Specifically, for common multimodal tasks in the financial field, this system provides comprehensive training support, including image-to-text (OCR) scenarios, such as recognizing text and numbers in financial statements; financial chart understanding, such as analyzing trend signals in candlestick charts and line charts; visual question answering (VQA), such as answering questions based on given financial scenario diagrams or illustrations; and tabular question answering, such as answering relevant indicator values ​​based on financial statement images. By introducing these tasks into training, the model can learn how to convert visual content into accurate language descriptions and answers. For example, the model can identify key values ​​and signals in charts and then provide analytical conclusions in natural language; or the model can extract text information from invoice or contract images through OCR and combine it with financial knowledge to answer user questions. To achieve image-text modal alignment, a cross-attention mechanism is used during training to fuse image features with text features, jointly optimizing the model's visual encoder and text decoder, enabling it to fully utilize image cues and align with text semantics when generating answers. To ensure excellent model fine-tuning performance and applicability to the financial industry, the system constructed a high-quality training sample set supported by financial data. The data construction process fully integrates the scenarios and data characteristics of the VisFinEval benchmark, extracting and labeling multimodal data from real financial business documents. Specific steps include: collecting financial images covering eight categories (such as financial relationship diagrams, candlestick charts, pie charts, financial statements, official seals, etc.) from financial research reports, annual reports, and professional examination question banks; then, domain experts develop question-and-answer pairs or descriptions for each image, forming image-text paired training samples. To ensure the professionalism and accuracy of the aligned samples, the system introduces a multi-layered data verification mechanism: first, a pre-trained model automatically generates preliminary image-text QA pairs, which are then categorized by the model classifier into specific financial business scenarios to ensure comprehensive coverage; subsequently, the generated data undergoes quality checks using a combination of automatic indicator filtering (such as answer accuracy and consistency detection) and manual review. This process is essentially an automated implementation of the VisFinEval data construction method, producing high-quality visual alignment training samples through a closed loop of task generation, data annotation, and quality review. Ultimately, the finely tuned model will achieve ideal alignment results in tasks such as OCR, chart analysis, and financial question answering, significantly improving the model's understanding and response capabilities to financial visual information.

[0104] To achieve fully automated model training and evaluation, the training module of this system adopts a modular design, breaking down the complex training process into several sub-components and providing end-to-end control from task creation to model output. These modules work together to form a complete training pipeline. Task Generation Module: Automatically generates a list of training tasks based on the assessment needs of financial scenarios. This module combines preset scenario templates and real-world business cases, utilizing large-scale models to assist in generating multi-turn dialogue question-and-answer, image and text description, and other training tasks. For example, the system can automatically create a series of question-and-answer tasks with financial statement images based on the "financial statement analysis" scenario, and arrange questions of different difficulties and types to ensure broad training coverage and close resemblance to real business processes.

[0105] The data construction module is responsible for constructing the dataset required for training, including extracting samples from the original data source, labeling, and formatting them. This module retrieves financial images and related text data from databases or file systems, matches and combines them according to the list provided by the task generation module, forming sample pairs of image + text input and expected output. Subsequently, data cleaning and enhancement (such as image resolution optimization and text regularization) are performed, and training and validation sets are split as needed. For automatically generated QA pairs, this module also calls a quality control subprocess to perform multi-dimensional verification and filtering of samples to remove inaccurate or inconsistent data, ensuring the quality of data entering the training process.

[0106] The training scheduling module automatically orchestrates and schedules the training process. Based on a pre-defined training plan, this module allocates datasets and model configurations to computing resources and controls the start, pause, and resumption of training jobs. The scheduling strategy can be dynamically adjusted based on task priority and resource availability, supporting sequential training of multiple models or parallel training of different model instances. In a multi-machine cluster environment, the training scheduling module is also responsible for coordinating job synchronization across nodes, enabling efficient distributed training.

[0107] The resource management module manages and allocates the hardware and software resources required for training. This module monitors the usage of resources such as GPUs, CPUs, and storage in the cluster and optimizes the deployment of training jobs accordingly. For example, when training models with high computational demands, it automatically selects multi-GPU servers or enables model parallelism; for smaller fine-tuning tasks, it appropriately organizes them for execution in a single-GPU or even CPU environment. The resource management module can also select the optimal operating environment based on the characteristics of different models (such as selecting GPU acceleration libraries for Transformer models), improving training efficiency and reducing costs.

[0108] Training Monitoring Module: This module tracks the metrics and status of the model training process in real time, enabling full-process monitoring and feedback. It records metrics such as loss and accuracy for each iteration and provides a visual interface for developers to view training curves. When anomalies occur (such as training divergence or performance stagnation), the monitoring module can issue alerts or automatically adjust parameters such as the learning rate. After training, the module generates a detailed report, including the model's final performance, convergence status, and comparative analysis with historical versions, providing a basis for subsequent model evaluation and deployment.

[0109] Through the collaborative division of labor among the above modules, this system constructs a fully automated model training pipeline. From task definition and data preparation to training execution and quality monitoring, each step is automated, reducing manual intervention and error rates. This design ensures the high efficiency and controllability of the model training process, enabling large-scale financial multimodal model training to run stably in complex environments. At the same time, the modular architecture facilitates expansion and maintenance, allowing for the addition of functional components according to new requirements, ensuring the system can quickly adapt and complete model training when facing emerging financial tasks.

[0110] The system's model training module is compatible with various multimodal prompting strategies, enabling specialized training for complex interactive prompts to improve model performance in reasoning processes and dialogue contexts. First, the system supports Chain-of-Thought (CoT) prompt training, which introduces examples containing reasoning steps into the training data or requires the model to output intermediate reasoning processes. By encouraging the model to "think step-by-step" during training, the model can output its thought process before providing a conclusion when faced with problems requiring multi-step calculations or logical deductions, thereby improving the accuracy of solving complex problems. For example, for questions requiring cross-year comparisons of financial indicators, the model will first explain the comparison steps before arriving at the final answer. VisFinEval evaluations include many multi-step numerical calculations and hypothesis deduction tasks (such as counterfactual reasoning under controlled conditions). For these tasks, the system uses CoT-style fine-tuning to teach the model to generate coherent intermediate reasoning chains, thereby reducing erroneous inferences and illusions in complex business process reasoning. Second, the training module fully supports prompt optimization in multi-turn interactive question-and-answer scenarios. That is, the model can not only accept one-time text and image questions but also handle a series of continuous dialogues, where the context includes the interweaving of preceding text and image information. This system enhances the model's ability to remember and understand dialogue history and context by incorporating multi-turn dialogue data (including multiple question-and-answer exchanges between the user and the model based on the same image, or consecutive questions involving multiple images) into the training corpus. After this training, the model can adjust subsequent responses based on previous rounds of questions and answers in practical applications, achieving context-sensitive analysis. Furthermore, compatibility with mixed image and text context prompts is also a key training focus. The system allows multiple images and text fragments to coexist in single-turn or multi-turn prompts, and guides the model to learn when to reference which image in a dialogue through special labeling and positional encoding. For example, in an investment decision dialogue, the user first provides a stock chart and asks a question; after the model answers, the user adds a text information (news summary) and asks a further question; the trained model will be able to integrate the aforementioned image and text information to provide an accurate answer. In summary, by leveraging the training compatibility with chained reasoning, multi-turn dialogue, and mixed image and text prompts, the model in this system can better understand user intent, track context, and perform multimodal reasoning when handling real financial business consultations, thereby improving the professionalism and reliability of interactive question answering.

[0111] This system is highly compatible with mainstream open-source multimodal training frameworks and models, aligning its training process with industry-standard solutions. Adopting an open architecture, the system seamlessly integrates with multimodal model frameworks such as OFA, BLIP-2, MiniGPT-4, and LLaVA. For visual language models developed based on these frameworks, the system provides corresponding adapters and data interfaces, enabling them to directly utilize the system's training data and scheduling mechanisms for fine-tuning. For example, for visual encoder + text decoder architectures like BLIP-2, the system can load pre-trained visual encoder weights and interface them with the internal language model, then use a unified data pipeline for alignment training. Similarly, for LLaVA models, the system pre-configures corresponding dialogue formats and image input interfaces, ensuring a training process consistent with native implementations. Various open-source models involved in the VisFinEval evaluation (such as InternVL, Molmo, and LLaVA) can all be directly supported for training and evaluation by this system. This broad framework compatibility allows the system to leverage the latest achievements of the open-source community, quickly introduce new models, and validate their performance in the financial field. Meanwhile, for model architectures with special modifications (such as variant models that introduce structured hints), the system can also support them through plug-in module extensions, maintaining flexible adaptability to diverse models. Regarding model fine-tuning, the training module provides various efficient fine-tuning techniques to reduce the cost of adapting large models to financial-specific tasks. First, the system supports LoRA (Low-Rank Adaptation) fine-tuning technology, which adjusts large models by inserting only a small number of trainable parameter matrices, thus allowing the model to learn new financial knowledge without fully training all weights. This enables even visual language models with billions of parameters to complete domain fine-tuning with limited computing power, accelerating training convergence and reducing GPU memory usage. Second, the system is compatible with multimodal adapter training, i.e., adding a lightweight visual adaptation module to the pre-trained model (such as connecting an image feature projection layer before the language model or training a small-scale cross-modal Transformer), training only this adapter part to achieve the model's absorption of new visual information. This method is similar to schemes such as MiniGPT-4, which connects the visual encoding output to the language model through a trainable layer, significantly improving training efficiency. This system allows users to choose from full model fine-tuning, LoRA fine-tuning, or adapter fine-tuning as needed through configuration support, balancing fine-tuning effectiveness and resource consumption. Furthermore, after training, the system can uniformly manage and compare model versions generated by different fine-tuning strategies, facilitating the selection of the optimal solution for practical business applications. Finally, the model training module fully considers heterogeneous hardware deployment strategies, ensuring efficient model training and inference across various computing platforms. The system has designed abstract computing interfaces for CPUs, GPUs, and various AI acceleration chips (such as TPUs, ASICs, and domestically developed chips).During training scheduling, suitable hardware resources are dynamically matched based on the model size and the selected fine-tuning method. For example, full training of large-parameter models prioritizes multi-GPU parallel computing, combined with distributed training strategies to shorten the training time. Lightweight fine-tuning can be completed in a single-GPU or even CPU environment, facilitating deployment on ordinary servers. For scenarios requiring cross-hardware collaboration (such as using GPUs for matrix operations and CPUs for data preprocessing), the system also features optimized pipeline arrangements to reduce data transfer overhead between different devices. Through support for heterogeneous hardware, models trained by this system can be flexibly deployed on cloud clusters or local servers, and even ported to embedded AI devices for inference, maximizing resource utilization and deployment flexibility while meeting performance requirements.

[0112] Figure 3 A flowchart illustrating a financial data processing apparatus provided in an embodiment of this application. Figure 3 As shown, the device includes: an acquisition module 301 for acquiring financial images; a generation module 302 for detecting the financial images using a target detection algorithm, identifying financial image regions, and annotating key elements within the financial image regions to generate an image annotation file; a matching module 303 for performing cross-modal feature matching on the image annotation file based on preset financial image classification labels using a cross-modal model to obtain classification results and classification confidence; and a correction module 304 for correcting the classification results and classification confidence by associating and fusing the key elements with the classification results through a data interface when the classification confidence is lower than a preset threshold.

[0113] Example 2: Embodiments of this application also provide a computer device, including: a memory storing an executable program; and a processor for running the program, wherein the program executes the methods in various embodiments of the present invention during runtime.

[0114] The aforementioned memory can refer to devices inside a computer used to store data and programs, including RAM, hard disks, etc. RAM can be used to temporarily store running programs and data, while hard disks can be used to store programs and data long-term. Memory enables the computer to read and write data and execute programs. The aforementioned processor is responsible for executing instructions in computer programs and performing data processing. It can also be responsible for controlling and executing various operations, including arithmetic operations, logical operations, and data transmission.

[0115] Example 3: Embodiments of this application also provide a computer-readable storage medium including a stored executable program, wherein, when the executable program is running, it controls the device where the computer-readable storage medium is located to perform the methods of various embodiments of the present invention.

[0116] The aforementioned computer storage media can refer to the media used in computer memory to store certain discontinuous physical quantities. Computer storage media mainly include semiconductors, magnetic cores, magnetic drums, magnetic tapes, laser discs, etc. Computer-readable storage media include stored programs, which can be a set of instructions that a computer can recognize and execute, running on an electronic computer to meet certain information needs.

[0117] Example 4: Embodiments of this application also provide a computer program product, including a computer program that, when executed by a processor, implements the methods of various embodiments of the present invention.

[0118] The aforementioned computer program products can refer to software programs that have been written, tested, and released, and can run on computers or other devices. Computer program products can include application programs, operating systems, utility software, etc., used to achieve specific functions or solve specific problems.

[0119] Example 5: Embodiments of this application also provide a computer program product, including a non-volatile computer-readable storage medium for storing a computer program that, when executed by a processor, implements the methods in various embodiments of the present invention.

[0120] The aforementioned non-volatile computer-readable storage medium can refer to a medium for storing data. Non-volatile computer-readable storage media can retain data without loss when power is off and can be used to store long-term data, such as operating systems, applications, and user files. Non-volatile storage media can include hard disk drives, solid-state drives, optical disks, and flash memory storage devices, etc.

[0121] Example 6: Embodiments of this application also provide a computer program that, when executed by a processor, implements the methods described in the various embodiments of the present invention.

[0122] The aforementioned computer program can refer to a set of instructions used to tell the computer to perform specific tasks or operations. Computer programs can be written by programmers using specific programming languages ​​and can include algorithms, data structures, logic, and control flow. Computer programs can be used for a variety of purposes, including application software, operating systems, etc.

[0123] In the above embodiments of the present invention, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0124] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.

[0125] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0126] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0127] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0128] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A financial data processing method, characterized in that, include: Acquire financial images; The financial image is detected using an object detection algorithm, the financial image region is identified, and key elements within the financial image region are labeled to generate an image labeling file. Based on preset financial image classification labels, a cross-modal model is used to perform cross-modal feature matching on the image annotation file to obtain classification results and classification confidence. When the classification confidence level is lower than a preset threshold, the key elements are correlated and fused with the classification result through a data interface to correct the classification result and the classification confidence level, specifically including: Obtain the key elements, classification results, and classification confidence scores; Iterate through the classification results and query the preset financial chart knowledge base to obtain the required element set and excluded element set corresponding to each classification result; Calculate the matching degree between the key element and the knowledge base. The matching degree includes the number of supporting evidence and the number of opposing evidence. The number of supporting evidence is the number of elements that match the key element with the required element set, and the number of opposing evidence is the number of elements that fall into the excluded element set. The classification result and classification confidence are adjusted based on the matching degree.

2. The financial data processing method according to claim 1, characterized in that, The acquisition of financial images includes: Record detailed source information for financial images, including at least one of the following: the issuing institution, report name, and publication date of the financial research report; the disclosing entity, year, and disclosure platform link of the corporate annual report; the name, version number, and open-source license type of the open-source dataset; and the domain name, data crawling time, and corresponding financial product code of the financial website. For detailed source information from different sources, a differentiated permission verification method is used for verification. If the detailed source information is compliant, the financial image is acquired and a copyright registration permission certificate ledger is generated. The permission certificate ledger includes a unique data identifier, source details, permission certificate document, registration time, and registrant.

3. The financial data processing method according to claim 2, characterized in that, The financial images include at least one of the following: line chart, bar chart, pie chart, financial relationship diagram, financial statement, supporting data table, official seal image, and candlestick chart.

4. The financial data processing method according to claim 1, characterized in that, Also includes: Based on the business objectives of each sub-scenario in the front-end, middle-end, and back-end, and according to the principle of difficulty stratification, scenario-specific prompt word templates are pre-configured. The scenario-specific prompt word templates include three elements: data type, task objective, and business constraints, which are used to limit the content generated by the large model to fit the financial logic. Financial images and corresponding text data are input into a visual-text collaborative reasoning model along with scene-specific prompt word templates to generate question-answer pairs, while data scale is ensured through quantity control. The question-answer pairs are classified and quality-screened using a financial semantic classification model to form a multimodal dataset covering multiple sub-scenarios. The quality screening includes: the question must contain financial terminology corresponding to the sub-scenarios, and the answer must have a unique and verifiable reasoning.

5. The financial data processing method according to claim 4, characterized in that, The process involves inputting financial images and corresponding text data, along with scene-specific prompt word templates, into a visual-text collaborative reasoning model to generate question-answer pairs. Simultaneously, data scale is ensured through quantity control, including: The corresponding text data is matched to the corresponding sub-scene template, and specific data information is filled into the prompt words according to the corresponding sub-scene template. The financial image is converted into a model-compatible input format to form an input package of visual data and structured prompt words. The input packet is fed into the visual-text collaborative reasoning model to obtain question-answer pairs; Determine whether the number of question-answer pairs for each sub-scenario meets the preset value. If not, regenerate question-answer pairs by supplementing the same type of input packets or adjusting the prompt word template until the requirements are met.

6. The financial data processing method according to claim 4, characterized in that, Also includes: The visual-text collaborative reasoning model is used to automatically score the multimodal dataset in multiple dimensions, retaining question-answer pairs that simultaneously meet the threshold requirements of five dimensions: image information density, QA semantic validity, data diversity, objectivity, and computational complexity. The filtered question-and-answer pairs and corresponding financial images are verified based on manual annotation methods and preset first data review standards. The verified question-and-answer pairs are decided based on the unanimous expert voting mechanism. When all experts unanimously determine that the data meets the preset second data review standard, the verified question-and-answer pairs are included in the final dataset.

7. The financial data processing method according to claim 1, characterized in that, The acquisition of financial images also includes: Four types of environmental disturbance data were generated based on financial images to test the robustness of the multimodal large language model under non-ideal visual conditions. The four types of environmental disturbance data include: Key information occlusion: Perform partial occlusion or blurring on the core decision information area in the financial image, with the range controlled within the first threshold range of the core decision information area, simulating scenarios such as document folding, stain coverage, or incomplete scanning. Redundant image interference: Visually identical but business-irrelevant image elements are superimposed on financial images, and the second threshold transparency is used to cover them, simulating scenarios of multiple reports on the same page or charts being mistakenly inserted. Missing information: Removed parts of the financial image related to the question-answering task; Irrelevant information overlay: Add text or graphic elements unrelated to the task to the edge of the financial image, occupying the third threshold visual space, to simulate handwritten annotations, advertising watermarks, or irrelevant text annotation scenarios.

8. A financial data processing device, characterized in that, include: The acquisition module is used to acquire financial images; The generation module is used to detect the financial image according to the target detection algorithm, identify the financial image region, and annotate the key elements in the financial image region to generate an image annotation file; The matching module is used to perform cross-modal feature matching on the image annotation file based on the preset financial image classification labels and using a cross-modal model to obtain the classification result and classification confidence. The correction module is used to correct the classification result and classification confidence when the classification confidence is lower than a preset threshold by associating and fusing the key elements with the classification result through a data interface. Specifically, it includes: Obtain the key elements, classification results, and classification confidence scores; Iterate through the classification results and query the preset financial chart knowledge base to obtain the required element set and excluded element set corresponding to each classification result; Calculate the matching degree between the key element and the knowledge base. The matching degree includes the number of supporting evidence and the number of opposing evidence. The number of supporting evidence is the number of elements that match the key element with the required element set, and the number of opposing evidence is the number of elements that fall into the excluded element set. The classification result and classification confidence are adjusted based on the matching degree.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the financial data processing method as described in any one of claims 1-7.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the financial data processing method as described in any one of claims 1-7.