Multi-modal financial information auditing method and device, equipment and medium

By preprocessing and intelligently reviewing multimodal financial information, the problem of low efficiency in traditional review methods is solved, achieving full coverage of multimodal information and efficient compliance review. It also supports flexible adaptation of dynamic rules, improving review efficiency and accuracy.

CN121961445APending Publication Date: 2026-05-01珠海盈米基金销售有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
珠海盈米基金销售有限公司
Filing Date
2025-12-09
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Traditional financial information review methods are inefficient, lack comprehensive rule coverage, have insufficient multimodal content recognition capabilities, and face difficulties in dynamically updating compliance rules.

Method used

By acquiring and preprocessing multimodal target information, data to be reviewed is generated, compliance review items are identified, and intelligent review tasks based on dynamic rules are invoked. Image text information is extracted by combining optical character recognition and image recognition models, semantic analysis is performed using language function models, accuracy verification is performed by calling financial data query tools, and rule prompts are optimized to improve review efficiency and accuracy.

Benefits of technology

It enables simultaneous recognition of multimodal information, ensuring that the audit covers all compliance risk points in text and images, improving the targeting and efficiency of the audit, supporting the flexible adaptation and efficient execution of dynamic compliance rules, and reducing manual inspection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961445A_ABST
    Figure CN121961445A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-mode financial information auditing method and device, equipment and a medium. The multi-modal financial information auditing method comprises the following steps: acquiring multi-modal target information, and preprocessing the multi-modal target information to generate to-be-audited data; identifying compliance auditing items associated with the to-be-audited data; and calling an intelligent auditing task based on a dynamic rule corresponding to each compliance auditing item, and performing compliance auditing on the to-be-audited data. By means of the mode, the efficiency advantages of the intelligent model in semantic comprehension and multi-modal processing can be exerted, manual check one by one is replaced, the auditing efficiency is improved, flexible adaptation and efficient execution of auditing rules are achieved, and the dynamic compliance requirement is met.
Need to check novelty before this filing date? Find Prior Art

Description

A method, apparatus, equipment and medium for verifying multimodal financial information Technical Field

[0001] This application relates to the field of big data analytics technology, and in particular to a method, apparatus, device, and medium for verifying multimodal financial information. Background Technology

[0002] Traditional review methods, faced with the ever-increasing volume of promotional materials and inspection content, require substantial manpower, significantly increasing platform operating costs. Furthermore, as the workload of manual review increases, review efficiency declines. Therefore, existing technologies for financial information review suffer from low efficiency, incomplete rule coverage, insufficient multimodal content recognition capabilities, and difficulties in dynamically updating compliance rules. Summary of the Invention

[0003] This application mainly provides a method, apparatus, equipment, and medium for reviewing multimodal financial information, in order to solve the problem of poor financial information review capabilities.

[0004] To address the aforementioned technical problems, this application adopts the following technical solution: providing a method for reviewing multimodal financial information, comprising: acquiring multimodal target information and preprocessing the multimodal target information to generate data to be reviewed; identifying compliance review items associated with the data to be reviewed; and respectively invoking intelligent review tasks based on dynamic rules corresponding to each compliance review item to conduct compliance review of the data to be reviewed.

[0005] In some embodiments, the step of acquiring multimodal target information and preprocessing the multimodal target information to generate data to be reviewed includes: receiving the multimodal target information uploaded manually or extracted from network inspections; extracting text information within images from the multimodal target information using optical character recognition and image recognition models; and adding the text information within images to the text information of the multimodal target information to form data to be reviewed in single-text form.

[0006] In some embodiments, the step of calling the dynamic rule-based audit tasks corresponding to each audit item to perform compliance audits on the data to be audited includes: dynamically adding rule prompts to each audit task to call and command the language function model; and performing semantic analysis through the language function model corresponding to each compliance audit item to obtain the compliance audit results.

[0007] In some embodiments, after performing semantic analysis through the language function model corresponding to each compliance audit item, the method further includes: in response to the language function model identifying a financial data item in the data to be audited, invoking the financial data query tool corresponding to each financial data item; and verifying the accuracy of the financial data through each financial data query tool.

[0008] In some embodiments, the language function model is a large language model.

[0009] In some embodiments, after the compliance audit of the data to be audited is performed, the method further includes: obtaining feedback information on the compliance audit results of manual review, and optimizing the description method of the rule prompt words.

[0010] In some embodiments, after identifying the compliance audit projects associated with the data to be audited, the method further includes: associating each compliance audit project and each audit task through the Dify workflow based on the number and type of the compliance audit projects associated with the data to be audited.

[0011] To address the aforementioned technical problems, another technical solution adopted in this application is: providing a multimodal financial information review device, comprising: a preprocessing module for acquiring multimodal target information and preprocessing the multimodal target information to generate data to be reviewed; an identification module for identifying compliance review items associated with the data to be reviewed; and an review module for respectively calling the intelligent review tasks based on dynamic rules corresponding to each compliance review item to conduct compliance review of the data to be reviewed.

[0012] This application also provides a computer device, the computer device comprising: a memory and at least one processor, the memory storing instructions; the at least one processor invokes the instructions in the memory to cause the computer device to perform the multimodal financial information verification method as described above.

[0013] This application also provides a computer-readable storage medium storing instructions that, when executed by a processor, implement the multimodal financial information verification method described above.

[0014] The beneficial effects of this application are as follows: Unlike existing technologies, this application discloses a method, apparatus, device, and medium for reviewing multimodal financial information. By acquiring multimodal target information and preprocessing it, data to be reviewed is generated. By collecting multiple types of target information, such as text and images, the limitations of a single data format are broken, covering review needs in complex scenarios. Simultaneous recognition of multimodal information is achieved, ensuring that the review covers all compliance risk points in text and images, avoiding the omission of image violations caused by traditional text-only reviews. By identifying the compliance review items associated with the data to be reviewed; and by initially identifying the types of compliance points in the content based on keywords, the compliance rules to be reviewed are accurately located, avoiding indiscriminate full-scale checks and improving the targeting and efficiency of the review. By separately calling the intelligent review tasks based on dynamic rules corresponding to each compliance review item, compliance reviews are performed on the data to be reviewed. Complex review tasks are broken down into standardized, configurable intelligent workflows, leveraging the efficiency advantages of AI in semantic understanding and multimodal processing, replacing manual item-by-item checks, improving review efficiency, and achieving flexible adaptation and efficient execution of review rules to meet dynamic compliance requirements. Attached Figure Description

[0015] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Among them: Figure 1 is a flowchart of an embodiment of the multimodal financial information review method provided by this application; Figure 2 is a flowchart of an embodiment of method step 100 as shown in Figure 1; Figure 3 is a flowchart of an embodiment of method step 300 as shown in Figure 1; Figure 4 is a flowchart of an embodiment of method step 320 as shown in Figure 3; Figure 5 is a multimodal financial information review device provided by this application; Figure 6 is a structural schematic diagram of an embodiment of computer equipment in this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0017] The terms "first," "second," and "third" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first," "second," or "third" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices.

[0018] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0019] Referring to Figure 1, Figure 1 is a flowchart of an embodiment of the multimodal financial information review method provided in this application. The multimodal financial information review method includes the following steps: 100: Obtain multimodal target information and preprocess the multimodal target information to generate data to be reviewed.

[0020] Multimodal target information is obtained through business system interfaces, covering various formats such as text files and image files. After preprocessing, the multimodal target information is encapsulated according to fixed fields, and the output is standardized data that can be directly read by subsequent review tasks.

[0021] Among them, multimodal financial information includes various types of content to be reviewed, such as promotional articles containing text paragraphs and promotional images uploaded manually, and WeChat official account articles crawled by web crawlers, covering all information carriers in both proactive submission and passive inspection scenarios.

[0022] The data to be reviewed is a unified format data formed after preprocessing multimodal target information. It contains text content and key information extracted from images and is the direct input object for subsequent compliance review.

[0023] Traditional review processes only text content, potentially missing hidden violations within images, such as superlative terms like "best product" in the corner of an image. By using OCR and image recognition technology, the preprocessing stage merges text within the image with the main text content, ensuring that the data to be reviewed includes all potential risk points.

[0024] For example, in a fund promotion article, the text did not mention any prohibited words, but the accompanying image contained the handwritten text "100% return". The system used OCR to extract the text and included it in the data to be reviewed, thus avoiding missed detection.

[0025] Further, referring to Figure 2, step 100 further includes the following steps: 110: receiving multimodal target information uploaded manually or extracted from network inspections.

[0026] Materials are received through business system interfaces, such as promotional materials for financial products and graphic advertisements, or publicly available content is passively extracted through network inspection tools, such as social media posts and web page promotion information, covering various information carriers such as text and images.

[0027] 120: Extract text information from images in multimodal target information using optical character recognition and image recognition models.

[0028] The Optical Character Recognition (OCR) model is used to extract text from images in multimodal information, such as recognizing the phrase "guaranteed principal and interest" in a poster. At the same time, the image recognition model is combined to locate the text region in the image, such as eliminating interference from complex backgrounds and focusing on the main text, to ensure the accuracy of text extraction.

[0029] Among them, the optical character recognition model uses deep learning algorithms to convert printed / handwritten text in images into computer-recognizable text data, thus overcoming the limitation that text in images is unreadable and converting hidden text in images into verifiable text, avoiding the omission of image violations caused by traditional text-only verification.

[0030] Image recognition models are used to identify key features in images, assisting optical character recognition (OCR) models in locating text and improving text extraction accuracy. Image recognition models help reduce the invalid recognition of non-text areas by OCR models, ensuring that the extracted text information within the image focuses more on the core content and reducing noise interference during subsequent review.

[0031] Text information within images is extracted using optical character recognition (OCR) and image recognition models. This transforms text in images that would otherwise be unverifiable into searchable and analyzable text data, bringing compliance risks hidden within images into the scope of review.

[0032] 130: Add the text information from the image to the text information of the multimodal target information to form a single text format of the data to be reviewed.

[0033] The text information extracted from the image in step 120 is merged with the original text information, redundant formats are removed, and the data is packaged in a unified structure to generate single-text data to be reviewed.

[0034] The single-text data to be reviewed is a unified structured text data formed by integrating the original text information and the "text information within the image". It is used to avoid the blind spots in the review caused by the separation of text and image information. The standardized data format can be directly called by the recognition module of the subsequent compliance review project without repeatedly parsing the multimodal original information, thus shortening the review process time.

[0035] By extracting text from images in multimodal information and merging it into single text data, the system achieves full integration of text and image information. This addresses the pain point of traditional auditing where text and images cannot be effectively identified, ensuring that the data to be audited covers all potential compliance risks in text and images, and laying a data foundation for subsequent accurate and efficient compliance audits.

[0036] 200: Identifies compliance audit projects associated with the data to be audited.

[0037] The system performs keyword and semantic analysis on the text-based data to be reviewed, identifying compliance risk-related elements within the content. When a specific financial product code is detected, it is marked as a financial data feature. Based on the extracted features, the system automatically associates them with pre-set compliance review items. Finally, it generates a list of compliance review items that need to be performed on the data to be reviewed, serving as a direct basis for subsequent intelligent review tasks.

[0038] Compliance audit projects are pre-categorized into specific compliance inspection categories based on the content characteristics of the data to be audited. Each category corresponds to a specific set of compliance rules, such as "Historical Return Promotion Standards", "Restrictions on the Expression of Expected Rate of Return", and "Verification of the Accuracy of Financial Data".

[0039] For example, when the content involves fund performance data, "historical return promotion" and "data authenticity" are compliance review items that need to be associated; when the data to be reviewed is a "fund product prospectus" and includes "historical net asset value curve", the system only calls the two review items "historical return promotion specifications" and "data authenticity verification", without checking irrelevant items such as "concentrated marketing time limit" and "compliance of advertising channels"; when the content is plain text promotional copy, only basic review items such as "use of extreme words" and "false promises" are triggered to avoid wasting resources.

[0040] For example, if a bank's promotional copy for a wealth management product mentions both "expected annualized return of 5.8%" and "number one in sales nationwide," the system will associate it with "expected rate of return promotion restrictions" and "limited word usage guidelines," respectively, while other unrelated items will not be triggered.

[0041] By clearly defining the associated audit projects, complex compliance checks can be broken down into specialized and modular sub-tasks. Subsequent steps can call the corresponding "dynamic rule intelligent audit task" for each project, achieving "precise tool matching" and "focused rule iteration".

[0042] 300: Invoke the intelligent audit tasks based on dynamic rules corresponding to each compliance audit project to conduct compliance audits on the data to be audited.

[0043] Based on the project list output during the compliance audit project identification phase, the corresponding intelligent audit task is retrieved from the system rule base. During task execution, the latest compliance rules for that audit project are injected into the audit model as a prompt, including rule description, violation examples, and output format. Dedicated tools are used to enhance audit capabilities for different audit projects. Each task independently outputs a compliance judgment, and the results are ultimately aggregated to form a comprehensive compliance report of the data to be audited.

[0044] Intelligent audit tasks based on dynamic rules are automated inspection processes that combine dynamically configurable compliance rules with intelligent tools to perform for specific compliance audit projects. Their core features are that rules can be updated without code modification, and the audit logic is deeply integrated with specific business scenarios.

[0045] By converting compliance rules into editable prompts instead of hardcoding them into the system's underlying layer, when regulations are updated, only the prompts for the corresponding audit tasks need to be modified, without needing to restart the system or iterate development.

[0046] Further, referring to Figure 3, step 300 further includes the following steps: 310: Dynamically add rule prompts in each audit task to call and command the language function model.

[0047] For identified compliance audit projects, the latest compliance rules corresponding to the project are converted into rule prompt words and dynamically injected into the language function model to enable command-based calls to the model.

[0048] For example, when reviewing projects related to historical earnings promotion, the rule prompts will explicitly include specific instructions such as "Check whether it is labeled 'Past performance does not guarantee future results'" and "Do not use extreme words such as 'best' or 'highest' to describe earnings," to ensure that the model performs the review task in accordance with compliance requirements.

[0049] Among them, rule prompts are instructions that translate the specific rule requirements of compliance audit projects into language functional models that can understand, including rule descriptions, violation examples, judgment criteria, and other elements.

[0050] For example, the rule prompt for the "Restriction on Expected Rate of Return Promotion" project might be: "Please check if the text contains expressions such as 'expected return' or 'projected return.' If so, confirm whether it is marked 'expected return is not actual return'; if it is not marked or promises to guarantee principal and return, it is considered a violation." The Language Function Model (LFM) is an intelligent model with natural language understanding and reasoning capabilities. Its core function is to perform semantic analysis on text based on rule prompts to identify compliance risk points. Its key feature is support for dynamic instruction input, allowing adjustment of the review logic through prompts without modifying the model's underlying code.

[0051] Optionally, the language function model is a large language model.

[0052] Rule prompts are stored in text format. When regulatory policies are updated, only the prompts for the corresponding audit projects need to be modified, without retraining the model or developing the system. Traditional keyword matching struggles to handle semantically complex expressions, while language function models, through contextual semantic understanding, can accurately determine whether an expression implies a violation. Compliance audit results include the location of the violation fragment and associated rules, allowing manual review to focus directly on risk points without needing to read the entire document.

[0053] 320: Semantic analysis is performed using the language function model corresponding to each compliance audit project to obtain the compliance audit results.

[0054] The language function model performs deep semantic analysis on the data to be reviewed based on the injected rule prompts. Specifically, it parses the key expressions in the content, compares the parsing results with the compliance rules in the prompts, and generates a structured compliance review result that includes "compliance / violation determination," "violation fragment location," and "risk level."

[0055] A single language function model can load rule prompts for multiple compliance audit projects simultaneously. Multi-dimensional checks can be completed by performing a single semantic analysis on the data to be audited, avoiding redundant calculations of scanning text separately for each rule in the traditional process.

[0056] Optionally, after performing semantic analysis through the language function model corresponding to each compliance audit item, the following steps are also included: 321: In response to the language function model identifying financial data items in the data to be audited, the financial data query tool corresponding to each financial data item is invoked.

[0057] When the language function model identifies a financial data item in the single-text data to be reviewed, the system automatically triggers the corresponding data verification process and calls the financial data query tool bound to that financial data item through the interface.

[0058] 322: Verify the accuracy of financial data through various financial data query tools.

[0059] The financial data query tool connects to authoritative data sources to verify financial data items in the data to be reviewed in real time.

[0060] By identifying financial data items and using query tools to verify their accuracy, the core solution addresses the compliance risk of discrepancies between financial promotional data and official filings, ensuring the authenticity and reliability of quantitative information in the data to be audited.

[0061] Optionally, after conducting a compliance review of the data to be reviewed, the process also includes: obtaining feedback information on the results of the manual review of compliance, and optimizing the description of rule prompts.

[0062] Human reviewers verify the compliance audit results output by the language function model, and mark cases of model misjudgment and the reasons for the misjudgment.

[0063] The system performs attribution analysis on manually submitted cases to pinpoint whether the root cause of the problem lies in inaccurate descriptions of the rule prompts. Based on the analysis results, the rule prompts are adjusted accordingly, including supplementing scenario conditions, refining judgment criteria, and clarifying exceptions.

[0064] The optimized rule prompts are updated to the system. The optimization effect is verified by backtesting historical misjudgment cases and reviewing new data. If there are still deviations, the above process is repeated to form a feedback loop.

[0065] By leveraging human feedback to expose vague wording or missing scenarios in rule prompts, optimizations can clarify rule boundaries. Cases missed in manual feedback often correspond to specific scenarios not covered by the rule prompts; optimization can incorporate these scenarios into the rule logic. Traditional rule updates rely on changes in regulatory policies, while optimizing prompts through human feedback enables rapid response to borderline cases that are not explicitly regulated but pose risks.

[0066] Optionally, after identifying the compliance audit projects associated with the data to be audited, the process also includes: associating each compliance audit project and each audit task through the Dify workflow based on the number and type of compliance audit projects associated with the data to be audited.

[0067] The Dify workflow first extracts the metadata of the data to be audited, identifies the compliance audit projects that need to be covered, and retrieves the corresponding audit tasks from the system's preset project task association rule library based on the type of audit project. A task queue is automatically generated according to the number of projects. If multiple projects are involved, the Dify workflow starts the associated audit tasks sequentially or in parallel according to preset priorities, and records the correspondence between projects and tasks to ensure that each audit project has a unique task to execute.

[0068] Among them, the Dify workflow is an automated workflow tool based on visual process orchestration capabilities, supporting rule configuration and task scheduling. Its core function is to automatically trigger and associate downstream tasks according to the attributes of input data through preset logic, realizing end-to-end process automation among data, projects, and tasks.

[0069] When the data to be audited is associated with multiple compliance audit projects, the Dify workflow can call the corresponding tasks in parallel instead of processing them serially in order. It replaces the traditional process of first manually judging the data type and then manually starting the corresponding task.

[0070] Optionally, semantic analysis is performed through the language function models corresponding to each compliance audit project, including: constructing a variant word library based on the pinyin self-similarity algorithm. Detecting variant expressions of illegal words in the data to be audited based on the variant word library through a pre-trained adversarial detection model in the financial field.

[0071] Sort out the financial illegal words clearly prohibited by regulations and high-frequency avoidance expressions; calculate the pinyin features of the basic words through the pinyin similarity algorithm, and generate variants such as homophone replacement, pinyin abbreviation, character replacement with similar shapes, and mixed coding; manually screen the generated variants and store them classified by "illegal risk level", "variant type", etc., to form a dynamically updatable variant word library.

[0072] The data to be audited is input into the model after text preprocessing; the model calls the variant word library and identifies variant expressions through semantic vector comparison; combines the context semantic analysis of the actual intention of the variant expressions, and outputs the determination of "violation / suspected violation" and the source of the variant words.

[0073] Among them, the variant word library is a database that stores financial illegal words and their variant forms generated manually / algorithmically. The core elements include "basic illegal words", "variant word types", "similarity scores", "risk levels", etc.

[0074] The adversarial detection model in the financial field is a text classification / named entity recognition model pre-trained based on financial industry corpora, which is specially optimized for illegal word variant avoidance behaviors such as homophone replacement and character confusion. Its core function is to identify variant expressions through semantic understanding rather than character matching.

[0075] The variant expressions of illegal words are forms that transform the basic illegal words through means such as character replacement, coding mixing, and semantic hinting to avoid traditional keyword detection. Common types include: homophone / near-homophone replacement: such as "steady profit" → "steady turn", "epidemic profit"; pinyin / letter mixing: such as "annualized return" → "nh return", "annualized shouyi"; character replacement with similar shapes / variant characters: such as "capital preservation" → "capital preservation", "capital wood"; semantic splitting hinting: such as "principal safety" → "the principal will not be less", "the invested money is guaranteed".

[0076] Traditional keyword detection relies on fixed character matching, which is easily bypassed by methods such as "homophonic substitution" and "similar-looking characters". However, variant word libraries generated based on pinyin similarity algorithms can cover common evasion methods. Combined with the semantic understanding capabilities of adversarial detection models, the false negative rate can be reduced.

[0077] Meanwhile, when new illegal variants appear, the variant vocabulary only needs to be updated through the pinyin similarity algorithm, without modifying the model structure. The model can then identify new variants through the dynamic vocabulary loading function, shortening the response cycle.

[0078] Optionally, rule prompts can be dynamically added to the review task to invoke and command language function models, including: calculating the complexity score of the review task based on the detection results of variant expressions of violating words in the data to be reviewed; and dynamically modifying the rule prompts based on the complexity score to invoke and command language function models of different scales.

[0079] First, the data to be reviewed is scanned using a financial domain adversarial detection model, outputting the core features of variant expressions. Based on the detection results, a complexity score, core scoring dimensions, and weights are generated using a preset algorithm. The level of detail of the rule prompts and the detection strategy are adjusted according to the complexity score to match the prompts with the task difficulty. Based on the modified rule prompts, the system automatically matches a language function model of the corresponding magnitude.

[0080] Complexity score is a quantitative indicator calculated based on the number, concealment, and type diversity of variant expressions of illegal words in the data to be reviewed. It is used to measure the technical difficulty of the review task and its core function is to provide a basis for decision-making on model resource allocation.

[0081] Optionally, the complexity score can be set to a range of 0-10. A higher score indicates a stronger intention to circumvent the system, requiring more sophisticated detection capabilities.

[0082] Among them, the different scales of the language function model are language model levels divided according to parameter size, semantic understanding ability, and computational resource consumption, and each level of model is optimized for review tasks of different complexities.

[0083] Specifically, the scale of language function models includes lightweight, medium-weight, and heavyweight.

[0084] To avoid wasting resources by using heavyweight models to handle all tasks, for example, low-complexity tasks are handled by lightweight models to reduce computational costs, while high-complexity tasks call heavyweight models to ensure detection accuracy.

[0085] The above describes the method for reviewing multimodal financial information in the embodiments of the present invention. The following describes the device for reviewing multimodal financial information in the embodiments of the present invention. Please refer to Figure 5. One embodiment of the device for reviewing multimodal financial information in the embodiments of the present invention includes: a preprocessing module 410, used to acquire multimodal target information and preprocess the multimodal target information to generate data to be reviewed.

[0086] The identification module 420 is used to identify compliance audit items associated with the data to be audited.

[0087] The audit module 430 is used to call the intelligent audit tasks based on dynamic rules corresponding to each compliance audit project to conduct compliance audits on the data to be audited.

[0088] Figure 5 above describes the feature extraction device in this embodiment of the invention from the perspective of modular functional entities. The following describes the computer device in this embodiment of the invention from the perspective of hardware processing.

[0089] Figure 6 is a schematic diagram of a computer device 500 provided in an embodiment of the present invention. The computer device 500 can vary significantly due to different configurations or performance characteristics. It may include one or more central processing units (CPUs) 510 (e.g., one or more processors) and a memory 520, and one or more storage media 530 (e.g., one or more mass storage devices) for storing application programs 533 or data 532. The memory 520 and storage media 530 may be temporary or persistent storage. The program stored in the storage media 530 may include one or more modules (not shown in the figure), each module including a series of instruction operations on the computer device 500. Furthermore, the processor 510 may be configured to communicate with the storage media 530 and execute the series of instruction operations in the storage media 530 on the computer device 500.

[0090] Computer device 500 may also include one or more power supplies 540, one or more wired or wireless network interfaces 550, one or more input / output interfaces 560, and / or one or more operating systems 531, such as Windows Server, MacOSX, Unix, Linux, FreeBSD, etc. Those skilled in the art will understand that the computer device structure shown in FIG6 does not constitute a limitation on the computer device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0091] The present invention also provides a computer device, the computer device including a memory and a processor, the memory storing computer-readable instructions, which, when executed by the processor, cause the processor to perform the steps of the multimodal financial information review method described in the above embodiments.

[0092] The present invention also provides a computer-readable storage medium, which can be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium, wherein the computer-readable storage medium stores instructions that, when the instructions are executed on a computer, cause the computer to perform the steps of the multimodal financial information verification method.

[0093] Unlike existing technologies, this application employs a manual feedback-driven rule suggestion optimization mechanism. It manually reviews and labels misjudged cases in the model, analyzes the reasons for misjudgments, supplements scenario conditions, and refines judgment criteria, forming a feedback optimization loop. This effectively improves the semantic accuracy of rule suggestions, reduces model comprehension bias, and enhances the rule's coverage of specific scenarios. This application utilizes the Dify workflow to automate the association between review projects and tasks, reducing manual intervention costs, supporting parallel review of multiple projects, and shortening the overall review cycle. Furthermore, this application provides a variant word library construction based on pinyin similarity algorithms and an adversarial detection model in the financial field. This breaks through the limitations of traditional keyword detection, improves the recognition rate of violation avoidance behaviors, effectively reduces the risk of misjudgment of professional terms, and facilitates dynamic updates of the variant word library. Through these technical means, facing new violation methods, the application rapidly generates new variants through the variant word library, supplements rule suggestions with manual feedback, and triggers advanced model detection through complexity scoring, significantly shortening the response cycle.

[0094] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on its differences from other embodiments. In particular, the storage medium embodiments and computer device embodiments are basically similar to the method embodiments, so the descriptions are relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0095] This application can be used in a wide range of general-purpose or specialized in-vehicle computing system environments or configurations. Examples include: personal computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, network PCs, minicomputers, and distributed computing environments including any of the above systems or devices.

[0096] In the several embodiments provided in this application, it should be understood that the disclosed methods and devices can be implemented in other ways. For example, the device embodiments described above are merely illustrative; multiple units or components may be combined or integrated into another system, or some features may be omitted or not performed.

[0097] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0098] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0099] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for verifying multimodal financial information, characterized in that, include: Acquire multimodal target information and preprocess the multimodal target information to generate data to be reviewed; Identify the compliance audit projects associated with the data to be audited; Each compliance audit project is invoked to perform a compliance audit on the data to be audited, based on dynamic rules.

2. The method for verifying multimodal financial information according to claim 1, characterized in that, The process of acquiring multimodal target information and preprocessing the multimodal target information to generate data to be reviewed includes: receiving the multimodal target information uploaded manually or extracted from network inspections; extracting text information within images from the multimodal target information using optical character recognition and image recognition models; and adding the text information within images to the text information of the multimodal target information to form data to be reviewed in single-text format.

3. The method for verifying multimodal financial information according to claim 1, characterized in that, The step of calling the dynamic rule-based audit tasks corresponding to each audit item to perform compliance audits on the data to be audited includes: dynamically adding rule prompts to each audit task to call and command the language function model; and performing semantic analysis through the language function model corresponding to each compliance audit item to obtain the compliance audit results.

4. The method for verifying multimodal financial information according to claim 3, characterized in that, After performing semantic analysis using the language function model corresponding to each compliance audit item, the method further includes: in response to the language function model identifying a financial data item in the data to be audited, calling the financial data query tool corresponding to each financial data item; and verifying the accuracy of the financial data using the financial data query tool.

5. The method for verifying multimodal financial information according to claim 4, characterized in that, The language function model is a large language model.

6. The method for verifying multimodal financial information according to claim 3, characterized in that, After conducting a compliance review on the data to be reviewed, the method further includes: obtaining feedback information on the results of the compliance review through manual review, and optimizing the description of the rule prompts.

7. The method for verifying multimodal financial information according to claim 1, characterized in that, After identifying the compliance audit projects associated with the data to be audited, the method further includes: using the Dify workflow to associate each compliance audit project with each audit task based on the number and type of the compliance audit projects associated with the data to be audited.

8. A multimodal financial information verification device, characterized in that, include: The preprocessing module is used to acquire multimodal target information and preprocess the multimodal target information to generate data to be reviewed; The identification module is used to identify the compliance audit items associated with the data to be audited; The audit module is used to call the intelligent audit tasks based on dynamic rules corresponding to each of the compliance audit projects to conduct compliance audits on the data to be audited.

9. A computer device, characterized in that, The computer device includes: a memory and at least one processor, the memory storing instructions; the at least one processor invokes the instructions in the memory to cause the computer device to perform the multimodal financial information auditing method as described in any one of claims 1-7.

10. A computer-readable storage medium storing instructions thereon, characterized in that, When the instruction is executed by the processor, it implements the method for reviewing multimodal financial information as described in any one of claims 1-7.