A method, device and medium for due diligence on a public account

By acquiring multi-source heterogeneous data, extracting text, image, and graph structure features, performing feature fusion, and using a hybrid risk classification model for assessment, a due diligence report is generated. This solves the problems of low efficiency and incomplete risk coverage in existing corporate account investigation methods, and achieves automated and standardized management.

CN122134444APending Publication Date: 2026-06-02天元大数据信用管理有限公司

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
天元大数据信用管理有限公司
Filing Date
2026-01-14
Publication Date
2026-06-02

AI Technical Summary

Technical Problem

Existing methods for investigating corporate accounts are inefficient, lack comprehensive risk coverage, and have highly subjective audit standards, making it difficult to achieve unified and standardized management.

Method used

By acquiring multi-source heterogeneous data, extracting text, image, and graph structure features, performing feature fusion to generate enterprise risk representation vectors, using a hybrid risk classification model for risk assessment, and generating due diligence reports.

Benefits of technology

It improved the efficiency of due diligence, solved the problems of limited risk insight dimensions and highly subjective review standards, and achieved automated and standardized management of due diligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122134444A_ABST
    Figure CN122134444A_ABST
Patent Text Reader

Abstract

The application discloses a kind of public account due diligence method, equipment and medium, it is related to the technical field of financial technology.The method comprises the following steps: obtaining multi-source heterogeneous data related to an enterprise;Feature extraction is performed on the multi-source heterogeneous data to obtain the multi-modal features corresponding to the enterprise;Feature fusion is performed on the vectorized text features, image features and graph structure features to generate the enterprise risk representation vector corresponding to the enterprise;The enterprise risk representation vector is input into a hybrid risk classification model to obtain the risk level of the enterprise output by the hybrid risk classification model;Based on the risk level, a due diligence report corresponding to the enterprise is generated.Through the above method, the problems of low efficiency, incomplete risk coverage and strong subjective judgment of existing investigation methods can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of financial technology, and in particular to a method, device and medium for due diligence of corporate accounts. Background Technology

[0002] Opening corporate accounts is the primary and crucial step for banks to serve corporate clients and establish financial partnerships. The rigor and reliability of its due diligence process directly impacts not only the bank's own credit, operational, and compliance risks, but also its operating cost structure and market reputation. However, the traditional due diligence model, which the banking industry has long relied on in this core process, has revealed numerous insurmountable structural flaws, specifically manifested in the following aspects: First, the operational model is highly dependent on manual labor, with complex processes and significant efficiency bottlenecks. Second, risk insight is limited in scope, lacking in-depth correlation and penetrating identification capabilities. Traditional risk assessment methods typically focus on the company's static information, making it difficult to deeply explore and dynamically analyze the complex and hidden equity networks and control relationships among the company, its legal representative, major shareholders, beneficial owners, and related parties. Third, the audit standards are highly subjective, making it difficult to achieve unified and standardized management. Summary of the Invention

[0003] This application provides a method, device, and medium for due diligence on corporate accounts to address the following technical problems: how to solve the problems of low efficiency, incomplete risk coverage, and strong subjective judgment in existing investigation methods.

[0004] In a first aspect, embodiments of this application provide a due diligence method for corporate accounts. The method includes: acquiring multi-source heterogeneous data related to a company, wherein the company is the applicant company for the corporate account; extracting features from the multi-source heterogeneous data to obtain multimodal features corresponding to the company, wherein the multimodal features include text features, image features, and graph structure features, the graph structure features being used to indicate the equity relationships of the company; performing feature fusion on the vectorized text features, image features, and graph structure features to generate a corporate risk representation vector corresponding to the company; inputting the corporate risk representation vector into a hybrid risk classification model to obtain a risk level corresponding to the company output by the hybrid risk classification model, wherein the hybrid classification model is used to perform risk assessment on the company based on the corporate risk representation vector to obtain the risk level corresponding to the company; and generating a due diligence report corresponding to the company based on the risk level.

[0005] Secondly, embodiments of this application also provide a due diligence device for corporate accounts, the device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a due diligence method for corporate accounts as described in the first aspect above.

[0006] Thirdly, embodiments of this application also provide a computer storage medium storing computer-executable instructions, which, when executed, implement a due diligence method for corporate accounts as described in the first aspect above.

[0007] The due diligence method, equipment, and medium for corporate accounts provided in this application have the following beneficial effects: In this embodiment, multi-source heterogeneous data related to the enterprise can be acquired, and then features are extracted from the multi-source heterogeneous data to obtain multimodal features, including text features, image features, and graph structure features. Next, the vectorized text features, image features, and graph structure features can be fused to generate an enterprise risk representation vector. This enterprise risk representation vector can then be input into a hybrid risk classification model to obtain the enterprise's risk level. Finally, based on the risk level, a due diligence report for the enterprise can be generated. In this way, through automated data collection, feature extraction, feature fusion, and risk assessment, the efficiency of due diligence can be improved. Furthermore, graph structure features can indicate the enterprise's equity relationships. Using graph structure features for risk assessment can solve the problem of a single dimension of risk insight and a lack of deep correlation and penetrating identification capabilities. Moreover, the due diligence report is generated based on the enterprise's risk level output by the hybrid risk classification model, rather than human judgment, which can solve the problem of strong subjectivity in existing methods' review standards and achieve unified and standardized management of due diligence. Attached Figure Description

[0008] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating a due diligence method for corporate accounts provided in this application embodiment; Figure 2 A technical architecture diagram of a due diligence method for corporate accounts provided in this application embodiment; Figure 3 This is a schematic diagram of the internal structure of a due diligence device for corporate accounts provided in an embodiment of this application. Detailed Implementation

[0009] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0010] The main drawbacks of existing methods for investigating corporate accounts are as follows: First, the work mode is highly dependent on manual labor, with complex processes and prominent efficiency bottlenecks.

[0011] The entire due diligence process, from material collection and preliminary verification to data entry and report writing, relies almost entirely on manual operations by account managers. Account managers must expend significant time and effort collecting, organizing, and verifying the massive amounts of paper and electronic materials provided by companies, including but not limited to business licenses, articles of association, equity structure diagrams, financial statements, identification of the actual controller, and proof of business premises. This process is not only tedious and repetitive, but also takes an average of several business days or even weeks per company, severely hindering the scalable expansion of banking business and the improvement of customer experience, while also resulting in high human and time costs.

[0012] Second, risk insight is limited in scope and lacks in-depth correlation and penetrating identification capabilities. Traditional risk assessment methods typically focus on the static information of the enterprise itself, making it difficult to deeply explore and dynamically analyze the complex and hidden equity networks and control relationships between the enterprise, its legal representative, major shareholders, beneficial owners, and related parties. Therefore, traditional methods often fall short in addressing high-risk behaviors such as the true controller concealed through complex equity designs, unfair transactions between related parties, and shell companies and shell financing aimed at circumventing regulations. This results in numerous blind spots in risk insight and prevents truly penetrating supervision.

[0013] Third, the subjective nature of due diligence standards makes it difficult to achieve unified and standardized management. Because due diligence conclusions largely depend on the individual professional experience, risk appetite, and sense of responsibility of the reviewers, significant differences can exist in the review criteria and judgment standards applied by different personnel, and even among the same person handling different cases. This pervasive subjectivity leads to instability and incomparability in risk assessment results, making it difficult for banks to establish and implement a unified, objective, and quantifiable set of risk management standards. This not only poses challenges to internal management but also easily triggers compliance and operational risks.

[0014] In practical applications, a fourth aspect is the lack of effective means to verify the authenticity of information, resulting in weak fraud risk prevention capabilities. Under the traditional model, auditors rely primarily on experience and basic cross-referencing to verify the authenticity of materials submitted by companies, lacking effective technical verification methods. For example, it is difficult to automate and reliably verify the authenticity of key documents such as business licenses and audit reports; information such as the actual operating address and production scale declared by companies often cannot be effectively verified remotely or on-site. This lack of verification capability makes the banking system highly susceptible to being deceived by meticulously crafted false materials.

[0015] While existing technological advancements have yielded some progress, none have systematically resolved the fundamental problems mentioned above. For instance, current OCR-based automated information processing solutions primarily address the digitization and automatic recognition of information on paper documents, improving data entry efficiency to some extent. However, the application of such technologies largely stops at preliminary information identification and entry, failing to delve into in-depth verification and risk correlation analysis. They are unable to effectively identify forged documents processed by technology, and lack the ability to conduct multi-dimensional cross-verification and comprehensive risk assessment of the identified information, thus making it difficult to independently support a rigorous due diligence process.

[0016] This application provides a technical solution for due diligence on corporate accounts. The technical solution proposed in this application will be described in detail below with reference to the accompanying drawings.

[0017] Figure 1 A flowchart illustrating a due diligence method for corporate accounts provided in this application embodiment. Figure 1 As shown in the embodiment of this application, a due diligence method for corporate accounts specifically includes the following steps: Step 101: Obtain multi-source heterogeneous data related to the enterprise.

[0018] The company mentioned here is the company that applied for the corporate account.

[0019] In this embodiment of the application, when conducting due diligence on a corporate bank account, multi-source heterogeneous data of the company applying for the account can be obtained first. In practical applications, multi-source heterogeneous data can include the company's declared data (basic company information, legal representative, etc.), as well as image materials such as business licenses, legal representative's ID cards, photos or videos of the business premises, and lease agreements uploaded by the company, without specific limitations. This can reduce information bias caused by a single data source and lay a solid data foundation for subsequent due diligence.

[0020] Step 102: Extract features from the multi-source heterogeneous data to obtain the multimodal features corresponding to the enterprise.

[0021] The multimodal features include text features, image features, and graph structure features, with the graph structure features used to indicate the equity relationships of the enterprise.

[0022] In this embodiment, feature extraction can be performed on the acquired multi-source heterogeneous data to obtain multimodal features corresponding to the enterprise. Text features, image features, and graph structure features among the multimodal features can reflect enterprise-related information from different perspectives. Text features are key information extracted from the multi-source heterogeneous data that can characterize text semantics or structure, and can be used to indicate the enterprise's text information. Image features are abstract descriptions of various information within images in the multi-source heterogeneous data, used to characterize the image's content, structure, style, and other attributes, and can be used to indicate the enterprise's image information. In practical applications, graph structure features describe the structural attributes of a graph, used to characterize the relationships between nodes and edges, the overall topological structure, and other characteristics of the graph. In this embodiment, the graph structure features obtained by feature extraction from multi-source heterogeneous data can indicate the enterprise's equity relationships. Thus, feature extraction can transform these multi-source heterogeneous data from different sources into multimodal features, comprehensively characterizing the enterprise from multiple dimensions, providing rich details for subsequent due diligence, and facilitating comprehensive consideration of multiple factors for risk assessment of the enterprise's corporate accounts, thus solving the problem of a single dimension of risk insight.

[0023] Step 103: Perform feature fusion on the vectorized text features, image features, and graph structure features to generate the enterprise risk representation vector corresponding to the enterprise.

[0024] In this embodiment, vectorized text features, image features, and graph structure features can be fused. For example, a fusion model based on a cross-modal attention mechanism can be used. This model can dynamically evaluate and weight the fusion of vectorized features from text, images, and graphs, capturing the complementarity between different modal information, and ultimately generating a comprehensive, unbiased, and unified corporate risk representation vector. In practical applications, other methods can be used for feature fusion, and no specific limitations are imposed. This overcomes the limitations of a single modality. By integrating information from three different forms—text, images, and graph structure data—a more comprehensive and accurate semantic representation (corporate risk representation vector) can be generated, providing a more comprehensive risk profile for the enterprise. Moreover, based on the complementarity of multimodal features, implicit information that is difficult to capture with a single feature can be discovered, enhancing judgment. This facilitates improving the accuracy of due diligence on corporate accounts.

[0025] Step 104: Input the enterprise risk representation vector into the hybrid risk classification model to obtain the risk level of the enterprise output by the hybrid risk classification model.

[0026] The hybrid classification model is used to assess the risk of an enterprise based on its risk representation vector, thereby obtaining the risk level of the enterprise.

[0027] In this embodiment, a hybrid classification model can be used to assess the risk of an enterprise based on its enterprise risk representation vector, thereby obtaining the corresponding risk level. In practical applications, hybrid classification models are typically built based on a large amount of historical data and advanced algorithms, enabling a relatively scientific and objective assessment of enterprise risk. The output of this hybrid classification model has a certain degree of accuracy and stability, providing a quantitative assessment indicator for enterprise risk. The risk level can be low, medium, or high, with no specific limitation. This eliminates the need for manual risk assessment, saving time and reducing costs. Moreover, using a hybrid classification model for enterprise risk assessment ensures the uniformity and standardization of risk assessment standards.

[0028] Step 105: Based on the risk level, generate a due diligence report for the enterprise.

[0029] In this embodiment, a due diligence report for the enterprise can be generated based on the aforementioned risk levels. By using risk levels as the basis for generating the due diligence report, the conclusions can be more accurate and reliable. This addresses the problem of subjective audit standards that rely heavily on the auditors' personal professional experience, risk appetite, and sense of responsibility.

[0030] In this embodiment, multi-source heterogeneous data related to the enterprise can be acquired, and then features are extracted from the multi-source heterogeneous data to obtain multimodal features, including text features, image features, and graph structure features. Next, the vectorized text features, image features, and graph structure features can be fused to generate an enterprise risk representation vector. This enterprise risk representation vector can then be input into a hybrid risk classification model to obtain the enterprise's risk level. Finally, based on the risk level, a due diligence report for the enterprise can be generated. In this way, through automated data collection, feature extraction, feature fusion, and risk assessment, the efficiency of due diligence can be improved. Furthermore, graph structure features can indicate the enterprise's equity relationships. Using graph structure features for risk assessment can solve the problem of a single dimension of risk insight and a lack of deep correlation and penetrating identification capabilities. Moreover, the due diligence report is generated based on the enterprise's risk level output by the hybrid risk classification model, rather than human judgment, which can solve the problem of strong subjectivity in existing methods' review standards and achieve unified and standardized management of due diligence.

[0031] In one possible implementation, acquiring enterprise-related multi-source heterogeneous data includes: The system obtains enterprise declaration data submitted by enterprises during the application process for opening corporate accounts through the first application programming interface (API). Obtain authoritative third-party data from the target platform's enterprise through the second API; Acquire the image data of the enterprise and perform standardized preprocessing on the image data; Obtain publicly available online data of the aforementioned company from publicly available internet channels; Data governance and quality management are performed on the enterprise declaration data, the third-party authoritative data, the image data, and the publicly available online data to obtain multi-source heterogeneous data.

[0032] In the above embodiments, the multi-source heterogeneous data mainly includes enterprise declaration data, third-party authoritative data, image data, and publicly available online data. In practical applications, enterprise declaration data can be obtained through a first Application Programming Interface (API), such as through a secure and encrypted API interface or a bank's front-end business system, receiving structured form data submitted by enterprises during the account opening application process, including basic enterprise information, legal representative, shareholder structure, main business, etc. Simultaneously, third-party authoritative data from the target platform can be obtained through a second API. In practical applications, dedicated lines or API connections can be established with certain target platforms to automatically acquire and periodically update authoritative data such as enterprise registration information, tax rating, credit records, legal proceedings, import and export customs declarations, ensuring the timeliness and credibility of the data. In the above embodiments, image data of the enterprise can also be acquired and standardized preprocessed. For example, standardized preprocessing can be performed on image materials uploaded by enterprises, such as business licenses, legal representative ID cards, photos or videos of business premises, and lease contracts. This includes standardizing image formats and sizes, and performing image noise reduction. These image materials from different sources, in different formats, and of different qualities are transformed into a unified and standardized form, enabling more efficient and accurate subsequent review, analysis, storage, and management. In the above embodiments, publicly available online data of enterprises can also be obtained from publicly available internet channels. For example, relying on a distributed web crawler framework, news information, social media sentiment, recruitment information, intellectual property announcements, and other publicly available online data related to the enterprise can be collected according to preset rules. After obtaining the above data, data governance and quality management can be performed on enterprise-reported data, third-party authoritative data, image data, and publicly available online data. This includes standardized cleaning, format unification, entity alignment, and quality verification, thereby obtaining multi-source heterogeneous data. This can solve problems such as data inconsistency, incompleteness, and duplication. In practical applications, a unique data identifier can be established for each enterprise within the multi-source heterogeneous data to facilitate storage and differentiation. The above methods can be used to collect, clean, and integrate various data from scattered and heterogeneous internal and external data sources to form unified, high-quality multi-source heterogeneous data, providing reliable data support for subsequent due diligence.

[0033] In one possible implementation, the step of extracting features from the multi-source heterogeneous data to obtain the multimodal features corresponding to the enterprise includes: The pre-trained language model is used to perform named entity recognition and semantic relation extraction on the unstructured text in the multi-source heterogeneous data to generate text features; Sentiment analysis models are used to analyze public online data from the multi-source heterogeneous data, and event extraction techniques are combined to identify risk events. Add the risk event to the text feature.

[0034] In the above embodiments, pre-trained language models can be used to perform named entity recognition and semantic relationship extraction on unstructured text from multi-source heterogeneous data to generate text features. For example, BERT or similar pre-trained language models can be used to perform deep learning on long texts such as company bylaws and purchase and sale contracts to automatically extract key entities such as "beneficiary owner," "main business," and "major trading counterparty" and their semantic relationships for deep semantic parsing. In the above embodiments, sentiment analysis models can also be used to perform public opinion analysis on publicly available online data from multi-source heterogeneous data, and combined with event extraction techniques to identify risk events. For example, sentiment analysis models can be used to determine the positive or negative tendencies of public opinion information, and combined with event extraction techniques, risk events such as "abnormal operations," "enforcement information," and "major negative reports" can be automatically identified and warned of, enabling continuous tracking of dynamic risks for enterprises. In practical applications, sentiment analysis aims to determine the emotional tendency expressed by text, such as positive, negative, or neutral. Sentiment analysis models can be based on dictionary rules (SenticNet), traditional machine learning models (Naive Bayes algorithm), deep learning models (convolutional neural networks), or hybrid models combining pre-trained language models (BERT-BiLSTM-Attention model), without specific limitations. Event extraction technology is a technique for automatically identifying and extracting event-related information from unstructured or semi-structured text, aiming to present the events described in the text in a structured form for easier computer understanding and processing. In the above embodiments, the obtained risk events can be added to the text features, thus enriching the content of the text features and facilitating a more comprehensive due diligence investigation of corporate accounts.

[0035] In practical applications, after obtaining the above text features, the text features can be vectorized to map the text features into a continuous numerical space, transforming semantic information into high-dimensional text feature vectors, enabling machines to process text through mathematical operations.

[0036] In one possible implementation, the step of extracting features from the multi-source heterogeneous data to obtain the multimodal features corresponding to the enterprise includes: A deep convolutional neural network model is used to identify the authenticity of image data in the multi-source heterogeneous data, and a first image feature is generated. A target detection model is used to identify elements in the image data of the multi-source heterogeneous data to obtain second image features; Based on the first image features and the second image features, the image features corresponding to the enterprise are obtained.

[0037] In the above embodiments, deep convolutional neural network models can be used to authenticate image data from multi-source heterogeneous data, generating first image features. For example, deep convolutional neural network models such as ResNet and EfficientNet can be used to perform microscopic feature analysis on images of key documents such as business licenses and ID cards. This can detect the consistency of the font library, the compliance of the seal outline and texture, the validity of the QR code / barcode, and the authenticity of the linked content, thereby effectively identifying counterfeit documents processed by PS, photocopying, and other techniques, and achieving authentication of abnormal fonts, irregular seals, and synthetic traces. In the above embodiments, object detection models can also be used to identify elements in image data from multi-source heterogeneous data, obtaining second image features. For example, object detection models such as YOLOv5 or Faster R-CNN can be used to locate and identify company nameplates, office environments, production equipment, etc., in images of business premises, generating second image features representing the actual business situation. This can be cross-compared with the declared information, providing intuitive visual evidence to support "on-site operation." In the above embodiments, image features including first image features and second image features can be obtained, which are not only rich in meaning, but also provide solid image information for subsequent risk assessment of enterprises.

[0038] In practical applications, cutting-edge computer vision technology can be used for high-precision image authentication, performing microscopic analysis of the physical and digital features of key documents such as business licenses and legal representative ID cards; natural language processing technology can be used for consistency analysis of text content, identifying logical contradictions and inconsistencies between different materials; and multiple authoritative data can be integrated for automated cross-comparison to construct a multi-layered, three-dimensional verification system. This allows for the automatic and accurate identification of fraudulent activities such as forged documents, fictitious business locations, and intentionally false statements, fundamentally enhancing banks' ability to defend against external fraud risks.

[0039] In one possible implementation, the step of extracting features from the multi-source heterogeneous data to obtain the multimodal features corresponding to the enterprise includes: Based on the company's equity chain, legal representative, and investment relationships, a knowledge graph of the company is constructed. The graph structure features in the knowledge graph are extracted using a graph neural network model.

[0040] In the above embodiments, a knowledge graph of the enterprise can be constructed based on its equity chain, legal representative, and investment relationships. This knowledge graph can penetrate to the ultimate individual shareholders, clearly presenting complex shareholding paths and control chains. Then, a graph neural network model is used to extract graph structure features from the knowledge graph. This technology can intelligently identify deliberately hidden abnormal equity structures such as cross-shareholding, circular investment, and hidden actual controllers, revealing potential related-party transactions and risks of profit transfer. In practical applications, the applicant enterprise can be used as the root node, and a dynamic relationship graph covering all levels of shareholders, legal representatives, subsidiaries, related parties, and ultimately the ultimate individual can be constructed through penetrating queries. Based on this, a graph neural network model is used to perform embedding representation learning on the nodes and edges in the graph, and the potential transmission path and intensity of risks in the relationship network are captured through message passing mechanisms, thereby extracting graph structure features for identifying implicit relationships, shell companies, and complex holding structures.

[0041] In one possible implementation, the hybrid classification model is used to perform risk assessment on the enterprise based on the enterprise risk representation vector to obtain the risk level corresponding to the enterprise, including: The enterprise risk representation vector is subjected to compliance review through a rules engine; If the compliance review is passed, a machine learning model is used to assess the risk representation vector of the enterprise to obtain the risk level of the enterprise and the confidence level corresponding to the risk level.

[0042] In the above embodiments, a rule engine can perform a hard compliance review based on regulatory policies on the enterprise risk representation vector. Then, if the compliance review is passed, a machine learning model is used to assess the risk of the enterprise risk representation vector, obtaining the enterprise's corresponding risk level and the confidence level. In practical applications, the rule engine can first interpret the enterprise risk representation vector, converting it into text, and then review it. The rule engine can embed a configurable compliance and business rule base to facilitate the execution of hard logical judgments. For example, it can automatically compare the enterprise and its shareholders with anti-money laundering blacklists in real time, implementing a "one-vote veto"; or verify whether the beneficial owner information is complete, automatically rejecting and prompting for supplementation if missing. The rule engine can ensure rigid compliance of business operations. The aforementioned machine learning model can be a hybrid machine learning model based on XGBoost and deep neural networks. This machine learning model takes the fused enterprise risk representation vector as input and can output a quantified risk level and corresponding confidence level through supervised learning. The aforementioned machine learning model can handle complex nonlinear relationships, discover weak risk signals that are difficult for humans to detect, and achieve accurate risk classification. In practical applications, the specific architecture of the machine learning model is not specifically limited. By combining the aforementioned rule engine (rigid, deterministic logic) with machine learning models (flexible, probabilistic judgment), both compliance bottom lines and risk insights are taken into account.

[0043] In one possible implementation, generating a due diligence report for the company based on the risk level further includes: Using model interpretability techniques, the Shapley value of each feature in the enterprise risk representation vector is calculated to quantify the contribution of the feature to the risk level. The top K features with a contribution rate higher than a preset threshold are back-mapped to the risk sources of the features, where K is an integer greater than 1; Generate a list of key evidence sorted by contribution, and integrate the list of key evidence into the due diligence report, wherein the list of key evidence includes the sources of risk.

[0044] In practical applications, once a machine learning model completes a risk assessment of a company (e.g., outputting "high risk" with 85% confidence), an interpretability analysis process can be automatically triggered. This involves first receiving the aforementioned company risk representation vector, multi-source heterogeneous data, and multimodal features. Then, interpretability techniques such as SHAP or LIME can be used to perform attribution analysis on the risk level output by the hybrid classification model. In practice, a pre-computed global SHAP interpreter can be used to calculate the Shapley value, analyze the contribution of each feature in the company risk representation to the prediction result, and then trace the key features back to their original sources of risk: the top K features with the highest contribution are mapped back to their original risk sources. For example, it might be found that the "fund transaction network density" feature (derived from a knowledge graph) and the "negative sentiment score" feature (derived from text features) are the main contributing factors. Next, methods such as LIME or Integrated Gradients can be used to analyze the text features, highlighting the words or sentences that contribute the most to the risk (e.g., "frequent," "guarantee," "overdue"). It can also analyze subgraph structures, using graph attention mechanisms or subgraph probing methods to identify and extract core correlation paths leading to risk (e.g., "the company is connected to a defaulting company through two layers of holding"). If images are involved, visualization techniques such as Grad-CAM can be used to generate heatmaps from the input image data, marking key areas in the image that affect the model's judgment (e.g., a suspicious stamp area on a document). In this way, for a specific company, it can be determined which text, which part of the image, or which relationship in the knowledge graph caused its high 'funding network density' or 'negative public opinion score'. Subsequently, a list of key evidence explaining the source of risk and ranked by contribution can be generated and integrated into the due diligence report.

[0045] The above methods can provide clear attribution analysis for each risk decision, generate easy-to-understand explanatory reports, and clearly identify the key factors that trigger risks, such as "the company is judged to be high-risk because of its high-frequency financial transactions with three deregistered companies," which greatly enhances the transparency and credibility of the decision-making process.

[0046] In practical applications, it enables digital traceability of the entire due diligence process, with complete records of every step of data retrieval, feature analysis, model reasoning, and rule judgment, ensuring full traceability of the decision-making process. Furthermore, the embedded interpretability technology clearly elucidates the logic and basis of each risk decision, generating easily understandable due diligence reports. This not only meets the stringent compliance requirements of certain financial regulatory agencies in the areas of Know Your Customer (KYC) and Anti-Money Laundering (AML), but also achieves absolute uniformity of audit standards within the institution, eliminating inconsistencies in standards and potential compliance risks arising from differences in personnel experience and subjective judgment.

[0047] In practical applications, the aforementioned due diligence report can include basic corporate information, a visual equity structure diagram, a detailed list of risk points, the final risk level, and specific handling recommendations. In this way, automatically generating a standardized due diligence report with a complete structure and rich illustrations can completely free up the productivity of manual report writing.

[0048] In practical applications, algorithmic models and rule engines can be used to output objective and consistent quantitative risk scores and highly structured due diligence reports. These reports not only include clear risk level conclusions but, more importantly, clearly list the basis for risk decisions and key evidence chains. This aims to provide bank approval personnel with highly consistent, transparent, and credible decision-making references, completely eliminating fluctuations in review standards caused by differences in personnel experience and subjective judgment. It systematically reduces operational and regulatory penalty risks arising from insufficient compliance review or inconsistent standards, thereby enhancing the robustness of the bank's overall compliance management system.

[0049] In one possible implementation, after generating the due diligence report corresponding to the enterprise based on the risk level, the method further includes: If the risk level is lower than the preset level, an automatic approval instruction is triggered to determine that the enterprise has passed the review; If the risk level is equal to or higher than the preset level, the company's application and the due diligence report will be pushed to the manual review queue corresponding to the risk level.

[0050] In the above embodiments, corresponding strategies can be executed based on the specific risk level. In practical applications, a highly available RESTful API service can be used to push the due diligence report and risk level to the bank's core business system or credit approval process engine in real time. Subsequently, according to preset strategies, subsequent process nodes such as "approved," "transferred to manual review," or "rejected" can be automatically triggered. For example, an automatic approval instruction can be triggered for low-risk applications, while medium- and high-risk applications can be pushed to the corresponding level of manual review queue, along with a complete due diligence report for decision-making reference. In this way, a seamless connection between due diligence and approval can be achieved.

[0051] Figure 2 By combining processes and components, the overall workflow and technical architecture of the due diligence method for corporate accounts in this application embodiment are systematically demonstrated.

[0052] Step 1: Acquire multi-source heterogeneous data related to the enterprise; Step 2: Obtain image features, text features, and graph structure features, and then perform feature fusion to obtain the enterprise risk representation vector; Step 3: Conduct a risk assessment; Step 4: Generate a due diligence report and proceed to the next step.

[0053] This application also provides a corporate account due diligence system, which can be used to implement the aforementioned corporate account due diligence method. In practical applications, the corporate account due diligence system takes multimodal data fusion as its core and adopts a dual-drive strategy of rule engine and machine learning model to achieve automation, intelligence, and standardization of the entire corporate account due diligence process. In practical applications, for example, the system may include a four-layer system architecture: data layer, analysis layer, decision layer, and application layer.

[0054] The data layer, as the foundation of the system, is responsible for collecting, cleaning, and integrating data from dispersed and heterogeneous internal and external data sources to form unified, high-quality data assets. The analysis layer, the system's "intelligent brain," is responsible for in-depth analysis and feature extraction of the multimodal information provided by the data layer, and for initially achieving cross-modal semantic association. The decision layer receives the high-dimensional features output by the analysis layer and, through fusion and reasoning, forms the final risk assessment and decision recommendations. The application layer encapsulates intelligent decision-making capabilities into services, directly empowering banking business processes and operational management.

[0055] Based on the same inventive concept, this application also provides a due diligence device for corporate accounts, the structure of which is as follows: Figure 3 As shown.

[0056] Figure 3 This is a schematic diagram of the internal structure of a due diligence device for corporate accounts provided in an embodiment of this application. Figure 3 As shown, the device includes: At least one processor 301; And a memory 302 that is communicatively connected to at least one processor; The memory 302 stores instructions that can be executed by at least one processor. The instructions are executed by at least one processor 301 to enable at least one processor 301 to perform the above-described due diligence method for corporate accounts.

[0057] In one possible implementation, the processor 301 is capable of performing the following actions: acquiring multi-source heterogeneous data related to an enterprise, wherein the enterprise is the applicant for a corporate account; extracting features from the multi-source heterogeneous data to obtain multimodal features corresponding to the enterprise, wherein the multimodal features include text features, image features, and graph structure features, and the graph structure features are used to indicate the equity relationship of the enterprise; performing feature fusion on the vectorized text features, image features, and graph structure features to generate an enterprise risk representation vector corresponding to the enterprise; inputting the enterprise risk representation vector into a hybrid risk classification model to obtain the risk level of the enterprise output by the hybrid risk classification model, wherein the hybrid classification model is used to perform risk assessment on the enterprise based on the enterprise risk representation vector to obtain the risk level of the enterprise; and generating a due diligence report corresponding to the enterprise based on the risk level.

[0058] Some embodiments of this application provide corresponding to Figure 1 A non-volatile computer storage medium storing computer-executable instructions, wherein the computer-executable instructions are configured to execute the aforementioned due diligence method for corporate accounts.

[0059] In one possible implementation, the computer-executable instructions are configured to perform: acquiring multi-source heterogeneous data related to the enterprise, wherein the enterprise is the applicant enterprise for a corporate account; extracting features from the multi-source heterogeneous data to obtain multimodal features corresponding to the enterprise, wherein the multimodal features include text features, image features, and graph structure features, wherein the graph structure features are used to indicate the equity relationship of the enterprise; performing feature fusion on the vectorized text features, image features, and graph structure features to generate an enterprise risk representation vector corresponding to the enterprise; inputting the enterprise risk representation vector into a hybrid risk classification model to obtain the risk level of the enterprise output by the hybrid risk classification model, wherein the hybrid classification model is used to perform risk assessment on the enterprise based on the enterprise risk representation vector to obtain the risk level of the enterprise; and generating a due diligence report corresponding to the enterprise based on the risk level.

[0060] The various embodiments in this application are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the embodiments for IoT devices and media are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0061] The systems, media, and methods provided in this application are one-to-one correspondences. Therefore, the systems and media also have similar beneficial technical effects as their corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be repeated here.

[0062] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0063] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0064] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0065] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0066] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.

[0067] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0068] Computer-readable media include both permanent and non-permanent, removable and non-removable media that can store information by any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.

[0069] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0070] The above description is merely an embodiment of this application and is not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.

Claims

1. A method for due diligence on corporate accounts, characterized in that, include: Acquire multi-source heterogeneous data related to enterprises, wherein the enterprises are those that applied for corporate accounts; Feature extraction is performed on the multi-source heterogeneous data to obtain the multimodal features corresponding to the enterprise. The multimodal features include text features, image features, and graph structure features. The graph structure features are used to indicate the equity relationship of the enterprise. The vectorized text features, image features, and graph structure features are fused to generate a corporate risk representation vector corresponding to the enterprise. The enterprise risk representation vector is input into the hybrid risk classification model to obtain the risk level of the enterprise output by the hybrid risk classification model. The hybrid classification model is used to perform risk assessment on the enterprise based on the enterprise risk representation vector to obtain the risk level of the enterprise. Based on the risk level, a due diligence report is generated for the corresponding company.

2. The method according to claim 1, characterized in that, The acquisition of multi-source heterogeneous data related to the enterprise includes: The system obtains enterprise declaration data submitted by enterprises during the application process for opening corporate accounts through the first application programming interface (API). Obtain authoritative third-party data from the target platform's enterprise through the second API; Acquire the image data of the enterprise and perform standardized preprocessing on the image data; Obtain publicly available online data of the aforementioned company from publicly available internet channels; Data governance and quality management are performed on the enterprise declaration data, the third-party authoritative data, the image data, and the publicly available online data to obtain multi-source heterogeneous data.

3. The method according to claim 2, characterized in that, The step of extracting features from the multi-source heterogeneous data to obtain the multimodal features corresponding to the enterprise includes: The pre-trained language model is used to perform named entity recognition and semantic relation extraction on the unstructured text in the multi-source heterogeneous data to generate text features; Sentiment analysis models are used to analyze public online data from the multi-source heterogeneous data, and event extraction techniques are combined to identify risk events. Add the risk event to the text feature.

4. The method according to claim 2, characterized in that, The step of extracting features from the multi-source heterogeneous data to obtain the multimodal features corresponding to the enterprise includes: A deep convolutional neural network model is used to identify the authenticity of image data in the multi-source heterogeneous data, and a first image feature is generated. A target detection model is used to identify elements in the image data of the multi-source heterogeneous data to obtain second image features; Based on the first image features and the second image features, the image features corresponding to the enterprise are obtained.

5. The method according to claim 1, characterized in that, The step of extracting features from the multi-source heterogeneous data to obtain the multimodal features corresponding to the enterprise includes: Based on the company's equity chain, legal representative, and investment relationships, a knowledge graph of the company is constructed. The graph structure features in the knowledge graph are extracted using a graph neural network model.

6. The method according to claim 1, characterized in that, The hybrid classification model is used to assess the risk of the enterprise based on the enterprise risk representation vector, and obtain the risk level corresponding to the enterprise, including: The enterprise risk representation vector is subjected to compliance review through a rules engine; If the compliance review is passed, a machine learning model is used to assess the risk representation vector of the enterprise to obtain the risk level of the enterprise and the confidence level corresponding to the risk level.

7. The method according to claim 1, characterized in that, Based on the risk level, the due diligence report for the corresponding enterprise is generated, and the report also includes: Using model interpretability techniques, the Shapley value of each feature in the enterprise risk representation vector is calculated to quantify the contribution of the feature to the risk level. The top K features with a contribution rate higher than a preset threshold are back-mapped to the risk sources of the features, where K is an integer greater than 1; Generate a list of key evidence sorted by contribution, and integrate the list of key evidence into the due diligence report, wherein the list of key evidence includes the sources of risk.

8. The method according to claim 1, characterized in that, After generating the due diligence report corresponding to the enterprise based on the risk level, the method further includes: If the risk level is lower than the preset level, an automatic approval instruction is triggered to determine that the enterprise has passed the review; If the risk level is equal to or higher than the preset level, the company's application and the due diligence report will be pushed to the manual review queue corresponding to the risk level.

9. A due diligence device for corporate accounts, characterized in that, The device includes: At least one processor; And, a memory communicatively connected to the at least one processor; The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform a due diligence method for a corporate account as described in any one of claims 1-8.

10. A computer storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed, they implement a due diligence method for corporate accounts as described in any one of claims 1-8.