Enterprise operation risk warning method, device, computer equipment and medium

Through multimodal model, feature extraction and alignment of enterprise transaction images and text data is generated to generate visual data links, which solves the problem of insufficient accuracy of enterprise operating risk warning and achieves more efficient risk identification and early warning.

CN119990782BActive Publication Date: 2025-07-08HANGZHOU MUNICIPAL PUBLIC SECURITY BUREAU HIGH-TECH IND DEV ZONE BRANCH HANGZHOU MUNICIPAL PUBLIC SECURITY BUREAU BINJIANG DISTRICT BRANCH +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510455401.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-07-08
Estimated Expiration
2045-04-11

AI Technical Summary

Technical Problem

The accuracy of enterprise operating risk warning in the prior art needs to be improved, and it is difficult to effectively identify and warn of potential risks.

Method used

Multimodal model is used to extract and align the enterprise's transaction image data and survey text data, and extract features of different granularity through convolutional neural network with adjustable hollow rate and regional attention mechanism to generate visual data links for risk warning.

Benefits of technology

It improves the accuracy and efficiency of enterprise operating risk warnings, can analyze and warn of potential risks more quickly and accurately, provide intuitive data display, and help regulatory authorities take effective measures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119990782B_ABST
    Figure CN119990782B_ABST
Patent Text Reader

Abstract

This application relates to the field of artificial intelligence technology, and discloses an enterprise operation risk warning method, device, computer equipment and medium. First, multi-frame transaction image data and investigation text data of the target enterprise are obtained; then, the first multi-modal model is used to perform correlation prediction on the first transaction image data and the investigation text data to obtain the first text-image correlation probability; then, the second multi-modal model is used to perform correlation prediction on the second transaction image data and the investigation text data to obtain the second text-image correlation probability; finally, a visual data chain is generated based on the first text-image correlation probability, the second text-image correlation probability, the investigation text data and the multi-frame transaction image data to perform operation risk warning on the target enterprise. Based on the multi-modal large model, feature alignment is performed on the transaction image data and the investigation text data, and data from different sources are comprehensively analyzed to achieve comprehensive mining of information, so as to obtain more accurate text-image correlation probability data and improve the accuracy of risk warning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular, to a method, device, computer device, and medium for early warning of enterprise operation risks. Background Art

[0002] Early warning of enterprise operation risks can timely identify potential operation risks, help regulatory authorities take effective measures to prevent the occurrence of illegal acts, and avoid financial risks caused by enterprise operation problems. However, the accuracy of early warning of enterprise operation risks in related technologies needs to be improved.

[0003] Therefore, there is an urgent need to propose a new method for early warning of enterprise operation risks. Summary of the Invention

[0004] This application provides a method, device, computer device, and medium for early warning of enterprise operation risks, which solves the technical problem that the accuracy of early warning of enterprise operation risks in related technologies needs to be improved, and achieves the technical effect of more accurate early warning of enterprise operation risks.

[0005] To achieve the above object, the main technical solutions adopted in this application include:

[0006] In a first aspect, an embodiment of this application provides a method for early warning of enterprise operation risks, and the method includes:

[0007] Obtain multi-frame transaction image data and investigation text data of a target enterprise; wherein, any transaction image data corresponds to a target image type;

[0008] Invoke a target multi-modal model corresponding to the target image type; wherein, the feature extraction granularity of the target multi-modal model matches the image content complexity of any transaction image data;

[0009] Input any transaction image data and the investigation text data into the target multi-modal model for feature extraction and feature alignment, and obtain the graphic-text correlation probability data between any transaction image data and the investigation text data;

[0010] Generate a visual data chain according to the graphic-text correlation probability data, the investigation text data, and the multi-frame transaction image data to conduct early warning of the operation risks of the target enterprise.

[0011] In the embodiments of the present application, based on transaction image data for different image types, corresponding multimodal large models are called for analysis, which can balance the accuracy of processing complex transaction image data and the efficiency of processing simple transaction images. Further, based on the first multimodal model, a convolutional neural network with adjustable dilation rate is used to globally extract microscopic features of the first granularity of the first transaction image data, and then correlated prediction is performed with the survey text data to obtain a more accurate first text-image correlation probability. Based on the second multimodal model, a regional attention mechanism is adopted to also globally extract macroscopic features of the second granularity of the second transaction image data, and then correlated prediction is performed with the survey text data to obtain a more accurate second text-image correlation probability. Finally, a visual data chain is generated based on the first text-image correlation probability, the second text-image correlation probability, the survey text data, and the transaction image data, facilitating office staff to more quickly and accurately analyze, judge, and warn of the business operation risks of the target enterprise.

[0012] In a second aspect, an embodiment of the present application provides a method for warning of business operation risks of an enterprise, the method including:

[0013] Obtain multi-frame transaction image data and survey text data of a target enterprise; wherein, the multi-frame transaction image data includes first transaction image data belonging to a first image type and second transaction image data belonging to a second image type; the first image type corresponds to a first multimodal model, the second image type corresponds to a second multimodal model, and the first multimodal model and the second multimodal model are in parallel;

[0014] Perform correlated prediction on the first transaction image data and the survey text data through the first multimodal model to obtain a first text-image correlation probability between the first transaction image data and the survey text data; wherein, the first multimodal model uses a convolutional neural network with adjustable dilation rate to globally extract microscopic features of the first granularity of the first transaction image data;

[0015] Perform correlated prediction on the second transaction image data and the survey text data through the second multimodal model to obtain a second text-image correlation probability between the second transaction image data and the survey text data; wherein, the second multimodal model uses a regional attention mechanism to also globally extract macroscopic features of the second granularity of the second transaction image data, and the first granularity is smaller than the second granularity;

[0016] Generate a visual data chain according to the first text-image correlation probability, the second text-image correlation probability, the survey text data, and the multi-frame transaction image data to perform a warning of the business operation risks of the target enterprise.

[0017] In a third aspect, an embodiment of the present application provides a device for warning of business operation risks of an enterprise, the device including:

[0018] A data acquisition module, configured to acquire multiple frames of transaction image data and survey text data of a target enterprise; wherein, any one of the transaction image data corresponds to a target image type.

[0019] A model invocation module, configured to invoke a target multi-modal model corresponding to the target image type; wherein, the feature extraction granularity of the target multi-modal model matches the image content complexity of any one of the transaction image data.

[0020] A data processing module, configured to input any one of the transaction image data and the survey text data into the target multi-modal model for feature extraction and feature alignment, so as to obtain graphic-text association probability data between any one of the transaction image data and the survey text data.

[0021] A data chain generation module, configured to generate a visualization data chain according to the graphic-text association probability data, the survey text data, and the multiple frames of transaction image data, so as to perform an operation risk warning on the target enterprise.

[0022] In a fourth aspect, an embodiment of the present application provides an enterprise operation risk warning device, and the device includes:

[0023] A data acquisition module, configured to acquire multiple frames of transaction image data and survey text data of a target enterprise; wherein, the multiple frames of transaction image data include first transaction image data belonging to a first image type and second transaction image data belonging to a second image type; the complexity of the first transaction image data is greater than the complexity of the second transaction image data.

[0024] A first data processing module, configured to invoke a first multi-modal model corresponding to the first image type, and perform an association prediction on the first transaction image data and the survey text data through the first multi-modal model, so as to obtain a first graphic-text association probability between the first transaction image data and the survey text data; wherein, the feature extraction granularity of the first multi-modal model matches the image content complexity of the first transaction image data.

[0025] A second data processing module, configured to invoke a second multi-modal model corresponding to the second image type, and perform an association prediction on the second transaction image data and the survey text data through the second multi-modal model, so as to obtain a second graphic-text association probability between the second transaction image data and the survey text data; wherein, the feature extraction granularity of the second multi-modal model matches the image content complexity of the second transaction image data.

[0026] A data chain generation module, configured to generate a visual data chain according to the first image-text association probability, the second image-text association probability, the survey text data, and the multi-frame transaction image data, so as to perform an operation risk warning on the target enterprise.

[0027] In a fifth aspect, an embodiment of the present application provides a computer device, including:

[0028] A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method described in any one of the above embodiments.

[0029] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method described in any one of the above embodiments.

[0030] In the embodiment of the present application, first, multi-frame transaction image data and survey text data of a target enterprise are obtained; then, a target multi-modal model corresponding to the target image type is called; then, any transaction image data and survey text data are input into the target multi-modal model for feature extraction and feature alignment to obtain the image-text association probability data between any transaction image data and survey text data; finally, a visual data chain is generated according to the image-text association probability data, the survey text data, and the multi-frame transaction image data, so as to perform an operation risk warning on the target enterprise. Based on the target multi-modal large model for feature extraction and feature alignment of the multi-frame transaction image data and the survey text data, data from different sources can be comprehensively analyzed, the information can be fully mined, so as to obtain more accurate image-text association probability data and improve the accuracy of risk warning. Description of the Drawings

[0031] In order to more clearly illustrate the specific embodiments of the present application or the technical solutions in the prior art, the following will briefly introduce the drawings required to be used in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0032] Figure 1a It is a flowchart of the enterprise operation risk warning method provided by the embodiment of this specification;

[0033] Figure 1b It is a schematic diagram of the visual data chain provided by the embodiment of this specification;

[0034] Figure 2 It is a flowchart of the enterprise operation risk warning method provided by the embodiment of this specification;

[0035] Figure 3 It is a flowchart of the enterprise operation risk warning method provided by the embodiments of this specification;

[0036] Figure 4 It is a flowchart of the enterprise operation risk warning method provided by the embodiments of this specification;

[0037] Figure 5 It is a flowchart of the enterprise operation risk warning method provided by the embodiments of this specification;

[0038] Figure 6 It is a flowchart of the enterprise operation risk warning method provided by the embodiments of this specification;

[0039] Figure 7 It is a flowchart of the enterprise operation risk warning method provided by the embodiments of this specification;

[0040] Figure 8 It is a schematic diagram of the enterprise operation risk warning device provided by the embodiments of this specification;

[0041] Figure 9 It is a schematic diagram of a computer structure provided by the embodiments of this specification. Detailed implementation manners

[0042] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Apparently, the described embodiments are some but not all of the embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the scope of protection of this application.

[0043] Enterprise operation risk warning can timely identify potential operation risks, provide effective decision-making support for regulatory authorities, help them take targeted measures to prevent and contain the occurrence of illegal acts, and avoid negative impacts such as financial risks and market fluctuations caused by enterprise operation problems. With the increasingly complex economic environment, the relevant technologies need to be improved in terms of the accuracy of enterprise operation risk warning to better cope with the increasing risk challenges.

[0044] Based on this, the present application proposes an enterprise operation risk early warning method. First, obtain multiple frames of transaction image data and investigation text data of the target enterprise; then call the target multimodal model corresponding to the target image type; next, input any transaction image data and investigation text data into the target multimodal model for feature extraction and feature alignment to obtain the graphic-text association probability data between any transaction image data and investigation text data; finally, generate a visual data chain based on the graphic-text association probability data, investigation text data, and multiple frames of transaction image data to conduct an operation risk early warning for the target enterprise. By calling the corresponding target multimodal large model for analysis based on the transaction image data for different target image types, it is possible to balance the accuracy of processing complex transaction image data and the efficiency of processing simple transaction images. Further, based on the target multimodal large model, feature extraction and feature alignment are performed on multiple frames of transaction image data and investigation text data, so that data from different sources can be comprehensively analyzed to achieve comprehensive information mining, thereby obtaining more accurate graphic-text association probability data. Finally, based on the visual data chain, office staff can analyze, judge, and early warn the operation risk of the target enterprise more quickly and accurately.

[0045] Further, in some embodiments, the enterprise operation risk early warning method may include: obtaining multiple frames of transaction image data and investigation text data of the target enterprise; determining the target image type to which any transaction image data belongs according to the image content complexity of any transaction image data; jointly inputting the any transaction image data and the investigation text data into a target multimodal model matching the target image type for graphic-text association prediction to obtain the graphic-text association probability data between the any transaction image data and the investigation text data; according to the graphic-text association probability data, selecting multiple frames of candidate transaction image data that meet the first-level screening requirements from the multiple frames of transaction image data; performing consistency verification on any candidate transaction image data and the structured business data of the target enterprise to select target image data that meet the second-level screening requirements from the multiple frames of candidate transaction image data; generating a visual data chain based on the target image data and the investigation text data to conduct an operation risk early warning for the target enterprise.

[0046] It should be noted that the first-level screening requirements may be requirements set for the graphic-text association situation, and the second-level screening requirements may be requirements set for the consistency situation between the transaction image and the enterprise business data. Using the graphic-text association probability data (such as the first graphic-text association probability, the second graphic-text association probability) to select candidate transaction image data that is graphically and textually associated with the investigation text data at the first level. Further, for the accuracy of information, the structured business data obtained by structuring the enterprise business data is verified for consistency with the candidate transaction image data to ensure the accuracy of the data and improve the credibility of the visual data chain.

[0047] According to an embodiment of the present application, an embodiment of an enterprise operation risk early warning method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0048] Please refer to Figure 1a , in this embodiment, an enterprise operation risk early warning method is provided, and the method includes:

[0049] S101. Obtain multi-frame transaction image data and investigation text data of the target enterprise.

[0050] Among them, the multi-frame transaction image data includes first transaction image data belonging to a first image type and second transaction image data belonging to a second image type. The complexity of the first transaction image data is greater than that of the second transaction image data. The first image type and the second image type are used to characterize different complexities of information in the transaction image data, such as simple and complex. The first image type corresponds to a first multimodal model, the second image type corresponds to a second multimodal model, and the first multimodal model and the second multimodal model are in parallel;

[0051] In some embodiments, the transaction image data of the target enterprise can be obtained from a bank transaction system, including first transaction image data belonging to a first image type and second transaction image data belonging to a second image type. In other embodiments, the first transaction image data belonging to the first image type and the second transaction image data belonging to the second image type can be generated according to the transaction data of the target enterprise obtained from the bank transaction system or the transaction data in the financial records of the target enterprise.

[0052] In some embodiments, the regulatory department investigates the various business activities of the target enterprise to generate investigation text data, which may include investigation reports, contract documents, financial data, tax data, employee information, partner information, communication record data, etc.

[0053] In some embodiments, both the first multimodal model and the second multimodal model can be multimodal models that process image and text data. The first multimodal model can process the first transaction image data and the investigation text data; the second multimodal model can process the second transaction image data and the investigation text data. The first multimodal model and the second multimodal model are in parallel and can process the corresponding types of transaction images simultaneously.

[0054] S103. Perform an association prediction on the first transaction image data and the survey text data through the first multi-modal model to obtain the first graphic-text association probability between the first transaction image data and the survey text data.

[0055] Among them, the first multi-modal model uses a convolutional neural network with an adjustable dilation rate to globally extract the microscopic features of the first granularity of the first transaction image data. The feature extraction granularity of the first multi-modal model matches the image content complexity of the first transaction image data. The first graphic-text association probability can be data used to quantify the relationship between the first transaction image data and the survey text data.

[0056] In some embodiments, the first transaction image data and the survey text data are input into the first multi-modal model for feature extraction. The feature extraction granularity matches the image content complexity of the first transaction image data to obtain the first transaction image features and the survey text features. Then, by mapping the first transaction image features and the survey text features into the same feature space, feature alignment is performed. Finally, the similarity between the first transaction image features and the survey text features (which can be the Euclidean distance or the cosine similarity) is compared, and the similarity is used as the first graphic-text association probability.

[0057] In some embodiments, the first multi-modal model can use a convolutional neural network with an adjustable dilation rate to globally extract the microscopic features of the first granularity of the first transaction image data. Specifically, the first multi-modal model realizes the extraction of microscopic features of different scales by dynamically adjusting the interval distance of the convolutional kernel (for example, adjusting the dilation rate from 1 to 3). Exemplarily, when the first transaction image data is relatively complex, different dilation rates can be used according to the number of detailed features of different regions of the first granularity in the image. When the detailed features of a certain region are lower than the preset threshold, a larger dilation rate is used; when the detailed features of a certain region are higher than the preset threshold, a larger dilation rate is used. By using a convolutional neural network with an adjustable dilation rate in the first multi-modal model, the feature extraction of the first transaction image data can take into account the accuracy of different regions. It can extract relatively rough features in regions with fewer detailed features through a high dilation rate, and extract sufficiently fine features in regions with more detailed features through a low dilation rate, thus ensuring the effectiveness of feature extraction and contributing to obtaining a more accurate first graphic-text association probability.

[0058] S105. Perform an association prediction on the second transaction image data and the survey text data through the second multi-modal model to obtain the second graphic-text association probability between the second transaction image data and the survey text data.

[0059] Among them, the second multi-modal model also extracts the macroscopic features of the second granularity of the second transaction image data globally by using the regional attention mechanism, and the first granularity is smaller than the second granularity. The feature extraction granularity of the second multi-modal model matches the image content complexity of the second transaction image data. The second graphic-text association probability can be data used to quantify the relationship between the second transaction image data and the investigation text data.

[0060] Among them, the feature extraction granularity of the second multi-modal model matches the image content complexity of the second transaction image data. The second graphic-text association probability can be data used to quantify the relationship between the second transaction image data and the investigation text data.

[0061] In some embodiments, the second transaction image data and the investigation text data are input into the second multi-modal model for feature extraction. The feature extraction granularity matches the image content complexity of the second transaction image data, obtaining the second transaction image features and the investigation text features. Then, by mapping the second transaction image features and the investigation text features into the same feature space, feature alignment is performed. Finally, the similarity between the second transaction image features and the investigation text features (which can be the Euclidean distance or the cosine similarity) is compared, and the similarity is used as the second graphic-text association probability.

[0062] In some embodiments, the second multi-modal model can use the regional attention mechanism to extract the macroscopic features of the second granularity of the second transaction image data globally. Specifically, the second multi-modal model automatically focuses on the macroscopic features with semantic importance in the second transaction image by calculating the association weights between the image regions of the second granularity and the keywords in the investigation text data. Exemplarily, when the second transaction image data is relatively simple, the second multi-modal model uses the regional attention mechanism to generate an attention heat map through a spatial transformation network. After dividing the image into 10×10 grid cells, the regions with a correlation higher than a threshold (such as 0.7) with the keywords in the text data are selected for feature enhancement. By adopting the regional attention mechanism, the second multi-modal model improves the model's recognition ability of the semantic correlation between the image and the text, improves the accuracy of the feature extraction of the second transaction image data, and helps to obtain a more accurate second graphic-text association probability.

[0063] S107. Generate a visual data chain based on the first graphic-text association probability, the second graphic-text association probability, the investigation text data, and the multi-frame transaction image data to conduct an operating risk warning for the target enterprise.

[0064] Among them, the visual data chain can be a data display chain generated by visually displaying the first graphic-text association probability, the second graphic-text association probability, the investigation text data, and the multi-frame transaction image data, which helps office staff intuitively understand and analyze the association relationships in the transaction image data and the investigation text data.

[0065] In some embodiments, a visualization data chain is generated based on a first graphic-text association probability, a second graphic-text association probability, survey text data, and multi-frame transaction image data. The visualization data chain can be presented graphically. In addition, the visualization data chain can also be presented in the form of a chart or other visual forms. Through the visualization data chain, complex data relationships can become more intuitive, which helps to improve the analysis and identification efficiency of the business operation risks of an enterprise and timely issue a business operation risk warning for the target enterprise.

[0066] In some embodiments, first associated image data associated with the survey text data is selected from multi-frame first transaction image data according to the first graphic-text association probability. Second associated image data associated with the survey text data is selected from multi-frame second transaction image data according to the second graphic-text association probability. The first associated image data, the second associated image data, and the survey text data are used for visual display to obtain a visualization data chain, so as to issue a business operation risk warning for the target enterprise.

[0067] In the above embodiments, based on the transaction image data for different image types, corresponding multi-modal large models are called for analysis, which can take into account both the accuracy of processing complex transaction image data and the efficiency of processing simple transaction image data. Further, based on the first multi-modal model, a convolutional neural network with an adjustable dilation rate is adopted to globally extract the microscopic features of the first granularity of the first transaction image data, and then associated prediction is performed with the survey text data to obtain a more accurate first graphic-text association probability. Based on the second multi-modal model, a regional attention mechanism is adopted to also globally extract the macroscopic features of the second granularity of the second transaction image data, and then associated prediction is performed with the survey text data to obtain a more accurate second graphic-text association probability. Finally, a visualization data chain is generated based on the first graphic-text association probability, the second graphic-text association probability, the survey text data, and the transaction image data, which facilitates office staff to more quickly and accurately analyze, judge, and warn of the business operation risks of the target enterprise.

[0068] In some embodiments, generating a visualization data chain based on a first graphic-text association probability, a second graphic-text association probability, survey text data, and multi-frame transaction image data to issue a business operation risk warning for the target enterprise includes: determining target image data associated with the survey text data from the multi-frame transaction image data according to the first graphic-text association probability and the second graphic-text association probability; generating a visualization data chain based on the target image data and the survey text data.

[0069] Among them, the transaction image data may be image information generated in the transaction activities of the target enterprise. The survey text data may be text information related to the business activities of the target enterprise obtained through investigation.

[0070] In some embodiments, the transaction image data of the target enterprise can be obtained from the bank transaction system, such as the charts output by the bank transaction system. In addition, charts required for the investigation can be generated based on the transaction data of the target enterprise obtained from the bank transaction system or the transaction data in the financial records of the target enterprise, and used as the transaction image data.

[0071] In some embodiments, the regulatory department conducts investigations on various business activities of the target enterprise to generate investigation text data, which may include investigation reports, contract documents, financial data, tax data, employee information, partner information, communication record data, etc. Exemplarily, the revenue data, profit data, liability data, etc. in the financial data can intuitively display the financial health of the target enterprise; the transaction terms, amounts, performance status, etc. in the contract documents can reflect the business transactions and credit status of the target enterprise; the tax payment situation in the tax records can reflect the compliance operation of the target enterprise.

[0072] In some embodiments, the level of the text-image association probability represents the degree of correlation between the transaction image data and the investigation text data. When the text-image association probability is high, it means that there is a high degree of matching and correlation between the transaction image data and the investigation text data in terms of content. Exemplarily, if a transaction image data shows the outflow of a large amount of funds of an enterprise, and the contract document in the investigation text data also mentions the payment of this fund, then the text-image association probability between them will be high. This high correlation indicates that this transaction behavior is associated with the business activities of the target enterprise, and may imply potential risks, such as unclear fund flow, abnormal large expenditures, etc., thus enabling early warning of the business risks of the target enterprise.

[0073] In some embodiments, a first text-image association probability threshold and a second text-image association probability threshold are set. When the first text-image association probability is higher than the first text-image association probability threshold, it is determined that the corresponding first transaction image data is the target image data associated with the investigation text data; when the second text-image association probability is higher than the second text-image association probability threshold, it is determined that the corresponding second transaction image data is the target image data associated with the investigation text data.

[0074] In some embodiments, to more intuitively display the relationship between target image data and survey text data, a visual data chain can be generated. The visual data chain can be displayed in various ways, such as a graphical interface, an interactive chart, etc. In the graphical interface, the transaction image data and the survey text data can be displayed in the form of nodes, and the connection lines between the nodes represent the first graphic-text association probability or the second graphic-text association probability between them. The thickness of the connection lines can represent the magnitude of the graphic-text association probability. Additionally, different ranges of graphic-text association probability values can be corresponded to different risk levels. For example, when the graphic-text association probability is less than the first risk threshold, this graphic-text association probability corresponds to a low risk level; when the graphic-text association probability is greater than or equal to the first risk threshold and less than the second risk threshold, this graphic-text association probability corresponds to a medium risk level; when the graphic-text association probability is greater than or equal to the second risk threshold, this graphic-text association probability corresponds to a high risk level. Different colors or icons can be used to identify different risk levels. For example, red represents a high risk level, yellow represents a medium risk level, and green represents a low risk level. In this way, office staff can quickly browse and understand the relationship between a large amount of data, and timely discover potential business risks of the target enterprise. Through the visual data chain, an intuitive and convenient way is realized to display the actual situation of the target enterprise, which is beneficial to the business risk early warning of the target enterprise.

[0075] In the above embodiments, first, target image data associated with the survey text data is determined in multiple frames of transaction image data according to the first graphic-text association probability data and the second graphic-text association probability data; secondly, a visual data chain is generated based on the target image data and the survey text data. By screening and visually displaying the transaction image data, the interference of redundant information is avoided, which helps to improve the accuracy and efficiency of subsequent analysis of the data chain.

[0076] In this embodiment, a method for business risk early warning of an enterprise is provided, and the method includes:

[0077] S110. Obtain multiple frames of transaction image data and survey text data of a target enterprise.

[0078] Among them, the target enterprise can be an enterprise that needs to be investigated by the regulatory department. The transaction image data can be image information generated in the transaction activities of the target enterprise. Any transaction image data corresponds to a target image type. The target image type is used to characterize the complexity of the information in the transaction image data, such as simple or complex. The survey text data can be text information related to the business activities of the target enterprise obtained through investigation.

[0079] In some embodiments, transaction image data of a target enterprise can be obtained from a bank transaction system, such as a chart output by the bank transaction system. In addition, a chart required for investigation can be generated based on the transaction data of the target enterprise obtained from the bank transaction system or the transaction data in the financial records of the target enterprise, and used as the transaction image data.

[0080] In some embodiments, the regulatory department conducts investigations on various business activities of the target enterprise to generate investigation text data, which may include investigation reports, contract documents, financial data, tax data, employee information, partner information, communication record data, etc.

[0081] S120. Invoke a target multimodal model corresponding to the target image type.

[0082] Among them, the target multimodal model can be a model that supports two modalities of image and text. The feature extraction granularity of the target multimodal model matches the complexity of the image content of any transaction image data. The feature extraction granularity can be the fineness of the target multimodal model in extracting image and text features.

[0083] In some embodiments, if the target image type belongs to a complex image type, the corresponding target multimodal model needs to extract more fine-grained features to more accurately capture the detailed information in the transaction image data; if the target image type belongs to a simple image type, the feature granularity extracted by the corresponding target multimodal model can be coarser to reduce the processing complexity and improve the processing performance of the transaction image data.

[0084] S130. Input any transaction image data and investigation text data into the target multimodal model for feature extraction and feature alignment to obtain the graphic-text association probability data between any transaction image data and the investigation text data.

[0085] Among them, transaction image features are obtained by extracting features from any transaction image data. Investigation text features are obtained by extracting features from the investigation text data. Feature alignment can be to map the transaction image features and the investigation text features into a shared feature space, so that the features of the two modalities can complement, compare, and synthesize with each other, thereby improving the analysis effect of multimodal data. The graphic-text association probability data can be the degree of association between the transaction image data and the investigation text data.

[0086] In some embodiments, the target multi-modal model includes an image processing branch and a text processing branch connected in parallel. First, use the image processing branch to extract features from any transaction image data to obtain the transaction image features of the any transaction image data; then use the text processing branch to extract features from the investigation text data to obtain the investigation text features of the investigation text data; then perform feature alignment based on the transaction image features and the investigation text features; finally, compare the similarity between the transaction image features and the investigation text features, which can be to calculate the Euclidean distance or cosine similarity between them, and use this similarity as the graphic-text association probability data.

[0087] S140. Generate a visual data chain based on the graphic-text association probability data, the investigation text data, and multiple frames of transaction image data to give an early warning of the business risks of the target enterprise.

[0088] Among them, the visual data chain can be a data display chain generated by visually displaying the graphic-text association probability data, the investigation text data, and multiple frames of transaction image data, which helps office staff intuitively understand and analyze the association relationships in the data.

[0089] In some embodiments, a visual data chain is generated based on the graphic-text association probability data, the investigation text data, and multiple frames of transaction image data. The visual data chain can be displayed in a graphical manner. Please refer to Figure 1b .. In addition, the visual data chain can also be displayed in the form of charts or other visual forms. Through the visual data chain, complex data relationships can become more intuitive, which helps to improve the analysis and identification efficiency of enterprise business risks and give an early warning of the business risks of the target enterprise in a timely manner.

[0090] In the above embodiments, first obtain multiple frames of transaction image data and investigation text data of the target enterprise; then call the target multi-modal model corresponding to the target image type; then input any transaction image data and investigation text data into the target multi-modal model for feature extraction and feature alignment to obtain the graphic-text association probability data between any transaction image data and investigation text data; finally, generate a visual data chain based on the graphic-text association probability data, the investigation text data, and multiple frames of transaction image data to give an early warning of the business risks of the target enterprise. Based on the transaction image data for different target image types, call the corresponding target multi-modal large model for analysis, which can take into account the accuracy of processing complex transaction image data and the efficiency of processing simple transaction images. Further, based on the target multi-modal large model, perform feature extraction and feature alignment on multiple frames of transaction image data and investigation text data, which can comprehensively analyze data from different sources, realize the comprehensive mining of information, and thus obtain more accurate graphic-text association probability data. Finally, based on the visual data chain, office staff can analyze, judge, and give an early warning of the business risks of the target enterprise more quickly and accurately.

[0091] Please refer to Figure 2 , in some embodiments, a visualization data chain is generated based on the graphic-text association probability data, the survey text data, and the multi-frame transaction image data, including:

[0092] S210. Determine the target image data associated with the survey text data in the multi-frame transaction image data according to the graphic-text association probability data.

[0093] S220. Generate a visualization data chain based on the target image data and the survey text data.

[0094] In some embodiments, a graphic-text association probability threshold is set. When the graphic-text association probability data is higher than the graphic-text association probability threshold, the corresponding target image data is determined as the target image data associated with the survey text data. Then, based on the survey text data and the filtered target image data associated with the survey text data, a visualization data chain can be generated.

[0095] In the above embodiments, first, the target image data associated with the survey text data is determined in the multi-frame transaction image data according to the graphic-text association probability data; second, a visualization data chain is generated based on the target image data and the survey text data. By filtering the transaction image data, the interference of redundant information is avoided, which helps to improve the accuracy and efficiency of subsequent data chain analysis.

[0096] Please refer to Figure 3 , in some embodiments, the first image type is a complex image type. The multi-frame transaction image data includes a transaction network graph and a transaction frequency heat map, and the transaction network graph and the transaction frequency heat map respectively correspond to the complex image type; the complex image type corresponds to the first multi-modal model. The first graphic-text association probability between the first transaction image data and the survey text data will be obtained by performing an association prediction on the first transaction image data and the survey text data through the first multi-modal model, including:

[0097] S310. Input the transaction network graph and the survey text data into the first multi-modal model for feature extraction and feature alignment to obtain the first association probability data between the transaction network graph and the survey text data.

[0098] S320. Input the transaction frequency heat map and the survey text data into the first multi-modal model for feature extraction and feature alignment to obtain the second association probability data between the transaction frequency heat map and the survey text data.

[0099] Among them, the first graphic-text association probability includes first association probability data and second association probability data. The first association probability data can be data used to quantify the relationship between the trading network graph and the survey text data. The second association probability data can be data used to quantify the relationship between the trading frequency heat map and the survey text data. The first multimodal model can be a fine-tuned CLIP model.

[0100] The trading network graph can be a network structure generated based on the fund flow data, related transaction data, and business cooperation relationships between the accounts of the target enterprise and other accounts, and is used to reveal the trading patterns, fund flows, and potential correlations between enterprises. From the details of the trading network graph, abnormal patterns of transactions between accounts, frequent fund transfers, and other unusual trading behaviors can be discovered. These minor changes may imply potential risks, correlations, or non-compliant operations. By deeply analyzing these details, it helps to better identify the business risks of the target enterprise.

[0101] The trading frequency heat map can be image data generated based on the trading time data, trading location data, trading volume data, trading partner data, trading type data (such as sales transactions, payment transactions, refunds, transfers, etc.) of the target enterprise. From the details of the trading frequency heat map, abnormal trading patterns and trends of the target enterprise can be discovered, such as abnormal trading activities frequently occurring in specific time periods or regions, sudden increases in specific accounts or trading types, and trading fluctuations significantly different from historical data.

[0102] In some embodiments, the trading network graph and the survey text data are input into the first multimodal model for feature extraction and feature alignment. By comparing the similarity between the trading image features and the survey text features, association probability data of abnormal fund flows, risk association probability data with other enterprises, potential illegal association probability data between accounts, etc. can be obtained. These association probability data constitute the first association probability data.

[0103] In some embodiments, the trading frequency heat map and the survey text data are input into the first multimodal model for feature extraction and feature alignment. By comparing the similarity between the trading image features and the survey text features, association probability data of abnormal trading activities, association probability data of sudden increases in specific accounts or trading types, association probability data of trading fluctuations, etc. can be obtained. These association probability data constitute the second association probability data.

[0104] In the above embodiments, first, the transaction network graph and the survey text data are input into the first multi-modal model for feature extraction and feature alignment to obtain the first correlation probability data between the transaction network graph and the survey text data; then, the transaction frequency heat map and the survey text data are input into the first multi-modal model for feature extraction and feature alignment to obtain the second correlation probability data between the transaction frequency heat map and the survey text data. By inputting the complex transaction image data and the survey text data into the first multi-modal model for feature extraction and alignment, the corresponding graphic-text correlation probability data can be obtained, thereby realizing the correlation analysis of the complex image data and the survey text data, and providing an accurate data basis for the business risk early warning of the target enterprise.

[0105] Please refer to Figure 4 , in some embodiments, the processing procedures of the transaction network graph and the transaction frequency heat map are the same; inputting the transaction network graph and the survey text data into the first multi-modal model for feature extraction and feature alignment to obtain the first correlation probability data between the transaction network graph and the survey text data includes:

[0106] S410. Perform multi-level feature extraction on the transaction network graph at multiple preset levels to obtain the image granularity features at each preset level.

[0107] S420. Perform multi-level feature extraction on the survey text data at multiple preset levels to obtain the text granularity features at each preset level.

[0108] S430. Perform hierarchical alignment and similarity calculation on the image granularity features and the text granularity features at multiple preset levels to obtain the correlation probability data at each preset level.

[0109] S440. Perform weighted summation based on the correlation probability data at each preset level to obtain the first correlation probability data.

[0110] Among them, the multi-level feature extraction can be to extract features at different levels from the transaction image data and the survey text data respectively, so that the first multi-modal model can deeply understand the transaction image data and the survey text data of complex image types.

[0111] In some embodiments, the first multi-modal model may be a text-image multi-modal model with a preset two-layer hierarchy. In the first layer, the first multi-modal model performs convolutional operations on the transaction network graph using a smaller scale to obtain the first-layer image granularity features; and uses word-based granularity to extract features from the survey text data to obtain the first-layer text granularity features. In the second layer, the first multi-modal model performs convolutional operations on the transaction network graph using a larger scale to obtain the second-layer image granularity features; and uses sentence-based granularity to extract features from the survey text data to obtain the second-layer text granularity features.

[0112] In some embodiments, the first multi-modal model maps the first-layer image granularity features and the first-layer text granularity features to a first feature space for feature alignment, and then performs similarity comparison to obtain the first-layer correlation probability data. Also, the first multi-modal model maps the second-layer image granularity features and the second-layer text granularity features to a second feature space for feature alignment, and then performs similarity comparison to obtain the second-layer correlation probability data.

[0113] In some embodiments, the first-layer correlation probability data corresponds to a first preset weight, and the second-layer correlation probability data corresponds to a second preset weight. The sum obtained by multiplying the first-layer correlation probability data by the first preset weight and adding the product of the second-layer correlation probability data multiplied by the second preset weight is the first correlation probability data.

[0114] It can be understood that the preset hierarchy can also be three layers, four layers or more layers. The more layers there are, the more helpful it is to obtain more accurate first correlation probability data, but the model processing performance will decrease. Therefore, the preset hierarchy should be reasonably set to balance the accuracy and processing performance.

[0115] In the above embodiments, by inputting the transaction network graph and the survey text data into the first multi-modal model for feature extraction and feature alignment, the image granularity features and the text granularity features can be respectively extracted at multiple preset levels, and hierarchical alignment and similarity calculation are performed, so as to obtain the correlation probability data at each level. Finally, by weighted summing the correlation probability data at these levels, the first correlation probability data is generated. This embodiment can effectively fuse complex types of image and text information, enhance the correlation analysis of image and text data, and thus provide more accurate decision-making support for the business risk warning of the target enterprise.

[0116] It should be noted that inputting the transaction frequency heat map and the survey text data into the first multi-modal model for feature extraction and feature alignment to obtain the second correlation probability data between the transaction frequency heat map and the survey text data includes: performing multi-level feature extraction on the transaction frequency heat map at multiple preset levels to obtain image granularity features at each preset level; performing multi-level feature extraction on the survey text data at the multiple preset levels to obtain text granularity features at each preset level; performing hierarchical alignment and similarity calculation on the image granularity features and the text granularity features at the multiple preset levels to obtain correlation probability data at each preset level; and performing weighted summation based on the correlation probability data at each preset level to obtain the second correlation probability data.

[0117] Please refer to Figure 5 , in some embodiments, the second image type is a simple image type; the multi-frame transaction image data includes a transaction timeline graph and a fund flow graph, and the transaction timeline graph and the fund flow graph respectively correspond to the simple image type; the simple image type corresponds to the second multi-modal model; performing correlation prediction on the second transaction image data and the survey text data through the second multi-modal model to obtain the second graphic-text correlation probability between the second transaction image data and the survey text data, including:

[0118] S510. Input the transaction timeline graph and the survey text data into the second multi-modal model for feature extraction and feature alignment to obtain the third correlation probability data between the transaction timeline graph and the survey text data.

[0119] S520. Input the fund flow graph and the survey text data into the second multi-modal model for feature extraction and feature alignment to obtain the fourth correlation probability data between the fund flow graph and the survey text data.

[0120] Among them, the second graphic-text correlation probability includes the third correlation probability data and the fourth correlation probability data. The third correlation probability data can be data used to quantify the relationship between the transaction timeline graph and the survey text data. The fourth correlation probability data can be data used to quantify the relationship between the fund flow graph and the survey text data.

[0121] The transaction timeline graph can show the transaction activities of an enterprise within a specific time period. The transactions in the graph are arranged in chronological order, marking each transaction (such as payment, receipt, deposit, withdrawal, etc.), and can reflect the trend, periodicity or suddenness of the fund flow.

[0122] The fund flow diagram can show the flow path of funds between different accounts or enterprises. Each node represents an account or entity, and each edge represents the flow path and amount of funds. Through this diagram, the source and destination of funds can be reflected.

[0123] In some embodiments, the transaction timeline diagram and the investigation text data are input into the second multimodal model for feature extraction and feature alignment. By comparing the similarity between the transaction image features and the investigation text features, the associated probability data of transaction patterns (such as large - amount fund flows, frequent small - amount transactions, etc.) and the associated probability data of time - node matching (such as special events occurring to the target enterprise in a certain time period in the investigation text data, and a large amount of fund flows on the transaction timeline diagram during this time period) can be obtained. These associated probability data constitute the third associated probability data.

[0124] In some embodiments, the fund flow diagram and the investigation text data are input into the second multimodal model for feature extraction and feature alignment. By comparing the similarity between the transaction image features and the investigation text features, the associated probability data of the relationship between fund flow and enterprise behavior (such as the target enterprise having illegal cross - border transactions) and the associated probability data of the relationship between account behavior and suspect individuals (such as a large amount of fund transactions between the target enterprise and suspect individuals) can be obtained. These associated probability data constitute the fourth associated probability data.

[0125] In the above - mentioned embodiments, first, the transaction timeline diagram and the investigation text data are input into the second multimodal model for feature extraction and feature alignment to obtain the third associated probability data between the transaction timeline diagram and the investigation text data; then, the fund flow diagram and the investigation text data are input into the second multimodal model for feature extraction and feature alignment to obtain the fourth associated probability data between the fund flow diagram and the investigation text data. By inputting simple transaction image data and investigation text data into the second multimodal model for feature extraction and alignment, the corresponding graphic - text associated probability data can be obtained, thereby realizing the correlation analysis of simple image data and investigation text data, and providing an accurate data basis for the business risk warning of the target enterprise.

[0126] Please refer to Figure 6 , in some embodiments, the processing procedures of the transaction timeline diagram and the fund flow diagram are the same; the second multimodal model includes a text encoder and an image encoder in parallel; inputting the transaction timeline diagram and the investigation text data into the second multimodal model for feature extraction and feature alignment to obtain the third associated probability data between the transaction timeline diagram and the investigation text data includes:

[0127] S610. Encode and transform the investigation text data through the text encoder to obtain the text semantic vector of the investigation text data.

[0128] S620. Encode and transform the transaction timeline graph through an image encoder to obtain the image feature vector of the transaction timeline graph.

[0129] S630. Use the contrastive learning method to align and calculate the similarity between the text semantic vector and the image feature vector to generate the third correlation probability data.

[0130] Among them, the contrastive learning method can be a method of learning feature representations by comparing the similarity between the text semantic vector and the image feature vector, which can map the text semantic vector and the image feature vector to a unified feature space.

[0131] In some embodiments, the text encoder can be a model based on the Transformer architecture (such as BERT, GPT), which can encode and transform the survey text data to obtain the text semantic vector of the survey text data. The survey text data is transformed into a vector form for further analysis, and subsequent feature alignment and similarity calculation can be performed. Specifically, the text encoder transforms the vocabulary, sentence structure, and semantic information in the survey text data into a high-dimensional vector, enabling the second multimodal model to understand the information in the survey text data and providing input data for the contrastive learning method.

[0132] In some embodiments, the image encoder can be a convolutional neural network (CNN), which can encode and transform the transaction timeline graph to obtain the image feature vector of the transaction timeline graph. Through the image encoder, the information in the transaction timeline graph is transformed into a feature vector for subsequent feature alignment and similarity calculation. Specifically, the image encoder extracts key features in the transaction timeline graph, such as node relationships, fund flow paths, etc., and transforms this information into a vector representation. Enabling the second multimodal model to understand the information in the transaction timeline graph and providing input data for the contrastive learning method.

[0133] In some embodiments, the second multimodal model includes a contrastive learning module, which can map the text semantic vector and the image feature vector to a high-dimensional feature space through the contrastive learning method, and make similar feature vectors close to each other in this space, while dissimilar ones are far from each other.

[0134] In some embodiments, first input the transaction timeline graph and the survey text data into the second multimodal model for feature extraction to obtain the text semantic vector and the image feature vector; then map these two vectors to a high-dimensional feature space through the contrastive learning module and calculate their similarity (which can be the Euclidean distance or cosine similarity), and finally use this similarity as the third correlation probability data.

[0135] In the above embodiments, by inputting the transaction timeline graph and the survey text data into the second multi-modal model, the text encoder and the image encoder in parallel are used to extract the text semantic vector and the image feature vector respectively, and the above vectors are aligned by the contrastive learning method to generate the third correlation probability data between the transaction timeline graph and the survey text data. This embodiment can effectively fuse simple types of image and text information, enhance the correlation analysis of image and text data, and thus provide more accurate decision-making support for the business risk warning of the target enterprise.

[0136] It should be noted that inputting the fund flow graph and the survey text data into the second multi-modal model for feature extraction and feature alignment to obtain the fourth correlation probability data between the fund flow graph and the survey text data includes: encoding and transforming the survey text data through the text encoder to obtain the text semantic vector of the survey text data; encoding and transforming the fund flow graph through the image encoder to obtain the image feature vector of the transaction timeline graph; using the contrastive learning method to align and perform similarity calculation on the text semantic vector and the image feature vector to generate the fourth correlation probability data.

[0137] Please refer to Figure 7 , in some embodiments, the target image type corresponding to any transaction image data is determined by the following method:

[0138] S710. Evaluate the complexity of any transaction image data to obtain the image complexity evaluation data of any transaction image data.

[0139] S720. Determine the target image type according to the target complexity range where the image complexity evaluation data is located and the complexity type relationship data.

[0140] Among them, the target image type includes the first image type and the second image type. The complexity type relationship data is used to describe the corresponding relationship between the complexity range and the image type of the transaction image data.

[0141] In some embodiments, the complexity of transaction image data can be evaluated by comprehensively evaluating texture complexity, object quantity, background noise detection, and image resolution. Exemplarily, the texture complexity of transaction image data can be evaluated by calculating texture features (such as gray-level co-occurrence matrix, LBP (Local Binary Pattern)). Generally, the higher the texture, the more complex the image content. Exemplarily, the number and distribution of objects in transaction image data can be identified through image segmentation or object detection algorithms (such as YOLOV8, Faster R-CNN, etc.). The more objects and the more overlapping regions usually mean higher complexity. Exemplarily, the noise components of transaction image data can be evaluated using the frequency domain features of the image (such as Fourier transform or wavelet transform). The more noise, the more complex the image. Exemplarily, the resolution of transaction image data can be determined by the number of pixels in width and height. When the resolution is higher, it usually contains more information and higher complexity.

[0142] In some embodiments, the image complexity evaluation data of any transaction image data is obtained by a scoring method. Specifically, the texture complexity evaluation, object quantity evaluation, background noise detection evaluation, and image resolution evaluation all have preset thresholds. When the evaluation result of each item is higher than the preset threshold, the score is recorded as 1, otherwise it is recorded as 0. The total score of the four evaluations is the image complexity evaluation data. Exemplarily, the complexity type relationship data is: when the image complexity evaluation data is greater than or equal to 3, the corresponding target image type is a complex image type, otherwise it is a simple image type. For an example of the complexity evaluation of transaction image data, please refer to Table 1.

[0143] Table 1 Example Table of Complexity Evaluation of Transaction Image Data

[0144]

[0145] In the above embodiments, by evaluating the complexity of transaction image data and combining the complexity type relationship data, the target image type of the transaction image data can be accurately determined. This embodiment can effectively distinguish transaction images of different complexities, thereby providing an accurate image type basis for subsequent calls to the target multi-modal model for analysis.

[0146] An embodiment of the present application also provides an enterprise operation risk warning device, which includes:

[0147] A data acquisition module for acquiring multiple frames of transaction image data and investigation text data of a target enterprise; wherein, the multiple frames of transaction image data include first transaction image data belonging to a first image type and second transaction image data belonging to a second image type; the first image type corresponds to a first multi-modal model, the second image type corresponds to a second multi-modal model, and the first multi-modal model and the second multi-modal model are in parallel;

[0148] The first data processing module is used to perform correlation prediction on the first transaction image data and the survey text data through the first multi-modal model, and obtain the first graphic-text correlation probability between the first transaction image data and the survey text data; wherein, the first multi-modal model uses a convolutional neural network with an adjustable dilation rate to globally extract the microscopic features of the first granularity of the first transaction image data;

[0149] The second data processing module is used to perform correlation prediction on the second transaction image data and the survey text data through the second multi-modal model, and obtain the second graphic-text correlation probability between the second transaction image data and the survey text data; wherein, the second multi-modal model uses a regional attention mechanism to also globally extract the macroscopic features of the second granularity of the second transaction image data, and the first granularity is smaller than the second granularity;

[0150] The data chain generation module is used to generate a visual data chain according to the first graphic-text correlation probability, the second graphic-text correlation probability, the survey text data and the multi-frame transaction image data, so as to perform an operating risk warning on the target enterprise.

[0151] Please refer to Figure 8 , an embodiment of the present application further provides an enterprise operating risk warning device 800, and the enterprise operating risk warning device 800 includes:

[0152] The data acquisition module 810 is used to acquire multi-frame transaction image data and survey text data of a target enterprise; wherein, any transaction image data corresponds to a target image type;

[0153] The model calling module 820 is used to call a target multi-modal model corresponding to the target image type; wherein, the feature extraction granularity of the target multi-modal model matches the image content complexity of any transaction image data;

[0154] The data processing module 830 is used to input any transaction image data and survey text data into the target multi-modal model for feature extraction and feature alignment, and obtain the graphic-text correlation probability data between any transaction image data and survey text data;

[0155] The data chain generation module 840 is used to generate a visual data chain according to the graphic-text correlation probability data, the survey text data and the multi-frame transaction image data, so as to perform an operating risk warning on the target enterprise.

[0156] In some embodiments, the data chain generation module 840 includes:

[0157] An image data determination unit, configured to determine target image data associated with the investigation text data from multiple frames of transaction image data according to the text-image association probability data;

[0158] A data chain generation unit, configured to generate a visualization data chain based on the target image data and the investigation text data.

[0159] In some embodiments, the first image type is a complex image type, the multiple frames of transaction image data include a transaction network graph and a transaction frequency heat map, and the transaction network graph and the transaction frequency heat map respectively correspond to the complex image type; the complex image type corresponds to a first multimodal model, and the data processing module 830 includes:

[0160] A first probability data acquisition unit, configured to input the transaction network graph and the investigation text data into the first multimodal model for feature extraction and feature alignment to obtain first association probability data between the transaction network graph and the investigation text data; and configured to input the transaction frequency heat map and the investigation text data into the first multimodal model for feature extraction and feature alignment to obtain second association probability data between the transaction frequency heat map and the investigation text data; wherein, the first text-image association probability includes the first association probability data and the second association probability data.

[0161] In some embodiments, the processing procedures of the transaction network graph and the transaction frequency heat map are the same, and the first probability data acquisition unit includes:

[0162] A feature extraction sub-unit, configured to perform multi-level feature extraction on the transaction network graph at multiple preset levels to obtain image granularity features at each preset level; and configured to perform multi-level feature extraction on the investigation text data at multiple preset levels to obtain text granularity features at each preset level;

[0163] An alignment and calculation sub-unit, configured to perform hierarchical alignment and similarity calculation on the image granularity features and the text granularity features at multiple preset levels to obtain association probability data at each preset level;

[0164] A weighted summation sub-unit, configured to perform weighted summation based on the association probability data at each preset level to obtain the first association probability data.

[0165] In some embodiments, the second image type is a simple image type, the multiple frames of transaction image data include a transaction timeline graph and a fund flow graph, and the transaction timeline graph and the fund flow graph respectively correspond to the simple image type; the simple image type corresponds to a second multimodal model; the data processing module 830 further includes:

[0166] A second probability data acquisition unit, configured to input the transaction timeline graph and the survey text data into a second multi-modal model for feature extraction and feature alignment, so as to obtain third correlation probability data between the transaction timeline graph and the survey text data; and configured to input the fund flow graph and the survey text data into the second multi-modal model for feature extraction and feature alignment, so as to obtain fourth correlation probability data between the fund flow graph and the survey text data; wherein, the second graphic-text correlation probability includes the third correlation probability data and the fourth correlation probability data.

[0167] In some embodiments, the processing processes of the transaction timeline graph and the fund flow graph are the same; the second multi-modal model includes a text encoder and an image encoder connected in parallel; the second probability data acquisition unit further includes:

[0168] An encoding subunit, configured to encode and transform the survey text data through the text encoder to obtain a text semantic vector of the survey text data; and configured to encode and transform the transaction timeline graph through the image encoder to obtain an image feature vector of the transaction timeline graph;

[0169] An alignment and calculation subunit, configured to use the contrastive learning method to align and calculate the text semantic vector and the image feature vector to generate the third correlation probability data.

[0170] In some embodiments, the enterprise operation risk warning device 800 further includes:

[0171] A complexity evaluation module, configured to evaluate the complexity of any transaction image data to obtain image complexity evaluation data of any transaction image data;

[0172] An image type determination module, configured to determine a target image type according to the target complexity range where the image complexity evaluation data is located and the complexity type relationship data; wherein, the complexity type relationship data is used to describe the correspondence between the complexity range and the image type of the transaction image data. The target image type includes a first image type and a second image type.

[0173] The further function descriptions of the above-mentioned various modules and units are the same as those in the corresponding above-mentioned embodiments, and will not be elaborated here.

[0174] The enterprise operation risk warning device in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0175] Please refer to Figure 9 , Figure 9is a schematic diagram of the structure of a computer device provided in an embodiment of the present application, such as Figure 9 As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 9 A processor 10 is taken as an example.

[0176] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.

[0177] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.

[0178] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0179] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.

[0180] The computer device further includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30, and the output device 20 may be connected through a bus or other means. Figure 9 Taking the connection through the bus as an example.

[0181] The input device 30 can receive input digital or character information and generate key signal inputs related to the user settings and function controls of the computer device, such as a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED), and a haptic feedback device (e.g., a vibration motor), etc. The above display device includes, but is not limited to, a liquid crystal display, a light-emitting diode, a display, and a plasma display. In some alternative embodiments, the display device may be a touch screen.

[0182] The embodiment of the present application also provides a computer-readable storage medium. The method according to the embodiment of the present application can be implemented in hardware, firmware, or be implemented as computer code that can be recorded on a storage medium, or be implemented as computer code that is originally stored in a remote storage medium or a non-transitory machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiment is implemented.

[0183] The embodiment of the present application provides a computer program product. The computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method of any embodiment of the present application.

[0184] Although the embodiments of the present application are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations all fall within the scope defined by the appended claims.

[0185] The systems, devices, modules or units illustrated in the above embodiments can be specifically implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer can be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or any combination of these devices.

[0186] For the convenience of description, the above devices are described by function as various units respectively. Of course, when implementing the present application, the functions of each unit can be implemented in one or more software and / or hardware.

[0187] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0188] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce a device for implementing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0189] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including an instruction device that implements the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0190] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one flow or more flows of the flowchart and / or one block or more blocks of the block diagram.

[0191] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the said element.

[0192] Each embodiment in this specification is described in a progressive manner, and the same or similar parts among the embodiments can be referred to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0193] The above are only the embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

[0194] Although the embodiments of the present application are described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations fall within the scope defined by the appended claims.

Claims

1. An enterprise operation risk early warning method, characterized in that, The method includes: Obtaining multi-frame transaction image data and investigation text data of a target enterprise; wherein, the multi-frame transaction image data includes first transaction image data belonging to a first image type and second transaction image data belonging to a second image type; the first image type corresponds to a first multi-modal model, the second image type corresponds to a second multi-modal model, and the first multi-modal model and the second multi-modal model are in parallel; Performing correlation prediction on the first transaction image data and the investigation text data through the first multi-modal model to obtain a first text-image correlation probability between the first transaction image data and the investigation text data, including: inputting the first transaction image data and the investigation text data into the first multi-modal model for feature extraction to obtain first transaction image features and investigation text features; performing feature alignment by mapping the first transaction image features and the investigation text features to the same feature space; comparing the similarity between the first transaction image features and the investigation text features, and taking the similarity as the first text-image correlation probability; wherein, the first multi-modal model uses a convolutional neural network with adjustable dilation rate to globally extract microscopic features of a first granularity of the first transaction image data; Performing correlation prediction on the second transaction image data and the investigation text data through the second multi-modal model to obtain a second text-image correlation probability between the second transaction image data and the investigation text data, including: inputting the second transaction image data and the investigation text data into the second multi-modal model for feature extraction to obtain second transaction image features and investigation text features; performing feature alignment by mapping the second transaction image features and the investigation text features to the same feature space; comparing the similarity between the second transaction image features and the investigation text features, and taking the similarity as the second text-image correlation probability; wherein, the second multi-modal model uses a regional attention mechanism to also globally extract macroscopic features of a second granularity of the second transaction image data, and the first granularity is smaller than the second granularity; Generating a visualization data chain according to the first text-image correlation probability, the second text-image correlation probability, the investigation text data, and the multi-frame transaction image data to perform an operating risk warning on the target enterprise.

2. The method according to claim 1, characterized in that, The generating a visualization data chain according to the first text-image correlation probability, the second text-image correlation probability, the investigation text data, and the multi-frame transaction image data to perform an operating risk warning on the target enterprise includes: Determining target image data associated with the investigation text data in the multi-frame transaction image data according to the first text-image correlation probability and the second text-image correlation probability; Generating the visualization data chain based on the target image data and the investigation text data.

3. The method according to claim 1, wherein The first image type is a complex image type; the multi-frame transaction image data includes a transaction network graph and a transaction frequency heat map, and the transaction network graph and the transaction frequency heat map respectively correspond to the complex image type; the complex image type corresponds to the first multi-modal model; Performing correlation prediction on the first transaction image data and the investigation text data through the first multimodal model to obtain a first graphic-text correlation probability between the first transaction image data and the investigation text data, including: Inputting the transaction network graph and the investigation text data into the first multimodal model for feature extraction and feature alignment to obtain first correlation probability data between the transaction network graph and the investigation text data; Inputting the transaction frequency heat map and the investigation text data into the first multimodal model for feature extraction and feature alignment to obtain second correlation probability data between the transaction frequency heat map and the investigation text data; wherein, the first graphic-text correlation probability includes the first correlation probability data and the second correlation probability data.

4. The method according to claim 3, characterized in that, The processing procedures for the transaction network graph and the transaction frequency heat map are the same; The step of inputting the transaction network graph and the investigation text data into the first multimodal model for feature extraction and feature alignment to obtain first correlation probability data between the transaction network graph and the investigation text data includes: Performing multi-level feature extraction on the transaction network graph at multiple preset levels to obtain image granularity features at each preset level; Performing multi-level feature extraction on the investigation text data at the multiple preset levels to obtain text granularity features at each preset level; Performing hierarchical alignment and similarity calculation on the image granularity features and the text granularity features at the multiple preset levels to obtain correlation probability data at each preset level; Performing weighted summation based on the correlation probability data at each preset level to obtain the first correlation probability data.

5. The method according to claim 1, wherein The second image type is a simple image type; the multi-frame transaction image data includes a transaction timeline graph and a fund flow graph, and the transaction timeline graph and the fund flow graph respectively correspond to the simple image type; the simple image type corresponds to a second multimodal model; Performing correlation prediction on the second transaction image data and the investigation text data through the second multimodal model to obtain a second graphic-text correlation probability between the second transaction image data and the investigation text data, including: Inputting the transaction timeline graph and the investigation text data into the second multimodal model for feature extraction and feature alignment to obtain third correlation probability data between the transaction timeline graph and the investigation text data; Inputting the fund flow graph and the investigation text data into the second multimodal model for feature extraction and feature alignment to obtain fourth correlation probability data between the fund flow graph and the investigation text data; wherein, the second graphic-text correlation probability includes the third correlation probability data and the fourth correlation probability data.

6. The method according to claim 5, wherein The processing procedures for the transaction timeline graph and the fund flow graph are the same; the second multimodal model includes a text encoder and an image encoder connected in parallel; Inputting the transaction timeline graph and the survey text data into the second multi-modal model for feature extraction and feature alignment to obtain third correlation probability data between the transaction timeline graph and the survey text data includes: Encoding and transforming the survey text data through the text encoder to obtain a text semantic vector of the survey text data; Encoding and transforming the transaction timeline graph through the image encoder to obtain an image feature vector of the transaction timeline graph; Using the contrastive learning method to align and calculate the similarity between the text semantic vector and the image feature vector to generate the third correlation probability data.

7. According to the method described in any one of claims 1 to 6, characterized in that, Determining the target image type corresponding to any transaction image data through the following method: Comprehensively evaluating the complexity of any transaction image data in terms of texture complexity evaluation, object quantity evaluation, background noise detection evaluation, and image resolution to obtain image complexity evaluation data of any transaction image data; Determining the target image type according to the target complexity range where the image complexity evaluation data is located and the complexity type relationship data; wherein, the complexity type relationship data is used to describe the corresponding relationship between the complexity range and the image type of the transaction image data; the target image type includes a first image type and a second image type.

8. An enterprise operation risk warning device, characterized in that, The device includes: A data acquisition module for acquiring multiple frames of transaction image data and survey text data of a target enterprise; wherein, the multiple frames of transaction image data include first transaction image data belonging to the first image type and second transaction image data belonging to the second image type; the first image type corresponds to a first multi-modal model, the second image type corresponds to a second multi-modal model, and the first multi-modal model and the second multi-modal model are in parallel; A first data processing module for performing correlation prediction on the first transaction image data and the survey text data through the first multi-modal model to obtain a first graphic-text correlation probability between the first transaction image data and the survey text data, including: inputting the first transaction image data and the survey text data into the first multi-modal model for feature extraction to obtain first transaction image features and survey text features; performing feature alignment by mapping the first transaction image features and the survey text features to the same feature space; comparing the similarity between the first transaction image features and the survey text features and using the similarity as the first graphic-text correlation probability; wherein, the first multi-modal model uses a convolutional neural network with an adjustable dilation rate to globally extract microscopic features of a first granularity of the first transaction image data; A second data processing module, configured to perform an association prediction on the second transaction image data and the investigation text data through the second multi-modal model, and obtain a second text-image association probability between the second transaction image data and the investigation text data, including: inputting the second transaction image data and the investigation text data into the second multi-modal model for feature extraction to obtain second transaction image features and investigation text features; performing feature alignment by mapping the second transaction image features and the investigation text features into the same feature space; comparing the similarity between the second transaction image features and the investigation text features, and using the similarity as the second text-image association probability; wherein, the second multi-modal model adopts a regional attention mechanism to also extract macroscopic features of a second granularity of the second transaction image data globally, and the first granularity is smaller than the second granularity; A data chain generation module, configured to generate a visualization data chain according to the first text-image association probability, the second text-image association probability, the investigation text data, and the multi-frame transaction image data, so as to perform an operation risk warning on the target enterprise.

9. A computer device, characterized in that, including: A memory and a processor, which are communicatively connected to each other. The memory stores computer instructions, and the processor executes the computer instructions to execute the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Multi-modal image-text matching model and construction method, device and application thereof

    CN115935199A

  • Service report generation method and device, equipment and medium

    CN119559297A