Enterprise operation risk early warning method and device, computer equipment and medium
By using multimodal models to extract and correlate the enterprise's transaction image data and survey text data, visual data links are generated, which solves the problem of insufficient accuracy of enterprise operating risk warning in the existing technology, and achieves more accurate risk identification and early warning.
Patent Information
- Application Number
- CN202510455401.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-11
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-04-11
AI Technical Summary
The accuracy of enterprise operation risk warning in the prior art needs to be improved, and it is difficult to effectively identify and prevent potential risks in enterprise operation.
By obtaining the multi-frame transaction image data and survey text data of the target enterprise, a multi-modal model matching the image type is called for feature extraction and feature alignment, a graph-text correlation probability data is generated, and a visual data link is generated based on this to perform risk warning.
It improves the accuracy of enterprise operating risk warning, can more accurately identify and prevent potential risks, and provides more detailed analysis and decision-making support.
Smart Images

Figure CN119990782A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a method, device, computer equipment and medium for early warning of business risk of an enterprise. Background Art
[0002] Enterprise operation risk warning can timely identify potential operation risks, help regulatory authorities take effective measures to prevent illegal activities, and avoid financial risks caused by enterprise operation problems. However, the accuracy of enterprise operation risk warning in related technologies needs to be improved.
[0003] Therefore, it is urgent to propose a new early warning method for enterprise operation risks. Summary of the invention
[0004] The present application provides a method, device, computer equipment and medium for early warning of business risk of an enterprise, which solves the technical problem in the related technology that the accuracy of early warning of business risk of an enterprise needs to be improved, and achieves the technical effect of more accurate early warning of business risk of an enterprise.
[0005] In order to achieve the above objectives, the main technical solutions adopted in this application include: In a first aspect, an embodiment of the present application provides a method for early warning of business risk of an enterprise, the method comprising: Acquire multiple frames of transaction image data and survey text data of a target enterprise; wherein any transaction image data corresponds to a target image type; Calling a target multimodal model corresponding to the target image type; wherein the feature extraction granularity of the target multimodal model matches the image content complexity of any transaction image data; Inputting the any transaction image data and the survey text data into the target multimodal model for feature extraction and feature alignment, and obtaining image-text association probability data between the any transaction image data and the survey text data; A visual data chain is generated based on the image-text association probability data, the survey text data and the multi-frame transaction image data to provide an early warning of business risks for the target enterprise.
[0006] In the embodiment of the present application, based on the transaction image data for different image types, the corresponding multimodal large model is called for analysis, which can take into account the accuracy of complex transaction image data processing and the efficiency of simple transaction image processing. Furthermore, based on the first multimodal model, a convolutional neural network with adjustable void rate is used to extract the microscopic features of the first granularity of the first transaction image data globally, and then the association prediction is performed with the survey text data to obtain a more accurate first image-text association probability. Based on the second multimodal model, a regional attention mechanism is used to extract the macroscopic features of the second granularity of the second transaction image data globally, and then the association prediction is performed with the survey text data to obtain a more accurate second image-text association probability. Finally, a visual data chain is generated based on the first image-text association probability, the second image-text association probability, the survey text data and the transaction image data, which facilitates office personnel to analyze, judge and warn the target enterprise's business risks more quickly and accurately.
[0007] In a second aspect, an embodiment of the present application provides a method for early warning of business risk of an enterprise, the method comprising: Acquire multiple frames of transaction image data and survey text data of a target enterprise; wherein the multiple frames of transaction image data include first transaction image data belonging to a first image type and second transaction image data belonging to a second image type; the first image type corresponds to a first multimodal model, the second image type corresponds to a second multimodal model, and the first multimodal model is connected in parallel with the second multimodal model; The first multimodal model is used to perform association prediction on the first transaction image data and the survey text data to obtain a first image-text association probability between the first transaction image data and the survey text data; wherein the first multimodal model uses a convolutional neural network with an adjustable dilation rate to globally extract microscopic features of a first granularity of the first transaction image data; The second multimodal model is used to perform association prediction on the second transaction image data and the survey text data to obtain a second image-text association probability between the second transaction image data and the survey text data; wherein the second multimodal model uses a regional attention mechanism to globally extract macro features of a second granularity of the second transaction image data, and the first granularity is smaller than the second granularity; A visual data chain is generated based on the first image-text association probability, the second image-text association probability, the survey text data and the multiple frames of transaction image data to provide an operational risk warning for the target enterprise.
[0008] In a third aspect, an embodiment of the present application provides an enterprise operation risk early warning device, the device comprising: A data acquisition module is used to acquire multiple frames of transaction image data and survey text data of a target enterprise; wherein any transaction image data corresponds to a target image type; A model calling module, used to call a target multimodal model corresponding to the target image type; wherein the feature extraction granularity of the target multimodal model matches the image content complexity of any transaction image data; A data processing module, used for inputting the any transaction image data and the survey text data into the target multimodal model to perform feature extraction and feature alignment, and obtaining image-text association probability data between the any transaction image data and the survey text data; The data link generation module is used to generate a visual data link based on the image-text association probability data, the survey text data and the multi-frame transaction image data to provide an early warning of the operating risks of the target enterprise.
[0009] In a fourth aspect, an embodiment of the present application provides an enterprise operation risk early warning device, the device comprising: A data acquisition module, used to acquire multiple frames of transaction image data and survey text data of a target enterprise; wherein the multiple frames of transaction image data include first transaction image data belonging to a first image type and second transaction image data belonging to a second image type; the complexity of the first transaction image data is greater than the complexity of the second transaction image data; a first data processing module, configured to call a first multimodal model corresponding to the first image type, and perform association prediction on the first transaction image data and the survey text data through the first multimodal model to obtain a first image-text association probability between the first transaction image data and the survey text data; wherein the feature extraction granularity of the first multimodal model matches the image content complexity of the first transaction image data; a second data processing module, configured to call a second multimodal model corresponding to the second image type, and perform association prediction on the second transaction image data and the survey text data through the second multimodal model to obtain a second image-text association probability between the second transaction image data and the survey text data; wherein the feature extraction granularity of the second multimodal model matches the image content complexity of the second transaction image data; The data link generation module is used to generate a visual data link based on the first image-text association probability, the second image-text association probability, the survey text data and the multi-frame transaction image data to provide an operational risk warning for the target enterprise.
[0010] In a fifth aspect, an embodiment of the present application provides a computer device, including: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method described in any of the above embodiments by executing the computer instructions.
[0011] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to enable a computer to execute the method described in any of the above embodiments.
[0012] In the embodiment of the present application, firstly, the multi-frame transaction image data and the survey text data of the target enterprise are obtained; then the target multimodal model corresponding to the target image type is called; then any transaction image data and the survey text data are input into the target multimodal model for feature extraction and feature alignment, and the image-text association probability data between any transaction image data and the survey text data is obtained; finally, a visual data chain is generated based on the image-text association probability data, the survey text data and the multi-frame transaction image data to provide an operating risk warning for the target enterprise. Based on the target multimodal large model, feature extraction and feature alignment of the multi-frame transaction image data and the survey text data can be performed, and data from different sources can be comprehensively analyzed to achieve comprehensive information mining, thereby obtaining more accurate image-text association probability data and improving the accuracy of risk warning. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0014] Figure 1a A flow chart of a business risk early warning method provided in an embodiment of this specification; Figure 1b A schematic diagram of a visual data link provided in an embodiment of this specification; Figure 2 A flow chart of a business risk early warning method provided in an embodiment of this specification; Figure 3 A flow chart of a business risk early warning method provided in an embodiment of this specification; Figure 4 A flow chart of a business risk early warning method provided in an embodiment of this specification; Figure 5 A flow chart of a business risk early warning method provided in an embodiment of this specification; Figure 6A flow chart of a business risk early warning method provided in an embodiment of this specification; Figure 7 A flow chart of a business risk early warning method provided in an embodiment of this specification; Figure 8 A schematic diagram of a business risk early warning device provided in an embodiment of this specification; Fig. 9 A schematic diagram of a computer structure provided in an embodiment of this specification. DETAILED DESCRIPTION
[0015] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present application.
[0016] Business risk warning can timely identify potential business risks, provide effective decision-making support for regulatory authorities, and help them take targeted measures to prevent and curb illegal activities, and avoid negative impacts such as financial risks and market fluctuations caused by business problems. As the economic environment becomes increasingly complex, the accuracy of business risk warnings in related technologies needs to be improved to better cope with the increasing risk challenges.
[0017] Based on this, the present application proposes a method for early warning of business risk of an enterprise, firstly obtaining multi-frame transaction image data and survey text data of a target enterprise; then calling a target multimodal model corresponding to the target image type; then inputting any transaction image data and survey text data into the target multimodal model for feature extraction and feature alignment, and obtaining the image-text association probability data between any transaction image data and survey text data; finally generating a visual data chain based on the image-text association probability data, survey text data and multi-frame transaction image data, so as to provide early warning of business risk for the target enterprise. Based on the transaction image data for different target image types, calling the corresponding target multimodal large model for analysis can take into account the accuracy of processing complex transaction image data and the efficiency of processing simple transaction images. Furthermore, based on the target multimodal large model, feature extraction and feature alignment are performed on multi-frame transaction image data and survey text data, and data from different sources can be comprehensively analyzed to realize comprehensive information mining, thereby obtaining more accurate image-text association probability data. Finally, based on the visual data chain, office staff can analyze, judge and warn the business risk of the target enterprise more quickly and accurately.
[0018] Furthermore, in some embodiments, the enterprise business risk warning method may include: acquiring multiple frames of transaction image data and survey text data of a target enterprise; determining a target image type to which any transaction image data belongs according to the image content complexity of any transaction image data; inputting the any transaction image data and the survey text data into a target multimodal model that matches the target image type to perform image-text association prediction, and obtain image-text association probability data between the any transaction image data and the survey text data; selecting multiple frames of candidate transaction image data that meet the first-level screening requirements from the multiple frames of transaction image data according to the image-text association probability data; performing consistency verification based on any candidate transaction image data and the structured business data of the target enterprise to select target image data that meet the second-level screening requirements from the multiple frames of candidate transaction image data; generating a visual data chain based on the target image data and the survey text data to provide a business risk warning for the target enterprise.
[0019] It should be noted that the first-level screening requirements may be requirements set for the association between images and texts, and the second-level screening requirements may be requirements set for the consistency between transaction images and enterprise business data. The image-text association probability data (such as the first image-text association probability and the second image-text association probability) are used to select candidate transaction image data associated with the image and text of the survey text data at the first level. Further, for the accuracy of information, the structured business data obtained by structured processing of enterprise business data is checked for consistency with the candidate transaction image data to ensure the accuracy of the data and improve the credibility of the visual data chain.
[0020] According to an embodiment of the present application, an embodiment of a method for early warning of business risks of an enterprise is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0021] See also Figure 1a In this embodiment, a method for early warning of business risk of an enterprise is provided, and the method includes: S101. Acquire multiple frames of transaction image data and survey text data of a target enterprise.
[0022] The multi-frame transaction image data includes first transaction image data belonging to a first image type and second transaction image data belonging to a second image type. The complexity of the first transaction image data is greater than the complexity of the second transaction image data. The first image type and the second image type are used to characterize different complexity levels of information in the transaction image data, such as simple and complex. The first image type corresponds to a first multimodal model, the second image type corresponds to a second multimodal model, and the first multimodal model is connected in parallel with the second multimodal model; In some embodiments, the transaction image data of the target enterprise may be obtained from the bank transaction system, including the first transaction image data belonging to the first image type and the second transaction image data belonging to the second image type. In other embodiments, the first transaction image data belonging to the first image type and the second transaction image data belonging to the second image type may be generated based on the transaction data of the target enterprise obtained from the bank transaction system or the transaction data in the financial records of the target enterprise.
[0023] In some implementations, the regulatory authorities investigate various business activities of the target enterprise and generate investigation text data, which may include investigation reports, contract documents, financial data, tax data, employee information, partner information, communication record data, etc.
[0024] In some embodiments, both the first multimodal model and the second multimodal model can be multimodal models that process image and text data. The first multimodal model can process the first transaction image data and the survey text data; the second multimodal model can process the second transaction image data and the survey text data. The first multimodal model and the second multimodal model are connected in parallel to simultaneously process the corresponding types of transaction images.
[0025] S103: Perform association prediction on the first transaction image data and the survey text data using a first multimodal model to obtain a first image-text association probability between the first transaction image data and the survey text data.
[0026] The first multimodal model uses a convolutional neural network with adjustable void rate to globally extract microscopic features of the first granularity of the first transaction image data. The feature extraction granularity of the first multimodal model matches the image content complexity of the first transaction image data. The first image-text association probability can be data used to quantify the relationship between the first transaction image data and the survey text data.
[0027] In some implementations, the first transaction image data and the survey text data are input into the first multimodal model for feature extraction, the feature extraction granularity matches the image content complexity of the first transaction image data, and the first transaction image features and the survey text features are obtained, and then the first transaction image features and the survey text features are mapped to the same feature space to perform feature alignment. Finally, the similarity of the first transaction image features and the survey text features is compared (which may be Euclidean distance or cosine similarity), and the similarity is used as the first image-text association probability.
[0028] In some embodiments, the first multimodal model may use a convolutional neural network with an adjustable hole rate to globally extract microscopic features of the first granularity of the first transaction image data. Specifically, the first multimodal model realizes the extraction of microscopic features of different scales by dynamically adjusting the spacing distance of the convolution kernel (for example, adjusting the hole rate from 1 to 3). Exemplarily, when the first transaction image data is relatively complex, different hole rates may be used according to the number of detail features of different regions of the first granularity in the image. When the detail features of a certain region are lower than a preset threshold, a larger hole rate is used; when the detail features of a certain region are higher than a preset threshold, a larger hole rate is used. By using a convolutional neural network with an adjustable hole rate in the first multimodal model, the accuracy of different regions can be taken into account for feature extraction of the first transaction image data. It can not only extract relatively rough features in regions with fewer detail features through a high hole rate, but also extract sufficiently fine features in regions with more detail features through a low hole rate, thereby ensuring the effectiveness of feature extraction and helping to obtain a more accurate first image-text association probability.
[0029] S105 . Perform association prediction on the second transaction image data and the survey text data by using a second multimodal model to obtain a second image-text association probability between the second transaction image data and the survey text data.
[0030] The second multimodal model adopts the regional attention mechanism to globally extract the macro features of the second granularity of the second transaction image data, and the first granularity is smaller than the second granularity. The feature extraction granularity of the second multimodal model matches the image content complexity of the second transaction image data. The second image-text association probability can be data used to quantify the relationship between the second transaction image data and the survey text data.
[0031] The feature extraction granularity of the second multimodal model matches the image content complexity of the second transaction image data. The second image-text association probability may be data used to quantify the relationship between the second transaction image data and the survey text data.
[0032] In some implementations, the second transaction image data and the survey text data are input into the second multimodal model for feature extraction, and the feature extraction granularity matches the image content complexity of the second transaction image data to obtain the second transaction image features and the survey text features, and then feature alignment is performed by mapping the second transaction image features and the survey text features to the same feature space. Finally, the similarity (which may be the Euclidean distance or the cosine similarity) of the second transaction image features and the survey text features is compared, and the similarity is used as the second image-text association probability.
[0033] In some embodiments, the second multimodal model can use a regional attention mechanism to globally extract macro features of the second granularity of the second transaction image data. Specifically, the second multimodal model automatically focuses on macro features of semantic importance in the second transaction image by calculating the association weights between the image area of the second granularity and the keywords in the survey text data. Exemplarily, when the second transaction image data is relatively simple, the second multimodal model uses a regional attention mechanism to generate an attention heat map through a spatial transformation network, divides the image into 10×10 grid units, and then screens out regions with a correlation with keywords in the text data that is higher than a threshold (such as 0.7) for feature enhancement. By adopting a regional attention mechanism, the second multimodal model improves the model's ability to recognize the semantic correlation between images and texts, improves the accuracy of feature extraction of the second transaction image data, and helps to obtain a more accurate second image-text association probability.
[0034] S107. Generate a visual data chain based on the first image-text association probability, the second image-text association probability, the survey text data, and the multi-frame transaction image data to provide an early warning of business risks for the target enterprise.
[0035] Among them, the visual data chain can be a data display chain generated by visually displaying the first image-text association probability, the second image-text association probability, the survey text data and multiple frames of transaction image data, which helps office personnel to intuitively understand and analyze the association relationship between transaction image data and survey text data.
[0036] In some embodiments, a visual data chain is generated based on the first image-text association probability, the second image-text association probability, the survey text data, and the multi-frame transaction image data, and the visual data chain can be displayed in a graphical manner. In addition, the visual data chain can also be displayed in a chart or other visual form. Through the visual data chain, complex data relationships can become more intuitive, which helps to improve the efficiency of analyzing and identifying business risks of enterprises, and timely provide business risk warnings to target enterprises.
[0037] In some embodiments, first associated image data associated with the investigation text data is selected from multiple frames of first transaction image data according to the first image-text association probability. Second associated image data associated with the investigation text data is selected from multiple frames of second transaction image data according to the second image-text association probability. The first associated image data, the second associated image data, and the investigation text data are used for visual display to obtain a visual data chain to provide an early warning of business risks for the target enterprise.
[0038] In the above embodiment, based on the transaction image data for different image types, the corresponding multimodal large model is called for analysis, which can take into account the accuracy of complex transaction image data processing and the efficiency of simple transaction image processing. Further, based on the first multimodal model, a convolutional neural network with adjustable void rate is used to extract the microscopic features of the first granularity of the first transaction image data globally, and then the association prediction is performed with the survey text data to obtain a more accurate first image-text association probability. Based on the second multimodal model, a regional attention mechanism is used to extract the macroscopic features of the second granularity of the second transaction image data globally, and then the association prediction is performed with the survey text data to obtain a more accurate second image-text association probability. Finally, a visual data chain is generated based on the first image-text association probability, the second image-text association probability, the survey text data and the transaction image data, which facilitates office staff to analyze, judge and warn the operating risks of the target enterprise more quickly and accurately.
[0039] In some embodiments, a visual data chain is generated based on a first image-text association probability, a second image-text association probability, survey text data, and multiple frames of transaction image data to provide business risk warnings for target enterprises, including: determining target image data associated with the survey text data in multiple frames of transaction image data based on the first image-text association probability and the second image-text association probability; and generating a visual data chain based on the target image data and the survey text data.
[0040] The transaction image data may be image information generated in the transaction activities of the target enterprise. The survey text data may be text information related to the business activities of the target enterprise obtained through survey.
[0041] In some embodiments, the target enterprise's transaction image data may be obtained from a bank transaction system, such as a chart output by the bank transaction system. In addition, the chart required for the investigation may be generated based on the target enterprise's transaction data obtained from the bank transaction system or the transaction data in the target enterprise's financial records as the transaction image data.
[0042] In some implementations, the regulatory authorities investigate various business activities of the target enterprise and generate investigation text data, which may include investigation reports, contract documents, financial data, tax data, employee information, partner information, communication record data, etc. For example, the revenue data, profit data, debt data, etc. in the financial data can intuitively show the financial health of the target enterprise; the transaction terms, amount, performance, etc. in the contract documents can reflect the business dealings and credit status of the target enterprise; the tax payment status in the tax records can reflect the compliance of the target enterprise.
[0043] In some embodiments, the image-text association probability represents the degree of correlation between the transaction image data and the survey text data. When the image-text association probability is high, it means that the transaction image data and the survey text data have a high degree of matching and correlation in content. For example, if a transaction image data shows the outflow of a large amount of funds from an enterprise, and the contract document in the survey text data also mentions the payment of this fund, then the image-text association probability between them will be high. This high correlation indicates that the transaction behavior is related to the business activities of the target enterprise, which may imply potential risks, such as unknown fund flows, abnormal large expenditures, etc., so as to achieve early warning of the target enterprise's business risks. .
[0044] In some embodiments, a first image-text association probability threshold and a second image-text association probability threshold are provided. When the first image-text association probability is higher than the first image-text association probability threshold, the corresponding first transaction image data is determined to be the target image data associated with the survey text data; when the second image-text association probability is higher than the second image-text association probability threshold, the corresponding second transaction image data is determined to be the target image data associated with the survey text data.
[0045] In some embodiments, in order to more intuitively display the relationship between the target image data and the survey text data, a visual data chain can be generated. The visual data chain can be displayed in a variety of ways, such as a graphical interface, an interactive chart, etc. In the graphical interface, the transaction image data and the survey text data can be displayed in the form of nodes, and the lines between the nodes represent the first image-text association probability or the second image-text association probability between them, and the thickness of the line can represent the size of the image-text association probability. In addition, the image-text association probability values of different ranges can also be corresponded to the risk level, for example, when the image-text association probability is less than the first risk threshold, the image-text association probability corresponds to a low risk level; when the image-text association probability is greater than or equal to the first risk threshold and less than the second risk threshold, the image-text association probability corresponds to a medium risk level; when the image-text association probability is greater than or equal to the second risk threshold, the image-text association probability corresponds to a high risk level. Different colors or icons can be used to identify different risk levels, for example, red represents a high risk level, yellow represents a medium risk level, and green represents a low risk level. In this way, office personnel can quickly browse and understand the relationship between a large amount of data and promptly discover the potential operating risks of the target enterprise. Through the visual data chain, the actual situation of the target enterprise can be displayed in an intuitive and convenient way, which is conducive to early warning of business risks for the target enterprise.
[0046] In the above embodiment, firstly, target image data associated with the investigation text data is determined in the multiple frames of transaction image data according to the first image-text association probability data and the second image-text association probability data; secondly, a visual data chain is generated based on the target image data and the investigation text data. By screening and visualizing the transaction image data, the interference of redundant information is avoided, which helps to improve the accuracy and efficiency of the subsequent data chain analysis.
[0047] In this embodiment, a method for early warning of business risk of an enterprise is provided, the method comprising: S110, obtaining multiple frames of transaction image data and survey text data of the target enterprise.
[0048] The target enterprise may be an enterprise that the regulatory authorities need to investigate. The transaction image data may be image information generated in the transaction activities of the target enterprise. Any transaction image data corresponds to a target image type. The target image type is used to characterize the complexity of the information in the transaction image data, such as simple or complex. The investigation text data may be text information related to the business activities of the target enterprise obtained through investigation.
[0049] In some embodiments, the target enterprise's transaction image data may be obtained from a bank transaction system, such as a chart output by the bank transaction system. In addition, the chart required for the investigation may be generated based on the target enterprise's transaction data obtained from the bank transaction system or the transaction data in the target enterprise's financial records as the transaction image data.
[0050] In some implementations, the regulatory authorities investigate various business activities of the target enterprise and generate investigation text data, which may include investigation reports, contract documents, financial data, tax data, employee information, partner information, communication record data, etc.
[0051] S120: Call a target multimodal model corresponding to the target image type.
[0052] The target multimodal model may be a model that supports both image and text modes. The feature extraction granularity of the target multimodal model matches the image content complexity of any transaction image data. The feature extraction granularity may be the degree of refinement of the target multimodal model in extracting image and text features.
[0053] In some embodiments, if the target image type is a complex image type, the corresponding target multimodal model needs to extract more fine-grained features in order to more accurately capture the detailed information in the transaction image data; if the target image type is a simple image type, the feature granularity extracted by the corresponding target multimodal model can be coarser in order to reduce the complexity of processing and improve the performance of transaction image data processing.
[0054] S130 , inputting any transaction image data and survey text data into a target multimodal model for feature extraction and feature alignment, and obtaining image-text association probability data between any transaction image data and survey text data.
[0055] Among them, feature extraction is performed on any transaction image data to obtain transaction image features. Feature extraction is performed on survey text data to obtain survey text features. Feature alignment can be mapping transaction image features and survey text features into a shared feature space, so that the features of the two modalities can complement, compare and integrate each other, thereby improving the analysis effect of multimodal data. The image-text association probability data can be the degree of association between transaction image data and survey text data.
[0056] In some embodiments, the target multimodal model includes an image processing branch and a text processing branch in parallel. First, the image processing branch is used to extract features of any transaction image data to obtain transaction image features of any transaction image data; then, the text processing branch is used to extract features of the survey text data to obtain survey text features of the survey text data; then, feature alignment is performed based on the transaction image features and the survey text features; finally, the similarity between the transaction image features and the survey text features is compared, which may be to calculate the Euclidean distance or cosine similarity between them, and the similarity is used as the image-text association probability data.
[0057] S140. Generate a visual data chain based on the image-text association probability data, the survey text data, and the multi-frame transaction image data to provide an early warning of the target enterprise's business risks.
[0058] Among them, the visual data chain can be a data display chain generated by visually displaying graphic-text association probability data, survey text data and multi-frame transaction image data, which helps office personnel to intuitively understand and analyze the association relationship in the data.
[0059] In some embodiments, a visual data chain is generated based on the image-text association probability data, the survey text data, and the multi-frame transaction image data. The visual data chain can be displayed in a graphical manner. Figure 1b In addition, the visual data chain can also be displayed in the form of charts or other visual forms. Through the visual data chain, complex data relationships can become more intuitive, which helps to improve the efficiency of analyzing and identifying business risks of enterprises and timely provide business risk warnings for target enterprises.
[0060] In the above embodiment, firstly, the multi-frame transaction image data and the survey text data of the target enterprise are obtained; then the target multimodal model corresponding to the target image type is called; then any transaction image data and the survey text data are input into the target multimodal model for feature extraction and feature alignment, and the image-text association probability data between any transaction image data and the survey text data is obtained; finally, a visual data chain is generated according to the image-text association probability data, the survey text data and the multi-frame transaction image data to provide an early warning of the target enterprise's business risk. Based on the transaction image data for different target image types, the corresponding target multimodal large model is called for analysis, which can take into account the accuracy of complex transaction image data processing and the efficiency of simple transaction image processing. Furthermore, based on the target multimodal large model, feature extraction and feature alignment are performed on the multi-frame transaction image data and the survey text data, and data from different sources can be comprehensively analyzed to achieve comprehensive information mining, thereby obtaining more accurate image-text association probability data. Finally, based on the visual data chain, office staff can analyze, judge and warn the business risks of the target enterprise more quickly and accurately.
[0061] See also Figure 2 In some embodiments, generating a visualization data chain based on the image-text association probability data, the survey text data, and the multi-frame transaction image data includes: S210: Determine target image data associated with the investigation text data in the multiple frames of transaction image data according to the image-text association probability data.
[0062] S220: Generate a visual data chain based on the target image data and the survey text data.
[0063] In some embodiments, a threshold value of image-text association probability is set. When the image-text association probability data is higher than the threshold value of image-text association probability, the corresponding target image data is determined to be the target image data associated with the survey text data. Then, a visual data chain can be generated based on the survey text data and the screened target image data associated with the survey text data.
[0064] In the above embodiment, firstly, target image data associated with the investigation text data is determined in the multi-frame transaction image data according to the image-text association probability data; secondly, a visual data chain is generated based on the target image data and the investigation text data. By screening the transaction image data, the interference of redundant information is avoided, which helps to improve the accuracy and efficiency of the subsequent data chain analysis.
[0065] See also Figure 3 In some embodiments, the first image type is a complex image type. The multi-frame transaction image data includes a transaction network map and a transaction frequency heat map, the transaction network map and the transaction frequency heat map correspond to complex image types respectively; the complex image type corresponds to the first multimodal model. The first multimodal model is used to perform association prediction on the first transaction image data and the survey text data to obtain a first image-text association probability between the first transaction image data and the survey text data, including: S310: Input the transaction network graph and the survey text data into a first multimodal model for feature extraction and feature alignment to obtain first association probability data between the transaction network graph and the survey text data.
[0066] S320: Input the transaction frequency heat map and the survey text data into the first multimodal model for feature extraction and feature alignment to obtain second association probability data between the transaction frequency heat map and the survey text data.
[0067] The first image-text association probability includes first association probability data and second association probability data. The first association probability data may be data for quantifying the relationship between the transaction network graph and the survey text data. The second association probability data may be data for quantifying the relationship between the transaction frequency heat map and the survey text data. The first multimodal model may be a fine-tuned CLIP model.
[0068] The transaction network map can be a network structure generated based on the fund flow data, related transaction data and business partnerships between the target enterprise's account and other accounts, which is used to reveal the transaction patterns, fund flows and potential correlations between enterprises. From the details of the transaction network map, we can find abnormal transaction patterns between accounts, frequent fund transfers and other unusual transaction behaviors. These small changes may indicate potential risks, correlations or non-compliant operations. By deeply analyzing these details, it is helpful to better identify the operating risks of the target enterprise.
[0069] The transaction frequency heat map can be image data generated based on the target enterprise's transaction time data, transaction location data, transaction volume data, transaction object data, transaction type data (such as sales transactions, payment transactions, refunds, transfers), etc. From the details of the transaction frequency heat map, it is possible to discover abnormal transaction patterns and trends of the target enterprise, such as abnormal transaction activities that frequently occur in a specific time period or region, sudden increases in specific accounts or transaction types, and transaction fluctuations that are significantly different from historical data.
[0070] In some embodiments, the transaction network graph and survey text data are input into the first multimodal model for feature extraction and feature alignment. By comparing the similarity between the transaction image features and the survey text features, the association probability data of abnormal capital flow, the risk association probability data with other enterprises, the probability data of potential illegal associations between accounts, etc. can be obtained. These association probability data constitute the first association probability data.
[0071] In some embodiments, the transaction frequency heat map and the survey text data are input into the first multimodal model for feature extraction and feature alignment. By comparing the similarity between the transaction image features and the survey text features, the association probability data of abnormal transaction activities, the association probability data of sudden increases in specific accounts or transaction types, the association probability data of transaction fluctuations, etc. can be obtained. These association probability data constitute the second association probability data.
[0072] In the above embodiment, the transaction network map and the survey text data are first input into the first multimodal model for feature extraction and feature alignment, and the first association probability data between the transaction network map and the survey text data is obtained; then the transaction frequency heat map and the survey text data are input into the first multimodal model for feature extraction and feature alignment, and the second association probability data between the transaction frequency heat map and the survey text data is obtained. By inputting complex transaction image data and survey text data into the first multimodal model for feature extraction and alignment, the corresponding image-text association probability data can be obtained, thereby realizing the association analysis of complex image data and survey text data, and providing an accurate data basis for early warning of business risks for target enterprises.
[0073] See also Figure 4 In some embodiments, the processing process of the transaction network map and the transaction frequency heat map is the same; the transaction network map and the survey text data are input into the first multimodal model for feature extraction and feature alignment to obtain the first association probability data between the transaction network map and the survey text data, including: S410, performing multi-level feature extraction on the transaction network graph at multiple preset levels to obtain image granularity features at each preset level.
[0074] S420, performing multi-level feature extraction on the survey text data at multiple preset levels to obtain text granularity features at each preset level.
[0075] S430: performing hierarchical alignment and similarity calculation on the image granularity features and the text granularity features at multiple preset levels to obtain association probability data at each preset level.
[0076] S440: Perform weighted summation based on the association probability data at each preset level to obtain first association probability data.
[0077] Among them, multi-level feature extraction can be to extract different levels of features from the transaction image data and the survey text data respectively, so that the first multimodal model can deeply understand the transaction image data and survey text data of complex image types.
[0078] In some embodiments, the first multimodal model may be a bimodal model of images and texts, with two preset levels. In the first level, the first multimodal model uses a smaller scale to perform convolution operations on the transaction network graph to obtain first-level image granularity features; uses word-based granularity to extract features from the survey text data to obtain first-level text granularity features. In the second level, the first multimodal model uses a larger scale to perform convolution operations on the transaction network graph to obtain second-level image granularity features; uses sentence-based granularity to extract features from the survey text data to obtain second-level text granularity features.
[0079] In some embodiments, the first multimodal model maps the first-level image granularity features and the first-level text granularity features to the first feature space for feature alignment, and then performs similarity comparison to obtain first-level association probability data. And the first multimodal model maps the second-level image granularity features and the second-level text granularity features to the second feature space for feature alignment, and then performs similarity comparison to obtain second-level association probability data.
[0080] In some embodiments, the first level association probability data corresponds to a first preset weight, and the second level association probability data corresponds to a second preset weight. The product of the first level association probability data multiplied by the first preset weight plus the product of the second level association probability data multiplied by the second preset weight is the first association probability data.
[0081] It is understandable that the preset levels can also be three, four or more levels. More levels help to obtain more accurate first association probability data, but the model processing performance will be reduced. Therefore, the preset levels should be set reasonably to strike a balance between accuracy and processing performance.
[0082] In the above embodiment, by inputting the transaction network map and the survey text data into the first multimodal model for feature extraction and feature alignment, the image granularity features and the text granularity features can be extracted at multiple preset levels, and hierarchical alignment and similarity calculation can be performed to obtain the association probability data at each level. Finally, the first association probability data is generated by weighted summation of the association probability data at these levels. This embodiment can effectively integrate complex types of image and text information, enhance the correlation analysis of image and text data, and thus provide more accurate decision support for the target enterprise's business risk warning.
[0083] It should be noted that the transaction frequency heat map and the survey text data are input into the first multimodal model for feature extraction and feature alignment to obtain the second association probability data between the transaction frequency heat map and the survey text data, including: performing multi-level feature extraction on the transaction frequency heat map at multiple preset levels to obtain image granularity features at each preset level; performing multi-level feature extraction on the survey text data at the multiple preset levels to obtain text granularity features at each preset level; performing hierarchical alignment and similarity calculation on the image granularity features and the text granularity features at the multiple preset levels to obtain association probability data at each preset level; and performing weighted summation based on the association probability data at each preset level to obtain the second association probability data.
[0084] See also Figure 5 In some embodiments, the second image type is a simple image type; the multi-frame transaction image data includes a transaction timeline diagram and a capital flow diagram, the transaction timeline diagram and the capital flow diagram correspond to simple image types respectively; the simple image type corresponds to a second multimodal model; performing association prediction on the second transaction image data and the survey text data through the second multimodal model to obtain a second image-text association probability between the second transaction image data and the survey text data includes: S510: Input the transaction timeline graph and the survey text data into the second multimodal model for feature extraction and feature alignment to obtain third association probability data between the transaction timeline graph and the survey text data.
[0085] S520: Input the capital flow diagram and the survey text data into the second multimodal model for feature extraction and feature alignment to obtain fourth association probability data between the capital flow diagram and the survey text data.
[0086] The second image-text association probability includes third association probability data and fourth association probability data. The third association probability data may be data for quantifying the relationship between the transaction timeline diagram and the survey text data. The fourth association probability data may be data for quantifying the relationship between the capital flow diagram and the survey text data.
[0087] The transaction timeline chart can show the transaction activities of an enterprise in a specific period of time. The transactions in the chart are arranged in chronological order, and each transaction (such as payment, collection, deposit, withdrawal, etc.) is marked, which can reflect the trend, periodicity or suddenness of capital flow.
[0088] A fund flow diagram can show the flow path of funds between different accounts or companies. Each node represents an account or entity, and each edge represents the path and amount of fund flow. Through this diagram, the source and destination of funds can be reflected.
[0089] In some embodiments, the transaction timeline diagram and the survey text data are input into the second multimodal model for feature extraction and feature alignment. By comparing the similarity between the transaction image features and the survey text features, transaction pattern association probability data (such as large-scale capital flows, frequent small-scale transactions, etc.), time node matching association probability data (such as special events occurring in a certain time period for the target enterprise in the survey text data, and a large amount of capital flows occurring in the transaction timeline diagram during the time period), etc. can be obtained. These association probability data constitute the third association probability data.
[0090] In some embodiments, the funds flow diagram and the investigation text data are input into the second multimodal model for feature extraction and feature alignment. By comparing the similarity between the transaction image features and the investigation text features, the association probability data between the funds flow and the corporate behavior (for example, the target company has illegal cross-border transactions), the association probability data between the account behavior and the suspect (for example, the target company has a large amount of funds transactions with the suspect), etc. can be obtained. These association probability data constitute the fourth association probability data.
[0091] In the above embodiment, the transaction timeline diagram and the survey text data are first input into the second multimodal model for feature extraction and feature alignment, and the third association probability data between the transaction timeline diagram and the survey text data is obtained; then the capital flow diagram and the survey text data are input into the second multimodal model for feature extraction and feature alignment, and the fourth association probability data between the capital flow diagram and the survey text data is obtained. By inputting simple transaction image data and survey text data into the second multimodal model for feature extraction and alignment, the corresponding image-text association probability data can be obtained, thereby realizing the association analysis of simple image data and survey text data, and providing an accurate data basis for early warning of business risks for target enterprises.
[0092] See also Figure 6 In some embodiments, the processing process of the transaction timeline diagram and the capital flow diagram is the same; the second multimodal model includes a text encoder and an image encoder in parallel; the transaction timeline diagram and the survey text data are input into the second multimodal model for feature extraction and feature alignment, and the third association probability data between the transaction timeline diagram and the survey text data is obtained, including: S610 , encoding and converting the survey text data through a text encoder to obtain a text semantic vector of the survey text data.
[0093] S620 , encoding and converting the transaction timeline graph through an image encoder to obtain an image feature vector of the transaction timeline graph.
[0094] S630: Use a contrastive learning method to align and perform similarity calculations on the text semantic vector and the image feature vector to generate third association probability data.
[0095] Among them, the contrastive learning method can be a method of learning feature representation by comparing the similarity between the text semantic vector and the image feature vector, which can map the text semantic vector and the image feature vector to a unified feature space.
[0096] In some embodiments, the text encoder can be a model based on the Transformer architecture (such as BERT, GPT), which can encode and transform the survey text data to obtain a text semantic vector of the survey text data. The survey text data is converted into a vector form for further analysis, and feature alignment and similarity calculation can be performed subsequently. Specifically, the text encoder converts the vocabulary, sentence structure and semantic information in the survey text data into a high-dimensional vector, so that the second multimodal model can understand the information in the survey text data and provide input data for the comparative learning method.
[0097] In some embodiments, the image encoder can be a convolutional neural network (CNN), which can encode and transform the transaction timeline graph to obtain an image feature vector of the transaction timeline graph. Through the image encoder, the information in the transaction timeline graph is converted into a feature vector for subsequent feature alignment and similarity calculation. Specifically, the image encoder extracts key features in the transaction timeline graph, such as node relationships, capital flow paths, etc., and converts this information into a vector representation. This enables the second multimodal model to understand the information in the transaction timeline graph and provide input data for the contrastive learning method.
[0098] In some embodiments, the second multimodal model includes a contrastive learning module, which can map the text semantic vector and the image feature vector to a high-dimensional feature space through a contrastive learning method, and make similar feature vectors close to each other in the space, and dissimilar ones far away from each other.
[0099] In some embodiments, the transaction timeline diagram and the survey text data are first input into the second multimodal model for feature extraction to obtain a text semantic vector and an image feature vector; then, the two vectors are mapped to a high-dimensional feature space through a contrastive learning module, and their similarity is calculated (which can be Euclidean distance or cosine similarity), and finally the similarity is used as the third association probability data.
[0100] In the above embodiment, by inputting the transaction timeline diagram and the survey text data into the second multimodal model, using the parallel text encoder and image encoder to extract the text semantic vector and image feature vector respectively, and aligning the above vectors through the contrast learning method, the third association probability data between the transaction timeline diagram and the survey text data is generated. This embodiment can effectively integrate simple types of image and text information, enhance the correlation analysis of image and text data, and thus provide more accurate decision support for the target enterprise's business risk warning.
[0101] It should be noted that the funds flow diagram and the survey text data are input into the second multimodal model for feature extraction and feature alignment to obtain fourth association probability data between the funds flow diagram and the survey text data, including: encoding and transforming the survey text data through the text encoder to obtain a text semantic vector of the survey text data; encoding and transforming the funds flow diagram through the image encoder to obtain an image feature vector of the transaction timeline diagram; and aligning and similarity calculating the text semantic vector and the image feature vector using a contrastive learning method to generate the fourth association probability data.
[0102] See also Figure 7 In some embodiments, the target image type corresponding to any transaction image data is determined in the following manner: S710 , performing complexity evaluation on any transaction image data to obtain image complexity evaluation data of any transaction image data.
[0103] S720: Determine the target image type according to the target complexity range of the image complexity evaluation data and the complexity type relationship data.
[0104] The target image type includes a first image type and a second image type. The complexity type relationship data is used to describe the corresponding relationship between the complexity range and the image type of the transaction image data.
[0105] In some embodiments, the complexity of transaction image data can be evaluated by comprehensively evaluating texture complexity, object quantity, background noise detection, and image resolution. Exemplarily, the texture complexity of transaction image data can be evaluated by calculating texture features (such as gray-level co-occurrence matrix, LBP (local binary pattern)). Generally, the higher the texture, the more complex the image content. Exemplarily, the number and distribution of objects in transaction image data can be identified by image segmentation or target detection algorithms (such as YOLOV8, Faster R-CNN, etc.). More objects and more overlapping areas usually mean higher complexity. Exemplarily, the frequency domain features of the image (such as Fourier transform or wavelet transform) can be used to evaluate the noise component of the transaction image data. The more noise, the more complex the image. Exemplarily, the resolution of the transaction image data can be determined by the number of pixels in width and height. The higher the resolution, the more information it contains and the higher the complexity.
[0106] In some embodiments, image complexity evaluation data of any transaction image data is obtained by scoring. Specifically, texture complexity evaluation, object quantity evaluation, background noise detection evaluation, and image resolution evaluation all have preset thresholds. When the evaluation result of each item is higher than the preset threshold, the score is recorded as 1, otherwise it is recorded as 0, and the total score of the four evaluations is the image complexity evaluation data. Exemplarily, the complexity type relationship data is: when the image complexity evaluation data is greater than or equal to 3, the corresponding target image type is a complex image type, otherwise it is a simple image type. Please refer to Table 1 for examples of complexity evaluation of transaction image data.
[0107] Table 1 Example table of transaction image data complexity evaluation In the above embodiment, by performing complexity evaluation on the transaction image data and combining the complexity type relationship data, the target image type of the transaction image data can be accurately determined. This embodiment can effectively distinguish transaction images of different complexities, thereby providing an accurate image type basis for subsequent call target multimodal model analysis.
[0108] The embodiment of the present application further provides an enterprise operation risk early warning device, which includes: A data acquisition module, used to acquire multi-frame transaction image data and survey text data of a target enterprise; wherein the multi-frame transaction image data includes first transaction image data belonging to a first image type and second transaction image data belonging to a second image type; the first image type corresponds to a first multimodal model, the second image type corresponds to a second multimodal model, and the first multimodal model is connected in parallel with the second multimodal model; A first data processing module, configured to perform association prediction on the first transaction image data and the survey text data through the first multimodal model to obtain a first image-text association probability between the first transaction image data and the survey text data; wherein the first multimodal model uses a convolutional neural network with an adjustable void rate to globally extract microscopic features of a first granularity of the first transaction image data; a second data processing module, configured to perform association prediction on the second transaction image data and the survey text data through the second multimodal model to obtain a second image-text association probability between the second transaction image data and the survey text data; wherein the second multimodal model adopts a regional attention mechanism to also globally extract macro features of a second granularity of the second transaction image data, and the first granularity is smaller than the second granularity; The data link generation module is used to generate a visual data link based on the first image-text association probability, the second image-text association probability, the survey text data and the multi-frame transaction image data to provide an operational risk warning for the target enterprise.
[0109] See also Figure 8 The embodiment of the present application further provides an enterprise operation risk early warning device 800, and the enterprise operation risk early warning device 800 includes: The data acquisition module 810 is used to acquire multiple frames of transaction image data and survey text data of the target enterprise; wherein any transaction image data corresponds to a target image type; A model calling module 820 is used to call a target multimodal model corresponding to a target image type; wherein the feature extraction granularity of the target multimodal model matches the image content complexity of any transaction image data; The data processing module 830 is used to input any transaction image data and survey text data into the target multimodal model for feature extraction and feature alignment, and obtain image-text association probability data between any transaction image data and survey text data; The data link generation module 840 is used to generate a visual data link based on the image-text association probability data, the survey text data and the multi-frame transaction image data to provide an early warning of the target enterprise's business risks.
[0110] In some embodiments, the data link generation module 840 includes: An image data determination unit, used to determine target image data associated with the investigation text data in the multiple frames of transaction image data according to the image-text association probability data; The data link generation unit is used to generate a visual data link based on the target image data and the survey text data.
[0111] In some embodiments, the first image type is a complex image type, the multi-frame transaction image data includes a transaction network map and a transaction frequency heat map, the transaction network map and the transaction frequency heat map correspond to the complex image type respectively; the complex image type corresponds to the first multimodal model, and the data processing module 830 includes: The first probability data acquisition unit is used to input the transaction network map and the survey text data into the first multimodal model for feature extraction and feature alignment to obtain first association probability data between the transaction network map and the survey text data; and is used to input the transaction frequency heat map and the survey text data into the first multimodal model for feature extraction and feature alignment to obtain second association probability data between the transaction frequency heat map and the survey text data; wherein the first image-text association probability includes first association probability data and second association probability data.
[0112] In some embodiments, the processing process of the transaction network graph and the transaction frequency heat map is the same, and the first probability data acquisition unit includes: The feature extraction subunit is used to perform multi-level feature extraction on the transaction network graph at multiple preset levels to obtain image granularity features at each preset level; and to perform multi-level feature extraction on the survey text data at multiple preset levels to obtain text granularity features at each preset level; An alignment and calculation subunit, used for performing hierarchical alignment and similarity calculation on the image granularity features and the text granularity features at multiple preset levels to obtain association probability data at each preset level; The weighted summation subunit is used to perform weighted summation based on the association probability data at each preset level to obtain first association probability data.
[0113] In some embodiments, the second image type is a simple image type, the multi-frame transaction image data includes a transaction timeline diagram and a capital flow diagram, the transaction timeline diagram and the capital flow diagram correspond to simple image types respectively; the simple image type corresponds to the second multimodal model; the data processing module 830 further includes: The second probability data acquisition unit is used to input the transaction timeline diagram and the survey text data into the second multimodal model for feature extraction and feature alignment to obtain third association probability data between the transaction timeline diagram and the survey text data; and is used to input the capital flow diagram and the survey text data into the second multimodal model for feature extraction and feature alignment to obtain fourth association probability data between the capital flow diagram and the survey text data; wherein the second image-text association probability includes the third association probability data and the fourth association probability data.
[0114] In some embodiments, the processing process of the transaction timeline diagram and the capital flow diagram is the same; the second multimodal model includes a text encoder and an image encoder in parallel; the second probability data acquisition unit also includes: The encoding subunit is used to encode and transform the survey text data through a text encoder to obtain a text semantic vector of the survey text data; and to encode and transform the transaction timeline graph through an image encoder to obtain an image feature vector of the transaction timeline graph; The alignment and calculation subunit is used to align and perform similarity calculation on the text semantic vector and the image feature vector using a contrastive learning method to generate third association probability data.
[0115] In some implementations, the enterprise operation risk early warning device 800 further includes: A complexity evaluation module, used to perform complexity evaluation on any transaction image data to obtain image complexity evaluation data of any transaction image data; The image type determination module is used to determine the target image type according to the target complexity range of the image complexity assessment data and the complexity type relationship data; wherein the complexity type relationship data is used to describe the corresponding relationship between the complexity range and the image type of the transaction image data. The target image type includes a first image type and a second image type.
[0116] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0117] The enterprise operation risk warning device in this embodiment is presented in the form of a functional unit, where the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0118] See also Fig. 9 , Fig. 9 is a schematic diagram of the structure of a computer device provided in an embodiment of the present application, such as Fig. 9As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Fig. 9 A processor 10 is taken as an example.
[0119] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0120] The memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiment.
[0121] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0122] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0123] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Fig. 9 The example of connecting through bus is taken in the following.
[0124] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0125] The embodiment of the present application also provides a computer-readable storage medium. The above method according to the embodiment of the present application can be implemented in hardware, firmware, or can be implemented as a computer code that can be recorded in a storage medium, or can be implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and will be stored in a local storage medium through a network download, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state hard disk, etc.; further, the storage medium can also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor, or hardware, the method shown in the above embodiment is implemented.
[0126] The embodiment of the present application provides a computer program product, which includes computer instructions, which are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method of any embodiment of the present application.
[0127] Although the embodiments of the present application are described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations are all within the scope defined by the appended claims.
[0128] The systems, devices, modules or units described in the above embodiments may be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smart phone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0129] For the convenience of description, the above device is described in various units according to their functions. Of course, when implementing the present application, the functions of each unit can be implemented in the same or multiple software and / or hardware.
[0130] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0131] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in one process or multiple processes in the flowchart and / or one box or multiple boxes in the block diagram.
[0132] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory produce a manufactured product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0133] These computer program instructions may also be loaded onto a computer or other programmable data processing device so that a series of operational steps are executed on the computer or other programmable device to produce a computer-implemented process, whereby the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0134] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0135] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0136] The above is only an embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
[0137] Although the embodiments of the present application have been described in conjunction with the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present application, and such modifications and variations are all within the scope defined by the appended claims.
Claims
1. A method for early warning of business risk, characterized in that: The method comprises: Acquire multiple frames of transaction image data and survey text data of a target enterprise; wherein the multiple frames of transaction image data include first transaction image data belonging to a first image type and second transaction image data belonging to a second image type; the first image type corresponds to a first multimodal model, the second image type corresponds to a second multimodal model, and the first multimodal model is connected in parallel with the second multimodal model; The first multimodal model is used to perform association prediction on the first transaction image data and the survey text data to obtain a first image-text association probability between the first transaction image data and the survey text data; wherein the first multimodal model uses a convolutional neural network with an adjustable dilation rate to globally extract microscopic features of a first granularity of the first transaction image data; The second multimodal model is used to perform association prediction on the second transaction image data and the survey text data to obtain a second image-text association probability between the second transaction image data and the survey text data; wherein the second multimodal model uses a regional attention mechanism to globally extract macro features of a second granularity of the second transaction image data, and the first granularity is smaller than the second granularity; A visual data chain is generated based on the first image-text association probability, the second image-text association probability, the survey text data and the multiple frames of transaction image data to provide an operational risk warning for the target enterprise.
2. The method according to claim 1, characterized in that The generating of a visual data chain according to the first image-text association probability, the second image-text association probability, the survey text data and the multiple frames of transaction image data to provide an early warning of business risks for the target enterprise includes: Determine target image data associated with the survey text data in the multiple frames of transaction image data according to the first image-text association probability and the second image-text association probability; The visualization data chain is generated based on the target image data and the survey text data.
3. The method according to claim 1, characterized in that The first image type is a complex image type; the multi-frame transaction image data includes a transaction network map and a transaction frequency heat map, the transaction network map and the transaction frequency heat map correspond to complex image types respectively; the complex image type corresponds to the first multimodal model; The performing association prediction on the first transaction image data and the survey text data by using the first multimodal model to obtain a first image-text association probability between the first transaction image data and the survey text data includes: Inputting the transaction network graph and the survey text data into the first multimodal model for feature extraction and feature alignment to obtain first association probability data between the transaction network graph and the survey text data; The transaction frequency heat map and the survey text data are input into the first multimodal model for feature extraction and feature alignment to obtain second association probability data between the transaction frequency heat map and the survey text data; wherein the first image-text association probability includes the first association probability data and the second association probability data.
4. The method according to claim 3, characterized in that The processing process of the transaction network graph and the transaction frequency heat map is the same; The step of inputting the transaction network graph and the survey text data into the first multimodal model for feature extraction and feature alignment to obtain first association probability data between the transaction network graph and the survey text data includes: Performing multi-level feature extraction on the transaction network graph at multiple preset levels to obtain image granularity features at each preset level; Performing multi-level feature extraction on the survey text data at the multiple preset levels to obtain text granularity features at each preset level; Performing hierarchical alignment and similarity calculation on the image granularity features and the text granularity features at the plurality of preset levels to obtain association probability data at each preset level; The first association probability data is obtained by performing weighted summation based on the association probability data at each preset level.
5. The method according to claim 1, characterized in that The second image type is a simple image type; the multi-frame transaction image data includes a transaction timeline diagram and a capital flow diagram, the transaction timeline diagram and the capital flow diagram respectively correspond to simple image types; the simple image type corresponds to the second multimodal model; The performing association prediction on the second transaction image data and the survey text data by using the second multimodal model to obtain a second image-text association probability between the second transaction image data and the survey text data includes: Inputting the transaction timeline graph and the survey text data into the second multimodal model for feature extraction and feature alignment to obtain third association probability data between the transaction timeline graph and the survey text data; The capital flow diagram and the survey text data are input into the second multimodal model for feature extraction and feature alignment to obtain fourth association probability data between the capital flow diagram and the survey text data; wherein the second image-text association probability includes the third association probability data and the fourth association probability data.
6. The method according to claim 5, characterized in that The processing process of the transaction timeline diagram and the capital flow diagram is the same; the second multimodal model includes a text encoder and an image encoder in parallel; The step of inputting the transaction timeline diagram and the survey text data into the second multimodal model for feature extraction and feature alignment to obtain third association probability data between the transaction timeline diagram and the survey text data includes: The survey text data is encoded and converted by the text encoder to obtain a text semantic vector of the survey text data; The transaction timeline graph is encoded and converted by the image encoder to obtain an image feature vector of the transaction timeline graph; The text semantic vector and the image feature vector are aligned and similarly calculated using a contrastive learning method to generate the third association probability data.
7. The method according to any one of claims 1 to 6, characterized in that: Determine the target image type corresponding to any transaction image data in the following way: Comprehensively evaluating texture complexity, object quantity, background noise detection and image resolution, performing complexity evaluation on any transaction image data, and obtaining image complexity evaluation data of any transaction image data; The target image type is determined according to the target complexity range of the image complexity assessment data and the complexity type relationship data; wherein the complexity type relationship data is used to describe the correspondence between the complexity range and the image type of the transaction image data; the target image type includes a first image type and a second image type.
8. An enterprise operation risk early warning device, characterized in that: The device comprises: A data acquisition module, used to acquire multi-frame transaction image data and survey text data of a target enterprise; wherein the multi-frame transaction image data includes first transaction image data belonging to a first image type and second transaction image data belonging to a second image type; the first image type corresponds to a first multimodal model, the second image type corresponds to a second multimodal model, and the first multimodal model is connected in parallel with the second multimodal model; a first data processing module, configured to perform association prediction on the first transaction image data and the survey text data through the first multimodal model to obtain a first image-text association probability between the first transaction image data and the survey text data; wherein the first multimodal model uses a convolutional neural network with an adjustable dilation rate to globally extract microscopic features of a first granularity of the first transaction image data; a second data processing module, configured to perform association prediction on the second transaction image data and the survey text data through the second multimodal model to obtain a second image-text association probability between the second transaction image data and the survey text data; wherein the second multimodal model adopts a regional attention mechanism to also globally extract macro features of a second granularity of the second transaction image data, and the first granularity is smaller than the second granularity; The data link generation module is used to generate a visual data link based on the first image-text association probability, the second image-text association probability, the survey text data and the multi-frame transaction image data to provide an operational risk warning for the target enterprise.
9. A computer device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the method according to any one of claims 1 to 7 by executing the computer instructions.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Multi-modal image-text matching model and construction method, device and application thereof
CN115935199A
Community risk level assessment method and system based on multi-modal data fusion
CN117726162A
Intelligent visualization and text association method for multi-modal knowledge graph
CN119441281A
Service report generation method and device, equipment and medium
CN119559297A
Method of training image-text retrieval model, method of multimodal image retrieval, electronic device and medium
US20220391587A1