Financial risk identification method and device, equipment and storage medium

By using a large vertical model in the financial field to identify risk labels and extract entity relationships, combined with cross-validation of relationship graphs, the problems of incomplete risk label coverage and high misjudgment rate in existing technologies are solved, achieving higher risk identification accuracy and explainability.

CN120807165APending Publication Date: 2025-10-17CHINA SOUTHERN POWER GRID CAPITAL HLDG CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511161353.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-19
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing financial risk identification methods find it difficult to extract comprehensive risk information from massive unstructured documents, resulting in incomplete risk label coverage and high misjudgment rate, which affects identification accuracy.

Method used

A vertical domain big model is used to identify risk labels and extract entity relationships from descriptive texts, and cross-validated with relationship graphs to improve the accuracy and completeness of risk identification.

Benefits of technology

Through the dual-channel cross-validation mechanism, the accuracy and interpretability of risk identification are significantly improved, and it can effectively deal with complex and hidden financial risks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120807165A_ABST
    Figure CN120807165A_ABST
Patent Text Reader

Abstract

The invention relates to a financial risk identification method and device, equipment and a storage medium. The financial risk identification method comprises the steps that a description text of a target project to be subjected to risk tag identification is acquired, and the description text is used for representing related information of a financial project to be invested; performing risk tag identification on the description text, and outputting a first identification result; performing entity recognition and entity relation extraction on the description text, and outputting an entity processing result; risk tag identification is carried out based on the entity processing result, a second identification result is obtained, and the first identification result and the second identification result are determined based on the same risk tag system; and determining a final risk identification result of the target project according to the first identification result and the second identification result. According to the method provided by the invention, the accuracy of risk identification in the financial field is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of computer, and in particular to a financial risk identification method, device and equipment and storage medium. BACKGROUND

[0002] In the field of financial investment, the related risk identification methods mostly depend on artificial research and judgment or simple models based on rules, and it is difficult to extract comprehensive risk information from massive unstructured documents, resulting in incomplete risk label coverage and high misjudgment rate, which seriously affects the accuracy of financial risk identification. SUMMARY

[0003] In order to solve the above technical problems, the present application provides a financial risk identification method, device, equipment and storage medium.

[0004] In a first aspect, the present application provides a financial risk identification method, comprising:

[0005] obtaining a description text of a target project to be identified by a risk label, wherein the description text is used to represent the related information of the financial project to be invested;

[0006] performing risk label identification on the description text, and outputting a first identification result;

[0007] performing entity identification and entity relationship extraction on the description text, and outputting an entity processing result;

[0008] performing risk label identification based on the entity processing result, and obtaining a second identification result, wherein the first identification result and the second identification result are determined based on the same risk label system;

[0009] determining the final risk identification result of the target project according to the first identification result and the second identification result.

[0010] In a second aspect, the present application provides a financial risk identification device, comprising:

[0011] an acquisition unit, configured to acquire a description text of a target project to be identified by a risk label, wherein the description text is used to represent the related information of the financial project to be invested;

[0012] a first identification unit, configured to perform risk label identification on the description text, and output a first identification result;

[0013] an entity processing unit, configured to perform entity identification and entity relationship extraction on the description text, and output an entity processing result;

[0014] The second identification unit is configured to identify a risk label based on the entity processing result to obtain a second identification result, wherein the first identification result and the second identification result are determined based on the same risk label system.

[0015] The determination unit is configured to determine a final risk identification result of the target project according to the first identification result and the second identification result.

[0016] In a third aspect, an electronic device is provided, comprising:

[0017] a memory;

[0018] a processor; and

[0019] a computer program;

[0020] The computer program is stored in the memory and configured to be executed by the processor to implement the method of the first aspect.

[0021] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program. The computer program is executed by a processor to implement the steps of the method of the first aspect.

[0022] The financial risk identification method provided by the present disclosure comprises the following steps: obtaining a description text of a target project to be identified by a risk label, wherein the description text is used to represent the related information of the financial project to be invested; performing risk label identification on the description text to output a first identification result; performing entity recognition and entity relationship extraction on the description text to output an entity processing result; performing risk label identification based on the entity processing result to obtain a second identification result, wherein the first identification result and the second identification result are determined based on the same risk label system; and determining a final risk identification result of the target project according to the first identification result and the second identification result. The method provided by the present application improves the accuracy of risk identification in the financial field. BRIEF DESCRIPTION OF DRAWINGS

[0023] The accompanying drawings, which are incorporated into and form part of the specification, illustrate embodiments consistent with the present disclosure and, together with the specification, serve to explain the principles of the present disclosure.

[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure or the prior art, the accompanying drawings needed to be used in the embodiments or prior art description will be briefly introduced. Obviously, for those skilled in the art, other drawings can also be obtained without creative labor based on these drawings.

[0025] Figure 1 A flowchart of a financial risk identification method provided by the present disclosure is shown in the figure;

[0026] Figure 2 Another flowchart of a financial risk identification method provided by an embodiment of the present disclosure is shown in FIG. 6.

[0027] Figure 3 A structural diagram of a financial risk identification device provided by an embodiment of the present disclosure is shown in FIG. 7.

[0028] Figure 4 A structural diagram of an electronic device provided by an embodiment of the present disclosure is shown in FIG. 8. DETAILED DESCRIPTION

[0029] In order to more clearly understand the above-mentioned purposes, features and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other without conflict.

[0030] In the following description, many specific details are set forth in order to provide a thorough understanding of the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments described in the specification are only some embodiments of the present disclosure, not all embodiments.

[0031] To solve the above technical problems, the present disclosure provides a financial risk identification method, which uses a vertical domain large model to greatly improve the accuracy of extracting entities and the relationship between entities from documents, and improve the completeness and accuracy of constructing a relationship graph. The risk features are extracted from the documents using the vertical domain large model, and the risk features are identified and classified. At the same time, the risk identification results inferred from the relationship graph are used to cross-verify the risk identification results output by the large model, and the two different verification methods are cross-used to further improve the identification ability of the risk label. One or more embodiments are described in detail as follows.

[0032] The financial risk identification method provided by the present disclosure can be applied to the risk identification scene in the field of financial investment. The method can be executed by a financial risk identification device, which can be realized by software and / or hardware, and the device can be integrated in an electronic device. The electronic device can include but is not limited to mobile terminals such as smartphones, notebook computers, digital broadcast receivers, personal digital assistants (PDA), tablet computers (Tablet PC), PMP (portable multimedia player), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, and the like, and fixed terminals such as digital televisions, desktop computers, smart home devices, and the like.

[0033] Figure 1A flowchart of a financial risk identification method provided by an embodiment of the present disclosure is shown in the figure, and specifically includes the following steps as shown in the figure: Figure 1

[0034] S101, acquire a description text of a target project to be subjected to risk label identification.

[0035] The description text is used to represent the related information of the financial investment project.

[0036] Understandably, the target project refers to a financial investment project. Specifically, the target user inputs a query request about the target project, and the query request is used to indicate the financial investment risk label identification of the target project. The query request includes a query prompt word or a query question of the target project, such as "the investment risk of project A", and also such as "which one is lower in risk, project A or project B?". Then, in response to the query request, the description text of the target project is acquired, wherein the description text refers to the related information of the target project, which is used to analyze the investment risk. For example, the description text includes the related information of the investment manager, the related information of the investment market, the related information of the enterprise finance, and the like. Other possible information for describing the investment project is not limited. In one possible case, when the query request involves multiple projects, the risk label identification is performed for each project, and based on the actual query request, the risk label identification results of each project are analyzed comprehensively to generate a reply to the query request.

[0037] Optionally, the description text of the target project to be subjected to risk label identification is acquired, and the acquisition can be realized through the following steps:

[0038] Acquire multi-modal data of the target project to be subjected to risk label identification, wherein the multi-modal data refers to financial industry data related to the target project; analyze the multi-modal data through a multi-modal large model to output the description text of the target project, wherein the multi-modal large model is applied to the financial field.

[0039] The multi-modal data includes at least two modal data of text, table, image, audio, and video.

[0040] ​Understandably, the multi-modal data of the target item is acquired, and the multi-modal data represents the data of the target item in the financial industry, wherein the multi-modal data includes at least two modal data of text, table, image, audio and video. Subsequently, the multi-modal data is parsed by a pre-trained multi-modal large model to output a description text of the target item, and the parsing refers to reading, analyzing, understanding and describing each modal data, and converting other modal data except the text into text content to obtain the description text, for example, reading table content and converting it into text content, describing image content and converting it into text content. The multi-modal large model is applied to the financial field, which can be understood as a large model in the financial vertical field. In addition to direct acquisition, the basic multi-modal large model can also be pre-trained and fine-tuned by a large amount of financial industry data, wherein the training data involved in the pre-training is not limited to financial field data, and the training data involved in the fine-tuning is concentrated in the financial field data.

[0041] Optionally, the multi-modal data of the target item to be subjected to risk label identification is acquired, which can be realized by the following steps:

[0042] The multi-source data of the target item is acquired, and the multi-source data refers to a data set collected from multiple information channels and having different modalities and structural forms, which is used to comprehensively represent the comprehensive state of the target item; data of different set modal types is extracted from the multi-source data to construct modal data sets corresponding to each set modal type; an association relationship between data of different set modal types is established based on actual source data to which the data belongs, so as to obtain a description text by associating and analyzing the multi-modal data; wherein the multi-modal data includes the modal data set and the association relationship.

[0043] Among them, the multi-source data includes at least two source data of financial data, economic data, operation data and multimedia data.

[0044] It can be understood that multi-source data refers to a data set collected from multiple data sources or multiple information channels, which has different modalities and structural forms, and can represent the comprehensive state of the project to a certain extent. The multi-source data includes at least two source data of financial data, economic data, operation data and multimedia data, for example, the financial data can be financial statements, the operation data can be news, and the multimedia data can be short videos. After obtaining the multi-source data of the target project, different modal types of data are extracted from the multi-source data, and modal data sets corresponding to different modal types are constructed, wherein each modal type can have a corresponding data set, and the multi-modal data includes at least two modalities of pictures, texts, tables, videos, etc. Subsequently, according to the actual source data to which the different modal data (i.e. data of different modal types) belongs, an association relationship between different modal data is established, so as to obtain a more comprehensive and more accurate description text based on the multi-modal data. In one embodiment, the news document includes multiple pictures and multiple texts. The multiple texts in the news are extracted and stored in the corresponding text data set, or at least part of the texts of some paragraphs are selected from the multiple texts and stored in the corresponding text data set. The multiple pictures in the news are extracted and stored in the corresponding picture data set, or some pictures are stored in the picture data set. At the same time, the multiple pictures and the multiple texts belong to the news document (i.e. the source data to which each modal data belongs, denoted as the attribution relationship), the positional relationship of the multiple pictures and the multiple texts in the entire news document, and the semantic relationship between the multiple pictures and the multiple texts. For example, the news document 1 includes picture a, picture b, text c and text d, wherein picture a, picture b, text c and text d belong to the news document 1 (the attribution relationship), text c is located between picture a and picture b (the positional relationship), and text c is a description statement about picture a (the semantic relationship). The association relationship between different modal data is not limited.

[0045] S102, risk label identification is performed on the description text, and a first identification result is output.

[0046] It can be understood that, on the basis of the above S101, the risk features in the description text are extracted, and risk label identification is performed according to the risk features to obtain a first identification result. The risk label identification is also a classification and grading of risks. The first identification result includes at least one risk type and a risk level of each risk type, for example, the risk type is credit risk, and the risk level is medium.

[0047] Optionally, risk label identification is performed on the description text, and a first identification result is output. Specifically, the following steps can be implemented:

[0048] The first large language model is used to perform risk label identification on the description text, and a first identification result is output.

[0049] It can be understood that a pre-trained first large language model is obtained, and the first large language model can be a large model in the financial vertical field. Then, the risk features in the document (the document is used to store the description text) are extracted by the first large language model, the risk is classified and graded, and a first identification result is obtained.

[0050] S103, entity recognition and entity relationship extraction are performed on the description text, and an entity processing result is output.

[0051] It can be understood that, on the basis of the above S101, entity recognition and entity relationship extraction are performed on the document storing the description text or the description text, and an entity processing result or an entity extraction result is obtained, wherein the entity processing result includes multiple entities and the relationship between the multiple entities, for example, the description text includes entity a and entity b, entity a is a residential area, entity b is a unit building in the residential area, and entity a and entity b have a subordinate relationship, which can be described as entity b being a unit building in entity a.

[0052] Optionally, entity recognition and entity relationship extraction are performed on the description text, and an entity processing result is output, which can be implemented by the following steps:

[0053] Entity recognition and entity relationship extraction are performed on the description text by the first large language model or the second large language model, and an entity processing result is output.

[0054] It can be understood that the entity and entity relationship extraction task can be implemented based on the first large language model, and can also be implemented based on the second large language model, and the first large language model and the second large language model are different. If it is implemented based on the first large language model, the extraction task and the risk classification and grading task can be called respectively. If it is implemented based on the second large language model, two large language models can be called respectively. It can be understood that whether based on one large language model or multiple large language models, the extraction task and the risk classification and grading task can be executed in parallel. In one embodiment, the entity extraction task and the first risk classification and grading task are executed in parallel based on one large language model, and then, after completing the entity extraction task, the second risk classification and grading task is executed based on the one large language model, wherein the first risk classification and grading task refers to risk classification and grading of the description text, and the second risk classification and grading task refers to risk classification and grading of the relationship graph, and the relationship graph is constructed based on the entity extraction result. In another embodiment, the entity extraction task and the first risk classification and grading task are executed in parallel based on two large language models, and then, after completing the entity extraction task, one of the large language models continues to execute the second risk classification and grading task.

[0055] S104, risk label recognition is performed based on the entity processing result, and a second identification result is obtained.

[0056] The first recognition result and the second recognition result are determined based on the same risk label system.

[0057] Understandably, on the basis of the above S103, after the extraction of the entity and the relationship between the entities in the description text is completed, the entity processing result is used as the key semantic feature input, and the preset risk mapping rule engine or machine learning classification model is used to automatically identify and generate the financial investment risk label related to the target project. The second recognition result refers to the risk label set generated by the semantic analysis of the entity processing result, which is in the form of structured data and includes one or more risk categories and their risk levels. The risk level can be divided into "low risk", "medium risk", "high risk" or "prohibit investment" levels according to the total risk score, or the confidence (0-1) of each label. For example, a risk category is "industry policy change risk", and the risk level is "high". In addition to the level represented by text, the level can also be represented by a numerical value. For example, within the risk level range of 0-1, 0.9 represents high risk, 0.5 represents medium risk, and 0.2 represents low risk. Other possible risk level representation methods are not limited.

[0058] The first recognition result and the second recognition result are determined based on the same risk label system, which includes multiple financial investment risk categories and risk levels included in each category. The output recognition result can include each risk category and the risk level of the category, or a specific category and the risk level of the specific category, or a risk category with a high risk level. Other possible output results are determined according to user needs.

[0059] Optionally, the risk label recognition based on the entity processing result obtains the second recognition result, which can be achieved through the following steps:

[0060] According to at least one entity and the relationship between at least one entity in the entity processing result, a relationship graph is constructed. The relationship graph is subjected to risk label recognition to obtain the second recognition result.

[0061] Understandably, a multi-dimensional relationship network graph (denoted as a relationship graph) is constructed for representing the internal structure and external association of the target project based on the structured entities and their semantic relationships extracted from the description text of the target project. The entity processing result includes entities and relationships between entities identified from multi-source data such as financial data, text files, false information, audio / video transcription content, etc. through natural language processing, information extraction or structured parsing technology, and the entity type includes at least one type of subject entity, attribute entity and event entity, etc. For example, the subject entity is the enterprise name, legal representative, shareholder, etc., the attribute entity is the financial indicator (asset-liability ratio, net profit), financing amount, financing round, industry classification, etc., and the event entity is "listed as an enforcement person", "major asset sale", etc. The relationship between entities includes at least one of the following: stock ownership relationship (such as "A company holds 60% of B company's stock ownership"), personnel concurrent relationship (such as "Zhang is the chairman of C company and the supervisor of D company"), upstream and downstream relationship in the supply chain (such as "G company is the main raw material supplier"), etc. Among them, A company and B company are entities, and other possible relationships between entities are not described. Subsequently, based on graph algorithm (such as betweenness centrality, community discovery) to measure the importance or vulnerability of each node in the relationship graph, the risk label is inferred, and the risk label includes risk category and risk level. The identification process specifically refers to a technical process for identifying potential systemic, conductive or hidden financial risks by using graph structure features and topology analysis technology. The second identification result highlights the technical advantages of graph analysis in discovering implicit risks, transmission paths and complex associations. In addition, the relationship graph includes multiple nodes and edges representing node relationships, where the node refers to the entity, and the edge of the node relationship refers to the semantic relationship between entities. The weight or attribute of the edge can reflect the relationship strength, risk level or frequency of occurrence.

[0062] S105, according to the first identification result and the second identification result, determining the final risk identification result of the target project.

[0063] Understandably, on the basis of the above steps, the preliminary risk identification result based on the original multi-source data (the first identification result) and the deep reasoning identification result based on the structured entity relationship graph (the second identification result) are fused and analyzed to generate a final risk identification result with higher accuracy, interpretability and robustness. The second identification result and the first identification result are complementary. The first identification result is a risk label set generated by directly performing end-to-end analysis on the description text of the target project, reflecting the surface semantics and explicit risk signals. The second identification result is a financial risk relationship graph constructed based on the entities and their relationships extracted from the description text, and the risk label set generated by graph traversal, rule matching or graph neural network reasoning, reflecting the deep structure, association transmission and implicit risk pattern.

[0064] The identification result of the risk label identification includes a risk type and a risk level of each risk type.

[0065] Optionally, according to the first identification result and the second identification result, a final risk label identification result of the target project is determined, which can be achieved through the following steps:

[0066] The risk levels of the same risk type in the first identification result and the second identification result are compared. If the risk levels of the same risk type in the first identification result and the second identification result are the same, the same risk type and the same risk level are determined as the final risk identification result of the target project. If the risk levels of the same risk type in the first identification result and the second identification result are different, the same risk type and the different risk levels are determined as the final risk identification result of the target project, and a pending mark is added to the risk identification result. The pending mark is used to indicate that a relevant person checks the risk identification result.

[0067] It can be understood that the risk categories / risk types and the risk levels of the same risk category in the first identification result and the second identification result are compared to determine the final financial risk label identification result. The risk identification result includes multiple risk labels, and the risk label includes a risk category and a risk level. In one possible case, the first identification result and the second identification result include each risk type and the risk level of each risk type. In this case, the risk levels of the same risk type are compared. If the risk levels of the same risk type are the same, the same risk level is determined as the final risk level. If the risk levels of the same risk type are different, the different risk levels are both determined as the final risk level (i.e., one risk type includes two risk levels), and a pending mark is added to the risk label. The pending mark is used to indicate that a relevant person checks the risk level to determine the final risk level, or the different risk levels are directly provided to the relevant user as a reference. In another possible case, the first identification result and the second identification result only output risk types of specific risk levels (such as medium and high). In this case, the same risk type and different risk types are determined. The determination of the risk level of the same risk type is described in the above embodiments. The different risk types and risk levels can be directly used as the final risk identification result, i.e., the risk types are union calculated. In addition, a weight can be set for each risk identification approach. For example, the second identification result determined based on the relationship graph is set to have a higher weight. In the case where the risk levels of the same risk type are different, the risk level with a higher weight is determined as the final risk level. The risk identification approach includes an approach of directly identifying based on a description text through a large language model and an approach 2 of identifying based on a relationship graph through a graph algorithm.

[0068] The financial risk identification method provided by the embodiments of the present disclosure aims to improve the accuracy, completeness and interpretability of risk judgment. By fusing two complementary identification paths, a double-channel cross-validation mechanism is constructed. Specifically, on the one hand, a financial vertical domain multi-modal large model is used to perform deep semantic analysis on the description text parsed from the multi-source data of the target project, directly identify the risk features therein, and generate a first identification result. On the other hand, based on the same large model, the description text is subjected to fine-grained entity and relationship extraction, key financial entities and their semantic relationships are identified, and a structured financial risk relationship graph is constructed therefrom; based on the topological structure and logical path of the relationship graph, through a graph reasoning or rule matching mechanism, composite and conductive risks caused by entity association are further identified, and a second identification result is generated. Subsequently, the first identification result (representing direct risk features) and the second identification result (representing graph reasoning risk features) are cross-validated, which significantly improves the completeness, accuracy and interpretability of risk identification, and effectively deals with complex, hidden and cross-subject risk identification in financial investment.

[0069] On the basis of the above embodiments, Figure 2 The flowchart of another financial risk identification method provided by the embodiments of the present disclosure is shown in the following steps: Figure 2

[0070] 1) obtaining multi-source data of a target project; 2) extracting multi-modal data from the multi-source data; 3) parsing the multi-modal data by a vertical domain multi-modal large model to obtain description text; 4) extracting and analyzing risk features in the description text by a vertical domain large language model to obtain a first risk identification result; 5) extracting entities and relationships between entities in the description text by the vertical domain large language model to obtain an entity extraction result; 6) constructing a relationship graph based on the entity extraction result; 7) inferring the risk represented by the relationship graph by a graph algorithm to obtain a second risk identification result; and 8) cross-validating the first risk identification result and the second risk identification result to obtain a final risk identification result.

[0071] It can be understood that the specific implementation of the above 1) to 8) is described with reference to the above embodiments, which will not be repeated here.

[0072] On the basis of the above examples, Figure 3 The structure diagram of a financial risk identification device provided by the embodiments of the present disclosure is shown in the following steps: Figure 3

[0073] ​​The acquisition unit 301 is configured to acquire a description text of a target project to be subjected to risk label identification, wherein the description text is used to represent related information of the financial project to be invested.

[0074] The first identification unit 302 is configured to perform risk label identification on the description text and output a first identification result.

[0075] The entity processing unit 303 is configured to perform entity identification and entity relationship extraction on the description text and output an entity processing result.

[0076] The second identification unit 304 is configured to perform risk label identification based on the entity processing result to obtain a second identification result, wherein the first identification result and the second identification result are determined based on the same risk label system.

[0077] The determination unit 305 is configured to determine a final risk identification result of the target project according to the first identification result and the second identification result.

[0078] Optionally, the first identification unit 302 is configured to:

[0079] perform risk label identification on the description text by using the first large language model and output the first identification result.

[0080] Optionally, the entity processing unit 303 is configured to:

[0081] perform entity identification and entity relationship extraction on the description text by using the first large language model or the second large language model and output the entity processing result.

[0082] Optionally, the acquisition unit 301 is configured to:

[0083] acquire multi-modal data of the target project to be subjected to risk label identification, wherein the multi-modal data refers to financial industry data related to the target project.

[0084] perform analysis on the multi-modal data by using a multi-modal large model to output the description text of the target project, wherein the multi-modal large model is applied to the financial field.

[0085] Optionally, the second identification unit 304 is configured to:

[0086] construct a relationship graph according to at least one entity and at least one relationship between entities in the entity processing result.

[0087] perform risk label identification on the relationship graph to obtain the second identification result.

[0088] The identification result of the risk label identification includes a risk type and a risk level of each risk type.

[0089] Optionally, the determining unit 305 is configured to:

[0090] compare the risk levels of the same risk types in the first recognition result and the second recognition result;

[0091] if the risk levels of the same risk types in the first recognition result and the second recognition result are the same, the same risk types and the same risk levels are determined as the final risk recognition result of the target project;

[0092] if the risk levels of the same risk types in the first recognition result and the second recognition result are different, the same risk types and the different risk levels are determined as the final risk recognition result of the target project, and the risk recognition result is marked with a pending mark, the pending mark being used to indicate that a related person checks the risk recognition result.

[0093] Optionally, the obtaining unit 301 is configured to:

[0094] obtain multi-source data of the target project, the multi-source data being a data set with different modalities and structural forms collected from multiple information channels and used to comprehensively represent a comprehensive state of the target project;

[0095] extract data of different set modal types from the multi-source data, and construct a modal data set corresponding to each set modal type;

[0096] establish an association relationship between the data of different set modal types based on actual source data to which the data of different set modal types belong, so as to perform association analysis on the multi-modal data to obtain description text;

[0097] The multi-modal data includes the modal data set and the association relationship.

[0098] The multi-source data includes at least two source data in financial data, economic data, operation data and multimedia data; and the multi-modal data includes at least two modal data in text, table, image, audio and video.

[0099] Figure 3 The financial risk recognition device of the illustrated embodiment can be used to execute the technical solutions of the above method embodiments, and the implementation principles and technical effects are similar, which will not be described here.

[0100] Figure 4 The structure schematic diagram of the electronic device provided by the embodiment of the present disclosure is provided. The following will be specifically referred to Figure 4, which shows a schematic structural diagram of an electronic device 400 suitable for implementing the embodiments of the present disclosure. The electronic device 400 in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable electronic devices, and the like, as well as fixed terminals such as digital TVs, desktop computers, smart home devices, and the like. Figure 4 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0101] like Figure 4 As shown, electronic device 400 may include a processing device 401 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes to implement the financial risk identification method of the embodiment described in the present disclosure according to a program stored in a read-only memory (ROM) 402 or a program loaded from a storage device 408 into a random access memory (RAM) 403. Various programs and data required for the operation of electronic device 400 are also stored in RAM 403. Processing device 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to bus 404.

[0102] Typically, the following devices may be connected to the I / O interface 405: an input device 406 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 407 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 408 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 409. The communication device 409 may allow the electronic device 400 to communicate with other devices wirelessly or by wire to exchange data. Although Figure 4 The electronic device 400 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.

[0103] In particular, the processes described above with reference to the flowcharts can be implemented as a computer software program according to embodiments of the present disclosure. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts, thereby implementing the financial risk identification method as described above. In such embodiments, the computer program can be downloaded and installed from a network by the communication device 409, or installed from the storage device 408, or installed from the ROM 402. When the computer program is executed by the processing device 401, the above-mentioned functions defined in the methods of embodiments of the present disclosure are performed.

[0104] It should be noted that the computer-readable medium described above in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium may, for example, be, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any suitable combination thereof. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in a baseband or as a carrier wave in a propagated data signal, in which the computer-readable program code is carried. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination thereof. The computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), or the like, or any suitable combination thereof.

[0105] In some embodiments, the client, server, or other computing machines can communicate using any known or later developed form of computer-readable media, including but not limited to wireless media, wire-based media, optical-based media, and the like. In some embodiments, the client, server, or other computing machines can communicate using any current or later developed network protocol, such as the HyperText Transfer Protocol (HTTP), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network ("LAN"), a wide area network ("WAN"), the Internet, and the like, as well as any current or later developed network.

[0106] The computer-readable medium described above can be included within the electronic device described above; or can exist separately from the electronic device, and not be assembled into the electronic device.

[0107] Optionally, when the one or more programs described above are executed by the electronic device, the electronic device can further perform other steps described in the embodiments above.

[0108] Computer program code for carrying out operations of the present disclosure can be written in any one or combination of one or more programming languages, including an object-oriented programming language such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network ("LAN") or a wide area network ("WAN"), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0109] The computer program product of the first aspect can include one or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform the operations of the first aspect. The computer program product of the first aspect can include a non-transitory computer-readable medium storing code that, when executed, causes a computer to perform operations for the first aspect.

[0110] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware, or by a combination of software and hardware. In some cases, the names of the units do not constitute a limitation on the units themselves.

[0111] The functions described in this specification can be implemented in hardware, software, or any combination thereof. In some embodiments, the functionality described can be provided as a service or software that can be used on a computer system, e.g., that is implemented in software, or that is provided by a service that is executed on a computer system. In some embodiments, the functions described can be implemented in software such as one or more computer programs.

[0112] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0113] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or gateway that includes a series of elements includes not only those elements, but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article or gateway. In the absence of further restrictions, the elements defined by the sentence "including a..." do not exclude the presence of other identical elements in the process, method, article or gateway that includes the elements.

[0114] The foregoing description is intended only to provide specific embodiments of the present disclosure, intended to enable those skilled in the art to understand and implement the present disclosure. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure is not intended to be limited to the embodiments described herein, but rather to be construed in the broadest manner consistent with the principles and novel features disclosed herein.

Claims

1. A financial risk identification method, characterized in that: include: Obtaining a description text of a target project to be identified for risk labels, wherein the description text is used to characterize relevant information of the financial project to be invested; Perform risk label recognition on the description text and output a first recognition result; Perform entity recognition and entity relationship extraction on the description text, and output entity processing results; performing risk label identification based on the entity processing result to obtain a second identification result, wherein the first identification result and the second identification result are determined based on the same risk label system; A final risk identification result of the target project is determined based on the first identification result and the second identification result.

2. The method according to claim 1, characterized in that The step of performing risk label identification on the description text and outputting a first identification result includes: Performing risk label recognition on the description text using a first language model, and outputting a first recognition result; The performing entity recognition and inter-entity relationship extraction on the description text and outputting entity processing results includes: Entity recognition and inter-entity relationship extraction are performed on the description text through the first largest language model or the second largest language model, and an entity processing result is output.

3. The method according to claim 1, characterized in that The step of obtaining the description text of the target project to be identified with a risk tag includes: Acquire multimodal data of a target project for which risk label identification is to be performed, wherein the multimodal data refers to financial industry data related to the target project; The multimodal data is parsed using a multimodal macro model to output a description text of the target project, wherein the multimodal macro model is applied in the financial field.

4. The method according to claim 1, wherein The performing risk label identification based on the entity processing result to obtain a second identification result includes: constructing a relationship graph based on at least one entity in the entity processing result and the relationship between the at least one entity; Risk label identification is performed on the relationship graph to obtain a second identification result.

5. The method according to claim 1, wherein The identification result of the risk label identification includes the risk type and the risk level of each risk type. The determining of the final risk label identification result of the target project based on the first identification result and the second identification result includes: comparing risk levels of the same risk type in the first identification result and the second identification result; If the risk level of the same risk type in the first identification result and the same risk level in the second identification result is the same, determining the same risk type and the same risk level as the final risk identification result of the target project; If the risk levels of the same risk type are different, the same risk type and different risk levels will be determined as the final risk identification result of the target project, and the risk identification result will be marked with a pending mark, which is used to instruct relevant personnel to verify the risk identification result.

6. The method according to claim 3, characterized in that The step of obtaining multimodal data of a target project to be subjected to risk label identification includes: Acquire multi-source data of the target project, wherein the multi-source data refers to a data set collected from multiple information channels and having different modalities and structures, and is used to comprehensively characterize the comprehensive status of the target project; Extracting data of different set modal types from the multi-source data, and constructing a modal data set corresponding to each set modal type; Establishing an association relationship between the data of different set modal types based on actual source data to which the data of different set modal types belong, so as to perform association analysis on the multimodal data to obtain the description text; The multimodal data includes the modality data set and the association relationship.

7. The method according to claim 6, characterized in that The multi-source data includes at least two source data from among financial data, economic data, operational data and multimedia data; the multi-modal data includes at least two modal data from among text, table, image, audio and video.

8. A financial risk identification device, characterized in that: include: An acquisition unit, configured to acquire a description text of a target project to be identified with a risk label, wherein the description text is used to represent relevant information of the financial project to be invested; a first recognition unit, configured to perform risk label recognition on the description text and output a first recognition result; An entity processing unit, configured to perform entity recognition and relationship extraction on the description text, and output entity processing results; a second identification unit, configured to perform risk label identification based on the entity processing result to obtain a second identification result, wherein the first identification result and the second identification result are determined based on the same risk label system; A determination unit is used to determine a final risk identification result of the target project based on the first identification result and the second identification result.

9. An electronic device, characterized in that: include: Memory; processor; as well as computer programs; The computer program is stored in the memory and is configured to be executed by the processor to implement the financial risk identification method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the financial risk identification method according to any one of claims 1 to 7 are implemented.