Prediction Method, Device, Equipment, Medium and Program Product for Credit Risk
By constructing a knowledge graph of credit subjects and using risk prediction models, combining the risk prediction results of target credit subjects and associated credit subjects, the problem of failure to effectively utilize the relationship between credit subjects in the existing technology is solved, and higher accuracy of credit risk prediction and timely risk transmission discovery is achieved.
Patent Information
- Application Number
- CN202110304717.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-22
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-03-22
AI Technical Summary
The prior art fails to effectively utilize the correlation between credit subjects in credit risk prediction, resulting in low accuracy of prediction results and the inability to detect the transmission of credit risks in a timely manner.
By obtaining the target characteristic data of the credit subject, a knowledge graph of the credit subject is constructed, and the associated credit subjects with an association relationship with the target credit subject is identified, and based on the risk prediction model, the credit risk prediction results of the target credit subject are determined.
It improves the accuracy of credit risk prediction, can promptly discover the transmission of credit risk between credit subjects, and enhances the ability to analyze the relationship between credit subjects.
Smart Images

Figure CN112927082B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and in particular, to a method, apparatus, device, medium, and program product for predicting credit risk. Background Art
[0002] With the prosperous development of the economy, a large number of credit entities as parties to credit relationships have emerged in the fields of market transactions and investments. Different credit entities not only play their respective roles in the credit system and undertake their respective socio-economic functions, but also form intricate relational connections among these originally independent credit entities and other credit entities through transactions and investment behaviors. Affected by these complex relational connections, when one or some credit entities exhibit credit-loss behaviors, the probability of credit risk occurring in the credit entities that have close relational connections with them will also increase accordingly. Effective credit risk prediction helps to make risk judgments and early warnings for credit entities in the fields of market transactions and investments. Therefore, it is very important to conduct in-depth analysis and mining of the relational connections existing among these credit entities to obtain the explicit and implicit relational connections existing among credit entities.
[0003] Related technologies also provide some solutions for predicting the credit risk of credit entities, but most of them obtain the prediction results of the credit risk of a credit entity by statistically analyzing and comparing the relevant data of the credit entity itself. Since the credit entity is statistically analyzed and compared as an independent individual, the relational connections formed between the credit entity and other credit entities are not effectively utilized, and the dimension of data analysis is low, resulting in low accuracy of the prediction results of credit risk and the inability to timely detect the transmission of credit risk among credit entities. Summary of the Invention
[0004] To achieve the above objectives, one aspect of the present disclosure provides a method for predicting credit risk, including: obtaining target feature data of a credit entity, where the credit entity includes a target credit entity and non-target credit entities, and the target feature data is used to characterize the credit risk of the credit entity; performing knowledge extraction on the target feature data to generate a knowledge graph of the target credit entity, where the knowledge graph is used to characterize the entities, attributes of each credit entity in the credit entity, and the relationships among the credit entities; based on the knowledge graph, determining associated credit entities having a relational connection with the target credit entity from the non-target credit entities; and determining a credit risk prediction result of the target credit entity based on a first risk prediction result of the target credit entity and a second risk prediction result of the associated credit entities.
[0005] Optionally, the knowledge extraction of the above target feature data to generate the knowledge graph of the above target credit subject includes: determining the data structure of the above target feature data, where different data structures correspond to different knowledge extraction logics; selecting the corresponding knowledge extraction logic according to the above data structure; and performing knowledge extraction on the above target feature data based on the above corresponding knowledge extraction logic to generate the knowledge graph of the above target credit subject.
[0006] Optionally, before generating the knowledge graph of the above target credit subject, the above method further includes: constructing a domain ontology for describing the association information of the above credit subject; defining the category to which the above domain ontology belongs and the identifiers belonging to the above category through the Web Ontology Language, where the above identifiers include entity identifiers, attribute identifiers, and relationship identifiers, and the above entity identifiers, attribute identifiers, and relationship identifiers are stored in a graph database; and constructing an entity description framework for describing the association information of the above domain ontology based on the above entity identifiers, attribute identifiers, and relationship identifiers, where the above entity description framework is used to generate a knowledge graph.
[0007] Optionally, the above data structure includes a structured data structure, and the above performing knowledge extraction on the above target feature data based on the above corresponding knowledge extraction logic to generate the knowledge graph of the above target credit subject includes: invoking the knowledge extraction middleware of the above graph database to perform knowledge extraction on the above target feature data to obtain target field information, where the above target field information includes first entity information, first attribute information, and first relationship information; and using the above target field information to fill the above entity description framework to generate the knowledge graph of the above target credit subject, where the above knowledge graph includes the above first entity information corresponding to the above entity identifier, the above first attribute information corresponding to the above attribute identifier, and the above first relationship information corresponding to the above relationship identifier.
[0008] Optionally, the above data structure includes an unstructured data structure, and the above performing knowledge extraction on the above target feature data based on the above corresponding knowledge extraction logic to generate the knowledge graph of the above target credit subject includes: annotating the above target feature data to obtain annotation sequence information, where the above annotation sequence information includes second entity information and second attribute information; extracting the above target feature data based on a weakly supervised learning extraction method to obtain second relationship information, where the above second relationship information is used to represent the relationship between any two credit subjects in the above credit subject; and generating the knowledge graph of the above target credit subject based on the above second entity information, the above second attribute information, and the above second relationship information, where the above knowledge graph includes the above second entity information corresponding to the above entity identifier, the above second attribute information corresponding to the above attribute identifier, and the above second relationship information corresponding to the above relationship identifier.
[0009] Optionally, generating the knowledge graph of the target credit subject based on the above-mentioned second entity information, the above-mentioned second attribute information, and the above-mentioned second relationship information includes: determining a confidence value corresponding to the above-mentioned second relationship information; obtaining a confidence threshold of the relationship information; based on the above-mentioned confidence threshold, extracting third relationship information from the above-mentioned credit subjects whose confidence values meet the above-mentioned confidence threshold; and using the above-mentioned second entity information, the above-mentioned second attribute information, and the above-mentioned third relationship information to fill the above-mentioned entity description framework to generate the knowledge graph of the target credit subject, and the above-mentioned knowledge graph includes the above-mentioned third relationship information corresponding to the above-mentioned relationship identifier.
[0010] Optionally, the above method further includes: reasoning about the above knowledge graph through the inference rules of the Web Ontology Language to improve the above knowledge graph; and / or performing consistency detection on the categories to which the above domain ontology belongs to clean abnormal above categories.
[0011] Optionally, determining the associated credit subject having an associated relationship with the above target credit subject from the above non-target credit subjects includes: obtaining the shortest path including the above target credit subject according to a preset path direction, and the above preset path direction includes the out-degree and in-degree directions; obtaining the community division result of the above knowledge graph through a preset community discovery algorithm, and there is an associated relationship between the credit subjects in the same community; based on the above shortest path including the above target credit subject and / or the community division result of the above knowledge graph, determining the associated credit subject having an associated relationship with the above target credit subject from the above non-target credit subjects.
[0012] Optionally, the method further includes: obtaining a risk prediction model; inputting the target feature data of the above target credit subject into the above risk prediction model to obtain the first risk prediction result of the above target credit subject; and inputting the target feature data of the above associated credit subject into the above risk prediction model to obtain the second risk prediction result of the above associated credit subject.
[0013] Optionally, the method further includes: obtaining training sample data, and the above training sample data includes the feature data of credit subjects with normal credit and the feature data of credit subjects with low credit; and training the above training sample data to obtain the above risk prediction model.
[0014] Optionally, the method further includes: updating the above risk prediction model based on the credit risk prediction result of the above target credit subject.
[0015] Optionally, determining the credit risk prediction result of the target credit subject based on the first risk prediction result of the target credit subject and the second risk prediction result of the associated credit subject includes: when the first risk prediction result indicates that the credit of the target credit subject is abnormal, determining that the credit risk prediction result of the target credit subject is high risk; or when the first risk prediction result indicates that the credit of the target credit subject is normal and the second risk prediction result indicates that there is an associated credit subject with abnormal credit among the associated credit subjects, determining that the credit risk prediction result of the target credit subject is high risk; or when the first risk prediction result indicates that the credit of the target credit subject is normal and the second risk prediction result indicates that there is no associated credit subject with abnormal credit among the associated credit subjects, determining that the credit risk prediction result of the target credit subject is low risk.
[0016] To achieve the above object, another aspect of the present disclosure provides a prediction device for credit risk, including: a first acquisition module for acquiring target feature data of a credit subject, where the credit subject includes a target credit subject and non-target credit subjects, and the target feature data is used to characterize the credit risk of the credit subject; a generation module for performing knowledge extraction on the target feature data to generate a knowledge graph of the target credit subject, where the knowledge graph is used to characterize the entities, attributes of each credit subject in the credit subject itself, and the relationships between the credit subjects; a first determination module for determining, based on the knowledge graph, an associated credit subject having an associated relationship with the target credit subject from the non-target credit subjects; and a second determination module for determining the credit risk prediction result of the target credit subject based on the first risk prediction result of the target credit subject and the second risk prediction result of the associated credit subject.
[0017] Optionally, the generation module includes: a first determination sub-module for determining the data structure of the target feature data, where different data structures correspond to different knowledge extraction logics; a selection sub-module for selecting the corresponding knowledge extraction logic according to the data structure; and a generation sub-module for performing knowledge extraction on the target feature data based on the corresponding knowledge extraction logic to generate the knowledge graph of the target credit subject.
[0018] Optionally, before generating the knowledge graph of the target credit subject, the above device further includes: a first construction module for constructing a domain ontology for describing the association information of the credit subject; a definition module for defining the category to which the domain ontology belongs and the identifiers belonging to the category by using the Web Ontology Language, where the identifiers include entity identifiers, attribute identifiers, and relationship identifiers, and the entity identifiers, attribute identifiers, and relationship identifiers are stored in a graph database; and a second construction module for constructing an entity description framework for describing the association information of the domain ontology based on the entity identifiers, attribute identifiers, and relationship identifiers, where the entity description framework is used to generate a knowledge graph.
[0019] Optionally, the data structure includes a structured data structure, and the generation sub-module includes: a first extraction unit for invoking the knowledge extraction middleware of the graph database to perform knowledge extraction on the target feature data to obtain target field information, where the target field information includes first entity information, first attribute information, and first relationship information; and a first generation unit for using the target field information to fill the entity description framework to generate the knowledge graph of the target credit subject, where the knowledge graph includes the first entity information corresponding to the entity identifier, the first attribute information corresponding to the attribute identifier, and the first relationship information corresponding to the relationship identifier.
[0020] Optionally, the data structure includes an unstructured data structure, and the generation sub-module includes: a labeling unit for labeling the target feature data to obtain labeling sequence information, where the labeling sequence information includes second entity information and second attribute information; a second extraction unit for extracting the target feature data based on a weakly supervised learning extraction method to obtain second relationship information, where the second relationship information is used to represent the relationship between any two credit subjects in the credit subject; and a second generation unit for generating the knowledge graph of the target credit subject based on the second entity information, the second attribute information, and the second relationship information, where the knowledge graph includes the second entity information corresponding to the entity identifier, the second attribute information corresponding to the attribute identifier, and the second relationship information corresponding to the relationship identifier.
[0021] Optionally, the second generating unit includes: a determining subunit, configured to determine a confidence value corresponding to the second relationship information; an obtaining subunit, configured to obtain a confidence threshold of the relationship information; an extracting subunit, configured to extract, based on the confidence threshold, third relationship information from the credit subjects whose confidence values meet the confidence threshold; and a generating subunit, configured to generate a knowledge graph of the target credit subject by filling the entity description framework with the second entity information, the second attribute information, and the third relationship information, where the knowledge graph includes the third relationship information corresponding to the relationship identifier.
[0022] Optionally, the apparatus further includes: an inference module, configured to infer the knowledge graph through inference rules of the Web Ontology Language to improve the knowledge graph; and / or a detection module, configured to perform consistency detection on the categories to which the domain ontology belongs to clean abnormal categories.
[0023] Optionally, the first determining module includes: a first obtaining sub-module, configured to obtain the shortest path including the target credit subject according to a preset path direction, where the preset path direction includes an out-degree direction and an in-degree direction; a second obtaining sub-module, configured to obtain a community division result of the knowledge graph through a preset community discovery algorithm, where there is an association relationship between credit subjects in the same community; and a second determining sub-module, configured to determine, based on the shortest path including the target credit subject and / or the community division result of the knowledge graph, an associated credit subject having an association relationship with the target credit subject from the non-target credit subjects.
[0024] Optionally, the apparatus further includes: a first obtaining module, configured to obtain a risk prediction model; a second obtaining module, configured to input the target feature data of the target credit subject into the risk prediction model to obtain a first risk prediction result of the target credit subject; and a third obtaining module, configured to input the target feature data of the associated credit subject into the risk prediction model to obtain a second risk prediction result of the associated credit subject.
[0025] Optionally, the apparatus further includes: a second obtaining module, configured to obtain training sample data, where the training sample data includes feature data of credit subjects with normal credit and feature data of credit subjects with low credit; and a training module, configured to train the training sample data to obtain the risk prediction model.
[0026] Optionally, the apparatus further includes: an updating module, configured to update the risk prediction model based on the credit risk prediction result of the target credit subject.
[0027] Optionally, the second determination module includes: a third determination sub-module, configured to determine that the credit risk prediction result of the target credit entity is a high risk when the first risk prediction result indicates that the credit of the target credit entity is abnormal; or a fourth determination sub-module, configured to determine that the credit risk prediction result of the target credit entity is a high risk when the first risk prediction result indicates that the credit of the target credit entity is normal and the second risk prediction result indicates that there is an associated credit entity with abnormal credit among the associated credit entities; or a fifth determination sub-module, configured to determine that the credit risk prediction result of the target credit entity is a low risk when the first risk prediction result indicates that the credit of the target credit entity is normal and the second risk prediction result indicates that there is no associated credit entity with abnormal credit among the associated credit entities.
[0028] To achieve the above object, another aspect of the present disclosure provides an electronic device, including: one or more processors, and a memory for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the credit risk prediction method as described above.
[0029] To achieve the above object, another aspect of the present disclosure provides a computer-readable storage medium storing computer-executable instructions that are used to implement the credit risk prediction method as described above when executed.
[0030] To achieve the above object, another aspect of the present disclosure provides a computer program, the computer program including computer-executable instructions that are used to implement the credit risk prediction method as described above when executed.
[0031] According to the embodiments of the present disclosure, based on a knowledge base that realizes data management and storage through link data, the implementation of the risk prediction of the knowledge graph technology in the vertical field of credit entities can at least partially solve the technical problems in the related art of the credit risk prediction method, where the credit entity is statistically analyzed and compared as an independent individual, without effectively using the association relationship existing between credit entities, resulting in low accuracy of the credit risk prediction result of the credit entity, difficult analysis of the association relationship of the credit entity, untimely discovery of the transmission of credit risk, and low data analysis dimension. Therefore, it can achieve the technical effects of easy expansion of analysis data, strong reasoning ability of association relationship, and easy search. BRIEF DESCRIPTION OF THE DRAWINGS
[0032] By reading the following detailed description with reference to the accompanying drawings, the above and other objects, features, and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the drawings, several embodiments of the present disclosure are shown by way of example and not limitation, wherein:
[0033] Figure 1 Schematically shows a system architecture applicable to the embodiments of the present disclosure;
[0034] Figure 2 Schematically shows a flowchart of a method for predicting credit risk according to an embodiment of the present disclosure;
[0035] Figure 3 Schematically shows an entity description framework diagram according to an embodiment of the present disclosure;
[0036] Figure 4 Schematically shows a flowchart of structured data knowledge extraction according to an embodiment of the present disclosure;
[0037] Figure 5 Schematically shows a flowchart of unstructured data knowledge extraction according to an embodiment of the present disclosure;
[0038] Figure 6 Schematically shows a flowchart of a method for predicting credit risk according to another embodiment of the present disclosure;
[0039] Figure 7 Schematically shows a flowchart of a method for predicting credit risk according to another embodiment of the present disclosure;
[0040] Figure 8 Schematically shows a block diagram of a device for predicting credit risk according to an embodiment of the present disclosure;
[0041] Figure 9 Schematically shows a schematic diagram of a computer-readable storage medium product suitable for implementing the method for predicting credit risk described above according to an embodiment of the present disclosure; and
[0042] Figure 10 Schematically shows a block diagram of an electronic device suitable for implementing the method for predicting credit risk described above according to an embodiment of the present disclosure. Detailed Embodiments
[0043] Hereinafter, embodiments of the present disclosure will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present disclosure. In the following detailed description, for the sake of explanation, numerous specific details are set forth in order to provide a comprehensive understanding of the embodiments of the present disclosure. However, it is obvious that one or more embodiments may be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily obscuring the concepts of the present disclosure.
[0044] The terms used herein are for the purpose of describing particular embodiments only and are not intended to limit the present disclosure. The terms such as "including" and "comprising" used herein indicate the presence of the above-mentioned features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components. All terms (including technical and scientific terms) used herein have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.
[0045] In cases where expressions such as "at least one of A, B, and C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.). In cases where expressions such as "at least one of A, B, or C, etc." are used, generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, or C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).
[0046] Some block diagrams and / or flowcharts are shown in the accompanying drawings. Some blocks or combinations of blocks in the block diagrams and / or flowcharts can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable credit risk prediction devices, so that when these instructions are executed by the processor, they can create a device for implementing the functions / operations illustrated in these block diagrams and / or flowcharts. The technology of the present disclosure can be implemented in the form of hardware and / or software (including firmware, microcode, etc.). In addition, the technology of the present disclosure can take the form of a computer program product on a computer-readable storage medium storing instructions, which can be used by or in conjunction with an instruction execution system.
[0047] Most of the related technologies obtain the prediction results of the credit risk of a credit subject by statistically analyzing and comparing the relevant data of the credit subject itself. Since the credit subject is regarded as an independent individual, the associated relationship formed between the credit subject and other credit subjects is not effectively utilized, and the dimension of data analysis is low, resulting in low accuracy of the prediction results of the credit risk and the inability to timely detect the transmission of the credit risk among credit subjects.
[0048] Therefore, the present disclosure provides a method for predicting credit risk, including a knowledge graph construction stage and a credit risk prediction stage. In the knowledge graph construction stage, first, target feature data of credit entities is obtained. The credit entities include target credit entities and non-target credit entities, and the target feature data is used to characterize the credit risk of the credit entities. Then, knowledge extraction is performed on the target feature data to generate a knowledge graph of the target credit entities, which is used to characterize the entities, attributes corresponding to each credit entity in the credit entities, and the relationships between the credit entities. In the credit risk prediction stage, first, based on the knowledge graph, associated credit entities having an association relationship with the target credit entity are determined from the non-target credit entities. Then, based on the first risk prediction result of the target credit entity and the second risk prediction result of the associated credit entities, the credit risk prediction result of the target credit entity is determined.
[0049] Risk prediction is an essential business process in various industries. For example, in the financial industry, it is necessary to monitor whether there is a default risk for the issuer of virtual resources. It should be noted that the method and device for predicting credit risk provided by the present disclosure can be used in the financial field and can also be used in any field other than the financial field. Therefore, the application fields of the method and device for predicting credit risk provided by the present disclosure are not limited.
[0050] Figure 1 Schematically shows a system architecture 100 applicable to the embodiments of the present disclosure. It should be noted that Figure 1 The shown is only an example of the system architecture to which the embodiments of the present disclosure can be applied, to help those skilled in the art understand the technical content of the present disclosure, but it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0051] As Figure 1 shown, the system architecture 100 according to this embodiment may include terminal devices 101, 102, 103, a network 104, and a server 105. The network 104 is used to provide a medium for communication links between the terminal devices 101, 102, 103 and the server 105. The network 104 may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.
[0052] Users can use the terminal devices 101, 102, 103 to interact with the server 105 through the network 104 to receive or send messages, etc. Various communication client applications may be installed on the terminal devices 101, 102, 103, such as shopping applications, web browser applications, search applications, instant messaging tools, email clients, social platform software, etc. (only as examples). The terminal devices 101, 102, 103 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers, and desktop computers, etc.
[0053] The server 105 can be a server that provides various services. For example, it can be a background management server (only for illustration) that supports the websites browsed by users using the terminal devices 101, 102, and 103. The background management server can analyze and process data such as user requests received, and feedback the processing results (such as web pages, information, or data obtained or generated according to user requests) to the terminal devices.
[0054] It should be noted that the method for predicting credit risk provided by the embodiments of the present disclosure can generally be executed by the server 105. Correspondingly, the device for predicting credit risk provided by the embodiments of the present disclosure can generally be set in the server 105. The method for predicting credit risk provided by the embodiments of the present disclosure can also be executed by a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. Correspondingly, the device for predicting credit risk provided by the embodiments of the present disclosure can also be set in a server or a server cluster different from the server 105 and capable of communicating with the terminal devices 101, 102, 103 and / or the server 105. It should be understood that Figure 1 the numbers of the terminal devices, the network, and the servers in
[0055] Figure 2 schematically shows a flowchart of the method for predicting credit risk according to an embodiment of the present disclosure.
[0056] As Figure 2 shown, the method 200 can include operation S210 to operation S240.
[0057] In operation S210, target feature data of a credit subject is obtained. In the present disclosure, the credit subject includes a target credit subject and a non-target credit subject. As the parties to a credit relationship, the credit subject is the bearer of the credit relationship and the actor of credit activities, such as an organizational entity, a personal entity, etc. Among them, the organizational entity can include, but is not limited to, an enterprise, an investment company, a bank, and the personal entity can include, but is not limited to, a key person of an enterprise, an investor.
[0058] According to an embodiment of the present disclosure, the target feature data is used to characterize the credit risk of the credit subject. The target feature data can be in-house data, such as credit data, financial data, and banking regulatory data, etc., or can be directly pulled from a specified database. For example, the credit data can be pulled from the database corresponding to the Credit Reference Center of the People's Bank of China, the financial data can be pulled from the database corresponding to a financial website, and the banking regulatory data can be pulled from the banking regulatory database of the China Banking Regulatory Commission. The target feature data can also be external data, such as legal data, public opinion data, industry regional data, real estate data, customs data, news data, etc.
[0059] In operation S220, knowledge extraction is performed on the target feature data to generate a knowledge graph of the target credit subject. In the present disclosure, the knowledge graph is used to represent the entities, attributes of each credit subject in the credit subject, and the relationships between the credit subjects. The knowledge graph is a structured symbolic representation of the objective physical world and is also a networked knowledge base. It is composed of entities with attributes linked by relationships, and the relationships also include their own attributes. From the perspective of graph theory, the knowledge graph is essentially a concept network, where its nodes represent entities in the objective physical world, and the edges represent various semantic relationships existing between the entities. The key point in constructing the knowledge graph lies in the mining of relationships between enterprises. There are various relationships between enterprises and between enterprises and individuals. Through these relationships, an enterprise relationship network, that is, an enterprise knowledge graph, can be constructed. Constructing an enterprise knowledge graph can help us mine the potential associations of enterprises from a large amount of messy data and generate an enterprise portrait.
[0060] In operation S230, based on the knowledge graph, associated credit subjects that have an associated relationship with the target credit subject are determined from non-target credit subjects. In the present disclosure, the associated credit subjects can be some of the non-target credit subjects or all of the non-target credit subjects.
[0061] In operation S240, based on the first risk prediction result of the target credit subject and the second risk prediction result of the associated credit subjects, the credit risk prediction result of the target credit subject is determined.
[0062] According to an embodiment of the present disclosure, the credit risk prediction result of the target credit subject is not only related to the first risk prediction result of the target credit subject but also related to the second risk prediction result of the associated credit subjects. The two jointly determine the credit risk prediction result, making efficient use of the bank's information system and enterprise big data in the Internet, integrating isolated data nodes into a unified knowledge base, breaking the enterprise isolated points, and realizing the interconnection and interoperability of customer enterprise information.
[0063] Through the embodiments of the present disclosure, the value of enterprise risk data is fully explored using big data to obtain the explicit and implicit association relationships between individuals and legal persons, identify groups composed of entities with certain common characteristics, calculate the process and probability of transmission of a certain event among associated entities, and provide a visual view of the litigation risks related to enterprise industry and commerce for customer managers. This can help banks timely predict potentially risky associated enterprises before granting loans, make early warnings and pre-judgments. At the same time, it can also help banks timely discover potential risks after granting loans, initiate the collection process in advance, and effectively reduce the losses of non-performing loans of banks.
[0064] As an alternative embodiment, extracting knowledge from target feature data to generate a knowledge graph of a target credit subject includes: determining the data structure of the target feature data, where different data structures correspond to different knowledge extraction logics; selecting the corresponding knowledge extraction logic according to the data structure; and based on the corresponding knowledge extraction logic, extracting knowledge from the target feature data to generate a knowledge graph of the target credit subject.
[0065] According to the embodiments of the present disclosure, the extraction methods of the target feature data entities with different data structures and the mutual relationships between entities are different. Since the target feature data refers to the data that can characterize the possibility of a monitored object having a credit default behavior, such as credit records, financial data, etc. Therefore, the data structure of the target feature data can be images, audio, text, and numbers, and the corresponding data structures can be structured data structures and unstructured data structures.
[0066] Through the embodiments of the present disclosure, different knowledge extraction logics are provided for different data structures, so that the construction of the knowledge graph is based on data with different data structures, and can more comprehensively and objectively reflect the association relationships between entities.
[0067] As an alternative embodiment, before generating the knowledge graph of the target credit subject, the method further includes: constructing a domain ontology for describing the association information of the credit subject; defining the category to which the domain ontology belongs and the identifiers belonging to the category through the Web Ontology Language, where the identifiers include entity identifiers, attribute identifiers, and relationship identifiers, and the entity identifiers, attribute identifiers, and relationship identifiers are stored in a graph database; and based on the entity identifiers, attribute identifiers, and relationship identifiers, constructing an entity description framework for describing the association information of the domain ontology, where the entity description framework is used to generate the knowledge graph.
[0068] The present disclosure introduces Ontology, that is, a formal specification of a shared conceptual model. In the computer field, Ontology describes knowledge at the semantic level and can be regarded as a general conceptual model that describes the knowledge of a certain discipline field. This model contains the basic terms in a certain discipline field and the relationships between the terms, or is called concepts and the relationships between concepts. Ontology is the consensus of the group and is a recognized set of concepts in the corresponding field, which is not equivalent to an individual. Through reasoning with the domain Ontology, the domain Ontology studies the concepts and the relationships between concepts in a specific field.
[0069] When the present disclosure abstracts and models the concept of enterprise Ontology according to the characteristics of the data involved in the field of enterprise association information, the properties set may include but are not limited to background information properties, enterprise operation status properties, main personnel properties, and historical risk properties. Further, corresponding sub-properties can also be set for each property, and the corresponding relationships between the property and the sub-properties are shown in Table 1.
[0070] In specific implementation, the Web Ontology Language (OWL) for semantic description of Ontology can be used to define sub-properties. For example, rdfs:subPropertyOf in OWL can be used to define sub-properties. It should be noted that the properties and their corresponding sub-properties shown in Table 1 are only illustrative. According to actual needs, properties different from those shown in Table 1 and sub-properties corresponding to the properties can be defined by oneself, and the present disclosure does not limit this.
[0071] Table 1
[0072]
[0073] When defining the data schema of Ontology in the OWL Ontology language, it is mainly necessary to define classes, sub-classes, properties, sub-properties, and the internal logical relationships of properties. In the present disclosure, through analyzing the relevant domain knowledge of enterprise association information, the classes, properties, and some sub-properties of enterprises are defined as shown in Table 2. It should be noted that the classes, properties, and some sub-properties of enterprises shown in Table 2 are only illustrative. According to actual needs, classes, properties, and some sub-properties of enterprises different from those shown in Table 2 can be defined by oneself, and the present disclosure does not make specific limitations on this.
[0074] As an alternative embodiment, in the present disclosure, in addition to defining classes and attributes, relationships between entities can also be defined. The relationships between entities can include relationship names, relationship types, and one-way relationships between entities. The one-way relationship can include a relationship start point and a relationship end point. Among them, the relationship type represents different types and scopes of relationships, including legal relationships, market relationships, and social relationships. Specifically, market relationships are currently mainly relationships generated by enterprise business operations. The scope of effect of social relationships is a type of relationship that is effective within the scope of personal social interactions. And legal relationships represent categories of relationships restricted by law. The relationships between entities are shown in Table 3. It should be noted that the relationships between entities shown in Table 3 are only illustrative, and different relationships between entities from those shown in Table 3 can be defined according to actual needs.
[0075] Table 2
[0076]
[0077] Table 3
[0078]
[0079] According to an embodiment of the present disclosure, using the classes and attributes defined by the ontology language, a Resource Description Framework (RDF) graph can be drawn. RDF is also an entity description framework graph, which is a data model represented using XML syntax, used to describe the characteristics of Web resources and the relationships between resources. RDF provides a general framework for expressing this information and enabling it to be exchanged between applications without losing semantics for occasions where information needs to be processed by applications rather than just displayed to people.
[0080] Figure 3 Schematically shows an entity description framework graph according to an embodiment of the present disclosure. As Figure 3As shown, in the RDF graph 300 of the present disclosure, a solid line can be used to identify the association relationship between class attributes and object attributes, and a dashed line can be used to identify the association relationship between object attributes and sub - attributes. For example, a solid line is used to identify the association relationship between the class attribute (Company) and the object attributes such as background information attribute, management state attribute, historical risk attribute, and key personnel attribute. A dashed line is used to identify the association relationship between the object attribute key personnel and its sub - attributes legalperson and share holder. In the present disclosure, the meanings of other class attributes, object attributes, and sub - attributes can be specifically referred to in Tables 1 - 3 mentioned above, and will not be elaborated here.
[0081] As an alternative embodiment, the data structure includes a structured data structure. Based on the corresponding knowledge extraction logic, knowledge extraction of target feature data to generate a knowledge graph of the target credit subject includes: invoking the knowledge extraction middleware of the graph database to perform knowledge extraction on the target feature data to obtain target field information, where the target field information includes first entity information, first attribute information, and first relationship information; and using the target field information to fill the entity description framework to generate a knowledge graph of the target credit subject, where the knowledge graph includes the first entity information corresponding to the entity identifier, the first attribute information corresponding to the attribute identifier, and the first relationship information corresponding to the relationship identifier.
[0082] In the present disclosure, a part of the structured data is raw data, which can be database data and Json data provided by the cooperative enterprises of the bank, and a part is additional data, which can be the encyclopedia web page data of relevant enterprises crawled through web crawlers. On the one hand, as additional data, the encyclopedia web page data can well supplement the raw data. For example, the field contents such as aliases and abbreviations in the encyclopedia web page data can expand the expression of entities, which is very helpful for establishing an entity synonym table in the following text. On the other hand, the encyclopedia web page data has strong real - time performance and can also play a certain role in updating the raw data.
[0083] In the present disclosure, the structured data not only includes the basic information data of entities, but also includes the relationship information data between entities. However, for both the basic information data of entities and the relationship information data between entities, the acquisition of these data requires corresponding the structured data with the entities and attributes defined in the graph database, and extracting the relationships between entities. In specific implementation, by analyzing the field information of the structured data, it can be found that the BasicInfo field in the Json data corresponds to the basic information of the entities in the graph database, and can include enterprise basic information data such as registered capital, establishment date, business status, industrial and commercial registration number, etc. The JudicialRisk field contains enterprise risk information, the StockholderInfo field contains the shareholder information of the enterprise, the KeyPersonInfo field contains the information of the key persons of the enterprise, the InvestListInfo field contains the external investment data of the enterprise, and the EnterpriseRelationship field contains branch and related transaction data.
[0084] In specific implementation, the middleware APOC of the graph database can be used to extract the structured data to obtain the entities and attributes corresponding to the structured data, and import the entities into the graph database and complete the mapping of attributes and the generation of relationships between entities. APOC supports the parsing of Json data. Therefore, for the data obtained from different data sources, it can be selected to first convert the data from different data sources into multiple data with a unified Json data format, and then implement the batch import of the data. It can also be selected to first convert the data from different data sources into multiple data with a unified Json data format and then splice them, and then implement the batch import of the data. After the batch import of the data is completed, the background information data, business status data, main personnel data, and historical risk data will be stored as the attributes of the enterprise nodes. According to the relationship information contained in the StockholderInfo, KeyPersonInfo, and InvestListInfo fields, the construction of various association relationships between entities can be completed, which can include but are not limited to the shareholding relationship of a person and an enterprise entity in a certain enterprise, the employment relationship of the key persons within the enterprise, and the investment relationship and related transaction relationship of an enterprise with other enterprises.
[0085] Figure 4 Schematically shows the flowchart of structured data knowledge extraction according to an embodiment of the present disclosure. As Figure 4 shown, the method 400 may include operation S410 to operation S4120.
[0086] In operation S410, the data is encoded into UTF-8 (a coding format where one byte contains 8 bits) through transcoding. In operation S420, the data format is unified into Json format. In operation S430, the APOC middleware is used to read the Json file. In operation S440, mention extraction is performed to obtain the mentions in the text, and the mention can be the aforementioned target field information. In operation S450, it is determined whether the entity exists in the knowledge base. If so, operation S460 is executed to update the entity attributes. If not, operation S470 is executed to create a new entity. In operation S480, the associated entities are traversed. In operation S490, it is determined whether the associated entity exists in the knowledge base. If not, operation S4100 is executed to create a new entity. If so, operation S4110 is executed to update the entity attributes. Finally, in operation S4120, the relationships between entities are established.
[0087] Through the embodiments of the present disclosure, the middleware of the graph database extracts structured data, imports the entities into the graph database and completes the mapping of attributes and the generation of relationships between entities, realizing the preparation and extraction of structured data, and providing support for the construction of the knowledge graph.
[0088] As an alternative embodiment, the data structure includes an unstructured data structure. Based on the corresponding knowledge extraction logic, knowledge extraction of the target feature data to generate a knowledge graph of the target credit subject includes: annotating the target feature data to obtain annotation sequence information, where the annotation sequence information includes second entity information and second attribute information; extracting the target feature data based on the extraction method of weakly supervised learning to obtain second relationship information, where the second relationship information is used to characterize the relationship between two credit subjects in the credit subject; and generating a knowledge graph of the target credit subject based on the second entity information, second attribute information and second relationship information, where the knowledge graph includes second entity information corresponding to the entity identifier, second attribute information corresponding to the attribute identifier, and second relationship information corresponding to the relationship identifier.
[0089] In the present disclosure, the unstructured text data adopts the classic Bidirectional Long Short-Term Memory Neural Network-Conditional Random Field (BiLSTM-CRF) model, which can convert the named entity recognition task into a labeling problem of the input sequence by character.
[0090] For example, the annotation example of the compensation relationship is as follows:
[0091] (Original text) 1. The defendants Zhang San and Li Si shall return the loan principal of 1,097,250.07 yuan to the plaintiff, a certain limited company, A Sub-branch, and pay the overdue interest of 82,129.67 yuan as of March 11, 2014, totaling 1,179,378.74 yuan.
[0092] (Annotation Result) First, [Defendants Zhang San and Li Si / payer] shall [return / act] to [Plaintiff A Sub-branch of Company XX / payee] the [loan principal / type] of [RMB 1,097,250.07 / amt], and [pay / act] the [arrears interest / type] of [RMB 82,129.67 / amt] as of March 11, 2014, with a total of [RMB 1,179,378.74 / total].
[0093] In this disclosure, after entity recognition for unstructured text data is completed, the obtained entities (including enterprise entities and / or person entities) are still discrete and unassociated nodes. Therefore, to construct a knowledge graph to form an association relationship network between enterprise entities and person entities, it is necessary to extract the relationships between entities. News data of relevant enterprises crawled by a crawler can be used to extract the association relationships between entities, which may include the employment relationship between a person entity and an enterprise entity, the investment relationship between enterprise entities, or the related transaction relationship related to equity transactions between enterprise entities.
[0094] Figure 5 Schematically shows a flowchart of knowledge extraction for unstructured data according to an embodiment of the present disclosure. As Figure 5 shown, the method 500 may include operations S510 to S590.
[0095] In operation S510, data preparation is performed. Specifically, prior data is imported during implementation. Since DeepDive realizes weakly supervised relation extraction, preferably, the training data is the already determined relationships between entities. Enterprise entities with existing transaction relationships can be used as training data, and this prior data mainly comes from equity transaction information in the knowledge base. In operation S520, data is stored in the database. The article to be extracted is imported. Specifically, during implementation, a large number of news texts crawled by the crawler are first converted into table files and placed in the input folder as the article to be extracted. Then, corresponding data tables are established in the main program file app.ddlog of DeepDive and imported into the database. In operation S530, natural language processing is carried out. Natural language processing methods are used to process the text to obtain sequences such as NER / POS / lexical dependencies of the text data. Taking sentences as units, the word segmentation, part-of-speech standard (POS), named entity recognition (NER), and syntactic analysis results of each sentence are returned, and these results are stored in the sentence table to prepare for subsequent feature extraction. In operation S540, match the known relationships of candidate entity pairs. According to the data table of entity pairs with known relationships, the known variable table is obtained. Define the data table for storing the results, and the labeled results are used as prior variables. Extract the text features of candidate entity pairs. First, a feature table is defined. The input of this feature table is the entity pair table and the text table, and the input and output attributes are in the main program file. The feature function is implemented by the ddlib library of DeepDive. After obtaining the window features, they are input into the feature table. The feature generation results are shown in Table 4. In operation S550, mention extraction is performed to obtain the mentions in the text, and the mention can be the aforementioned target field information. In operation S560, part of the data is labeled based on rules. Specifically, for the samples, positive example samples and negative example samples are marked. The prior data of candidate entities and known relationships are associated, and corresponding labels are assigned to part of the data through rules. First, a label table is defined to store the supervised data for labeling. Then, the prepared database data is imported into the table, the id of the rule is set and the corresponding weight is set. If the data has high credibility, a higher weight can be set for it. Correspondingly, if the data has low credibility, a lower weight can be set for it. Then, the labeling function is called to store the extracted data in the label table. It should be noted that since different rules may cover the same entity pair, resulting in different, even opposite, labeling results, in this disclosure, the labels between entity pairs need to be unified, and the labeling results are added. After completing the labeling using different rules, the labeling results can be statistically analyzed to obtain the labels.
[0096] Table 4
[0097]
[0098] In operation S570, candidate entity pairs are obtained. For entity extraction and generation of candidate entity pairs, a total variable table is obtained. In this step, candidate entities in the text need to be extracted. When extracting the relationship between enterprises and persons, candidate entities of enterprises and candidate entities of persons need to be obtained, and candidate entity pairs are generated. Combining with the known variable table, a total variable table is generated. In operation S580, feature extraction of entity pairs can obtain a feature table. In operation S590, a factor graph is constructed based on the total variable table and the feature table to obtain variable confidence. A factor graph is constructed and a probability model is generated. The entity pairs and the feature table are connected. Through the connection of feature factors, global learning of the weights of these features is performed. The program is compiled and the program is executed to generate a probability model.
[0099] As an alternative embodiment, generating a knowledge graph of a target credit subject based on second entity information, second attribute information, and second relationship information includes: determining a confidence value corresponding to the second relationship information; obtaining a confidence threshold of the relationship information; based on the confidence threshold, extracting third relationship information from the credit subjects whose confidence values meet the confidence threshold, and using the second entity information, second attribute information, and third relationship information to fill an entity description framework to generate a knowledge graph of the target credit subject, where the knowledge graph includes third relationship information corresponding to a relationship identifier.
[0100] According to an embodiment of the present disclosure, an information extraction tool DeepDive can be used to perform weakly supervised extraction of relationships between entities, and import the relationship extraction results with confidence higher than a preset threshold into a knowledge base.
[0101] According to an embodiment of the present disclosure, after completing the extraction of the association relationship between entities by DeepDive, the confidence result of the association relationship between entities can be finally obtained, as shown in Table 5. When the preset threshold of the confidence is 0.85, the relationship data with confidence not lower than 0.85 is considered as the relationship data with high reliability in the extraction results, and it can be imported into the knowledge base. Correspondingly, the relationship data with confidence lower than 0.85 is considered as the relationship data with low reliability in the extraction results, and it may not be imported into the knowledge base.
[0102] Table 5
[0103]
[0104] As an alternative embodiment, the method further includes: reasoning about the knowledge graph through inference rules of the Web Ontology Language to improve the knowledge graph; and / or performing consistency detection on the categories to which the domain ontology belongs to clean abnormal categories.
[0105] In specific implementation, the inference engine of the Jena inference machine can be used for ontology inference. In the first step, the most important data structure in the inference machine is the Model model object. The present disclosure constructs and initializes the model object using the factory class provided by Jena, including the top-level data schema ontology and triple knowledge. In the second step, a custom inference machine is generated. In specific implementation, it can be implemented based on the register provided by Jena and associated with the Model model object to obtain an InfModel, thereby endowing the model object with inference capabilities. In the third step, according to business requirements, the program interface provided by the inference machine is used for inference. The data of the Jena inference machine includes two parts. One part of the data is the triple knowledge in RDF format, that is, the domain enterprise entities, person entities, and the relationships between entities. The other part of the data is the upper-layer data constraint model, that is, the ontology information. By judging the hierarchical relationship of categories, it is determined whether entities have a superordinate or subordinate relationship, and the categories to which the entities belong are improved. Through ontology inference, it is detected whether two incompatible types are defined when defining the enterprise entity in terms of category. Through custom rule inference, the implicit relationships between entities are inferred and complemented. According to the constructed domain ontology of enterprise association information and Jena inference rules, the subclass relationship of categories can be defined using the special vocabulary subClassOf, and the implicit superordinate or subordinate relationship between categories can be obtained based on its transitivity. By introducing an ontology inference machine for ontology inference, the implicit category where a certain entity is located is inferred, and the category to which the entity belongs is complemented. The detection of inconsistencies can be verified through the verification interface of Jena, and the inconsistencies in the data schema definition categories can be obtained. The Jena inference machine performs rule inference by applying the Semantic Web Rule Language SWRL. By writing corresponding rules on the inference machine, multiple rules can also be customized in the inference rules for inference.
[0106] Figure 6 Schematically shows a flowchart of a method for predicting credit risk according to another embodiment of the present disclosure. As Figure 6 shown, the method may include operation S610 to operation S640.
[0107] In operation S610, custom rule reasoning is performed. Specifically, in implementation, implicit triple relationships between entities in the field of enterprise association information can be obtained based on custom rule reasoning. By formulating rules for reasoning, generally, more knowledge is inferred through existing RDF triple knowledge and designed rules. First, corresponding rules need to be designed, and then these rules are transformed into Jena reasoning statements according to the ontology and the reasoning specifications provided by Jena. The style of the reasoning file can generally be expressed as [Rule name: (RDF triple) (RDF triple) → (inferred triple)]. The designed rules are called by the rule name, and reasoning is performed in combination with the existing triple knowledge. Corresponding rules are designed according to the characteristics of the data in the field of enterprise association information as follows.
[0108] (1) Enterprise: has_share(X, Y): - enterprise: control(X, Y)
[0109] (2) Enterprise: has_transaction(Y, Z): - enterprise: has_share(X, Y), enterprise: has_share(X, Z)
[0110] (3) enterprise: subsidiary(X, Y): - enterprise: subsidiary(X, Y), enterprise: subsidiary(Y, Z)
[0111] The defined rules are respectively: Those who actually control an enterprise are also shareholders of the enterprise; If an enterprise holds shares in two enterprises respectively, then there is also a transaction relationship between these two enterprises; If the subsidiary of enterprise X is Y and the subsidiary of Y is Z, then Z is also a subsidiary of X.
[0112] In operation S620, ontology super-subclass reasoning is performed. Specifically, in implementation, the relationships between categories and the transitivity of categories are defined in the ontology. Between two defined categories, if Human is a subclass of Living Thing, then there is a super-subclass relationship between Human and Living Thing. In the inference engine, subClassOf is used to define subclasses and subPropertyOf is used to define sub-properties to determine whether there is a super-subclass relationship between concept 1 and concept 2. To determine the super-subclass relationship between the two concepts "real estate enterprise" and "enterprise", the inference engine needs to traverse all the upper categories defined for the "real estate enterprise" category. If it is found that the concept of "enterprise" is defined upstream of it, then it is determined that there is a super-subclass relationship between these two concepts.
[0113] In operation S630, entity category completion is performed. Specifically, in the ontology language, corresponding relationships can be defined between categories, such as the subcategory relationship, the disjoint relationship, etc. Generally, when an entity is imported into the knowledge base, only its belonging to a certain category is defined. However, according to the relationships between categories defined in the ontology, an entity may also belong to other categories. By applying an OWL reasoner, the reasoning is completed to supplement the categories where the entity is located. In the original data, a certain group belongs to the category of real estate enterprises. After supplementing the category, the group also belongs to the category of enterprises and the category of real estate enterprises at the same time.
[0114] In operation S640, category inconsistency detection is performed. Specifically, when the ontology is defined, two classes are defined and the relationship between them is the disjoint relationship. However, an entity belongs to these two categories, which indicates that there is an abnormality in a triple knowledge. It is necessary to find and return these two pieces of knowledge through inconsistency detection and clean the abnormal data among them. In the data schema, the relationship between the enterprise category and the person category is defined as disjoint, that is, a mutually exclusive relationship. Traverse the triple knowledge in the reasoner, return all entities with inconsistent category settings in the entities, and then clear the abnormal data.
[0115] Through the embodiments of the present disclosure, the knowledge graph of enterprise association information is supplemented and improved through knowledge base reasoning. Some hidden knowledge and conclusions need to be obtained through knowledge reasoning. The reasoning engine of the Jena reasoner can be used for ontology reasoning to supplement and improve the constructed knowledge graph.
[0116] As an alternative embodiment, determining an associated credit entity having an association relationship with a target credit entity from non-target credit entities includes: obtaining the shortest path including the target credit entity according to a preset path direction, where the preset path direction includes the out-degree direction and the in-degree direction; obtaining the community division result of the knowledge graph through a preset community discovery algorithm, where credit entities in the same community have an association relationship; and determining an associated credit entity having an association relationship with the target credit entity from non-target credit entities based on the shortest path including the target credit entity and / or the community division result of the knowledge graph.
[0117] Specifically, when implemented, it is determined whether there is a close relationship between an enterprise node and a low-credit entity through the shortest path and the social network analysis method, and the characteristics of the external relationships of the enterprise are extracted. A risk control model is obtained by learning the characteristics of normal credit enterprises and low-credit enterprises. The characteristic data of the three dimensions of the basic attributes, historical risks, and external relationships of the enterprise are used as the input of the risk control model to identify the enterprise risks.
[0118] Among them, the characteristic variables of the basic enterprise attributes are used to describe the basic information of the enterprise, which can include the registered capital, establishment time, business status, taxpayer qualification, paid-in capital, personnel scale, number of insured persons, tax rating, number of financing times, amount of financing, number of investment times, and amount of investment. In addition, some derived variables can also be obtained as shown in Table 6.
[0119] Among them, the variables of the enterprise historical risks depict the risk information of the enterprise, including the number of litigation-related information, the number of administrative penalty times, the number of equity pledge times, the amount of equity pledge, the number of chattel mortgage times, and the number of legal representative changes, as shown in Table 7.
[0120] Among them, the variables of the enterprise external relationships depict whether the enterprise is closely associated with other high-risk nodes in the graph network structure and its characteristic information in the network, including the number and proportion of high-risk entities in the first-degree, second-degree, and third-degree relationships in the network, the number and proportion of high-risk entities in the first-level and second-level communities, and the degree and betweenness of the current enterprise node in the network, as shown in Table 8.
[0121] Table 6
[0122]
[0123] Table 7
[0124]
[0125] Table 8
[0126]
[0127] According to the embodiments of the present disclosure, the distance relationship between enterprise entities and other entities is judged by the shortest path algorithm. Considering that the entities closely related to high-risk nodes may also have a greater risk, the number and proportion of high-risk nodes included within the third-degree relationship of entity nodes can be statistically counted based on investment relationships and related party transaction relationships as part of the input features of the risk control model.
[0128] In the present disclosure, implementing the judgment of the distance relationship between enterprise entities and other entities through the shortest path algorithm in the Neo4j graph database requires relying on Cypher syntax for query. According to the relationship predicates that need to be evaluated in the query statement, different query plans may be generated for planning the shortest path in the Cypher query statement. If the relationship predicates can be evaluated when searching for paths internally in the Neo4j graph database, then the fast bidirectional breadth-first search algorithm will be used for searching in the Neo4j graph database. Therefore, when there are relationship predicates in the path, the correct shortest path query result can always be returned based on this fast search algorithm.
[0129] However, in the actual search process, for example, when finding the shortest path of enterprise relationships, if each node carries the label of the enterprise, or there is no corresponding search attribute in the path, this fast algorithm cannot be used to find the shortest path. In this case, Neo4j may have to use the slower exhaustive depth-first traversal algorithm to find the shortest path. This means that the query plan will execute a degradation scheme in the shortest path query with non-universal predicates. For example, in the query statement, a query statement containing an existential predicate condition is used to query the result, and the query condition statement requires that at least one node contains an attribute name with a specified value. Such a query may not be able to return results through the fast search algorithm. In this case, Neo4j will fall back to using exhaustive search to enumerate all paths and return the results. There may be an order-of-magnitude difference in the running times of these two algorithms, so it is very important to ensure the use of the fast method for queries that are concerned about the return time. When the query plan selects exhaustive search, the exhaustive search is still only executed when the fast algorithm fails to find any matching paths. In some cases, falling back to the traversal search may consume a large amount of resources and take a lot of time. For example, in the case where there is no shortest path between two nodes, it is necessary to set forbid_exhaustive_shortestpath to true to avoid response timeouts.
[0130] Use Cypher query statements in Neo4j to query the shortest path, and obtain the risk characteristic information of the current enterprise based on the shortest path. Query the shortest path between the current enterprise node and the enterprises and individuals in the low credit score list, and the path length does not exceed 3. You can query the out-degree and in-degree directions of the current enterprise separately. The shorter the obtained path length, the closer the current enterprise is to the individuals or enterprises in the low credit score list, indicating that the current enterprise node may have a higher business credit risk. Through the shortest path, the number of high-risk entities associated with the enterprise in the relationship of path lengths from 1 to 3 degrees can be obtained, as well as the proportion of high-risk entities associated with the enterprise in the relationship of path lengths from 1 to 3 degrees.
[0131] The metrics of social network analysis mainly include the degree of nodes in the network, the paths between nodes, the betweenness centrality of nodes or edges, and the community division in the network, etc. Analyzing the role of relevant metrics in the enterprise association information knowledge graph through social network analysis methods is of great significance for risk control in the graph network. In the enterprise association information knowledge graph, analyze the meaning of the node degree. An enterprise node with a larger degree in the network structure of the graph indicates that the current enterprise has more transactions or investment behaviors with other enterprises, representing that the current enterprise node has a greater influence in the network.
[0132] Optionally, the degree of a node refers to all n nodes in the network and the current node The number of entities connected by a path. In a directed graph, it includes two directions: out-degree and in-degree. If there is an edge connecting the current node to the node the degree value is equal to 1, and when there is no connection, the degree value is equal to 0. The calculation formula for the node degree is as follows:
[0133]
[0134] Optionally, for the paths between nodes, the shortest path is usually of concern. The shortest path is the shortest one among all connected paths between the current node and a certain node in the network. Based on the shortest path, the intimacy between two nodes can be obtained . The definition of intimacy is that the sum of the shortest paths from the current node to all other nodes in the network is called the intimacy of the current node . The concept of intimacy is used to describe the closeness or distance between the current node and other nodes in the network. The calculation formula for intimacy is as follows:
[0135]
[0136] Optionally, the node betweenness and edge betweenness, collectively called betweenness, are mainly used to describe the intermediacy. The concepts of node betweenness and edge betweenness are similarly defined. The prerequisite for calculating betweenness is to calculate the shortest paths between all nodes in the network. The node betweenness of the current node is defined as the number of paths among these shortest paths that contain the current node. Similarly, the edge betweenness of a certain edge is the number of paths among these shortest paths that contain the current edge. Betweenness or intermediacy is a relatively important indicator in social network analysis. Intermediacy reflects the connectivity of a node or a certain edge. In the enterprise association information graph, if the betweenness of the current node is large, it means that the current node is in the middle of multiple organizations in the flow graph of enterprise transactions or investments, which also indicates to a certain extent the importance of the current node in the network. Community discovery in the network: The process of community discovery is similar to the clustering process. Eventually, the network model is divided into several communities. Nodes in the same community are closely related, and nodes in different communities are distantly related. Common community discovery algorithms include the GN (Girvan-Newman) algorithm, Louvain algorithm, etc. According to whether there are duplicate node elements between communities, the algorithms can be divided into two categories: overlapping communities and non-overlapping communities. In the enterprise association information graph, through community discovery, the network nodes are divided into communities, and the enterprise and person entities closely related to the current node are obtained. In the risk control process, if a certain node and multiple network black production nodes are in the same community, it can be determined that the node may be in a black production group, and its risk coefficient is higher.
[0137] Through the embodiments of the present disclosure, not only can the overall characteristics be described, but all characteristic variables of the enterprise entity can also be divided into three categories, namely, the enterprise basic attribute category, the enterprise historical risk category, and the enterprise external relationship category. It is best to extract the features so that fewer characteristic variables can be selected to achieve a better model effect, so that the selected characteristic variables can reflect all aspects of the data and achieve faster operating efficiency.
[0138] As an optional embodiment, the method also includes: obtaining a risk prediction model; inputting the target feature data of the target credit subject into the risk prediction model to obtain a first risk prediction result of the target credit subject; and inputting the target feature data of the associated credit subject into the risk prediction model to obtain a second risk prediction result of the associated credit subject.
[0139] As an optional embodiment, the method further includes: acquiring training sample data, wherein the training sample data includes feature data of credit subjects with normal credit and feature data of credit subjects with low credit; and training the training sample data to obtain a risk prediction model.
[0140] Conduct risk assessment on enterprise nodes to determine whether the current enterprise entity has high risk information. In specific implementation, a suitable evaluation model can be constructed for judgment. Samples for training the evaluation model can come from enterprise information in the knowledge graph and its association with other entities. The required negative samples are low-credit enterprise nodes in the entity, mainly from the public data of the enterprise information disclosure system. These entities with high risks are saved by establishing a low-credit entity list. The low-credit entity list can include individuals with poor credit ratings or high-risk enterprises. Among them, individuals with poor credit ratings can mainly include persons subject to execution for dishonesty, persons subject to travel or consumption restrictions, and some online black industries involved in the constructed knowledge graph. High-risk enterprises can mainly include enterprises listed in the list of seriously illegal and dishonest enterprises, enterprises with administrative penalties or illegal acts, and enterprises listed in the list of abnormal operations.
[0141] As an optional embodiment, the method further includes: updating the risk prediction model based on the credit risk prediction result of the target credit subject.
[0142] The risk prediction model can be updated through the credit risk prediction results of the target credit entity, and the real-time update of the risk prediction model can be realized, which is conducive to improving the accuracy of credit risk prediction.
[0143] As an alternative embodiment, determining the credit risk prediction result of the target credit entity based on the first risk prediction result of the target credit entity and the second risk prediction result of the associated credit entity includes: when the first risk prediction result indicates that the credit of the target credit entity is abnormal, determining that the credit risk prediction result of the target credit entity is high risk; or when the first risk prediction result indicates that the credit of the target credit entity is normal and the second risk prediction result indicates that there is an associated credit entity with abnormal credit among the associated credit entities, determining that the credit risk prediction result of the target credit entity is high risk; or when the first risk prediction result indicates that the credit of the target credit entity is normal and the second risk prediction result indicates that there is no associated credit entity with abnormal credit among the associated credit entities, determining that the credit risk prediction result of the target credit entity is low risk.
[0144] Through the embodiments of the present disclosure, isolated data nodes are integrated into a unified knowledge base, the value of enterprise risk data is fully mined, the enterprise isolated points are broken, the interconnection and intercommunication of customer enterprise information are realized, the enterprise big data of the bank is efficiently utilized, the explicit and implicit association relationships between individuals and legal persons are deeply mined, groups composed of entities with certain common characteristics are identified, the process and probability of transmission of a certain event between associated entities are calculated, etc., so as to provide a visual view of the enterprise industrial and commercial litigation risks for the customer manager. It helps the bank to timely predict potentially risky associated enterprises in the pre-loan stage, make early warnings and pre-judgments, and helps the bank to timely discover potential risks in the post-loan stage, initiate the collection process in advance, and effectively reduce the losses of non-performing loans of the bank.
[0145] Figure 7 Schematically shows a flowchart of a method for predicting credit risk according to another embodiment of the present disclosure. As Figure 7 shown, the prediction method may include operation S710 to operation S750.
[0146] In operation S710, construct a domain ontology in the field of enterprise association information, and complete the design of classes and attributes under the enterprise ontology concept through ontology language design.
[0147] In operation S720, perform knowledge extraction on structured data. Extract structured data through the graph database middleware APOC, import entities into the graph database and complete the mapping of attributes and the generation of relationships between entities.
[0148] In operation S730, design a method for entity recognition and relationship extraction for unstructured text data for knowledge extraction. Use the information extraction tool DeepDive to perform weakly supervised extraction of relationships between entities, and import the results with high confidence into the knowledge base.
[0149] In operation S740, the knowledge graph of enterprise association information is supplemented and improved through knowledge base reasoning. Hierarchical reasoning of categories is performed for upper and lower positions and category completion through multi-level relationships of categories. The inconsistency of the definition of enterprise entity categories is detected through ontology reasoning. The implicit relationships between entities are inferred and completed through custom rule reasoning.
[0150] In operation S750, based on the knowledge graph of enterprise association information, the close relationship between enterprise nodes and entities with low credit is judged through the shortest path and social network analysis method, and the characteristics of enterprise external relationships are extracted. The characteristic data of three dimensions, namely enterprise basic attributes, enterprise historical risks, and enterprise external relationships, are used as the input of the risk control model. The risk control model discriminates the enterprise risks by learning the characteristics of normal enterprises and enterprises with low credit.
[0151] Through the embodiments of the present disclosure, the enterprise risk prediction method and system based on the risk graph perform knowledge mining and reconstruction on enterprise structured, unstructured and other heterogeneous data, construct a risk graph in the field of enterprise association information, analyze the qualifications and credit of the target enterprise, and whether there is a close association between the target enterprise and enterprises with low credit. Establish an enterprise risk model with the enterprise as the core, provide the ability to analyze problems from multiple relationship perspectives, deeply explore the potential relationships between individuals and the value behind the data, improve the value density of risk information, deeply explore the explicit and implicit association relationships between individuals and legal persons, and obtain valuable information such as hidden association relationships between enterprises, so as to help banks evaluate risks and optimize decisions.
[0152] Figure 8 The block diagram of the prediction device for credit risk according to an embodiment of the present disclosure is schematically shown. As Figure 8 shown, the device 800 may include a first acquisition module 810, a generation module 820, a first determination module 830, and a second determination module 840.
[0153] The first acquisition module 810 is configured to acquire target feature data of a credit subject. The credit subject includes a target credit subject and a non-target credit subject, and the target feature data is used to characterize the credit risk of the credit subject. Optionally, the first acquisition module 810 may be configured to execute Figure 2 the operation S210 described herein, which will not be elaborated herein.
[0154] The generation module 820 is configured to perform knowledge extraction on the target feature data to generate a knowledge graph of the target credit subject, and the knowledge graph is used to characterize the entities, attributes of each credit subject in the credit subject, and the relationships between each credit subject. Optionally, the generation module 820 may be configured to execute Figure 2 the operation S220 described herein, which will not be elaborated herein.
[0155] The first determination module 830 is configured to determine, based on the knowledge graph, an associated credit entity having an association relationship with the target credit entity from non-target credit entities. Optionally, the first determination module 830 may be configured to perform, for example, Figure 2 the operation S230 described, which will not be elaborated herein.
[0156] The second determination module 840 is configured to determine a credit risk prediction result of the target credit entity based on a first risk prediction result of the target credit entity and a second risk prediction result of the associated credit entity. Optionally, the second determination module 840 may be configured to perform, for example, Figure 2 the operation S240 described, which will not be elaborated herein.
[0157] As an optional embodiment, the generation module 820 may include: a first determination sub-module configured to determine a data structure of the target feature data, where different data structures correspond to different knowledge extraction logics; a selection sub-module configured to select a corresponding knowledge extraction logic according to the data structure; and a generation sub-module configured to perform knowledge extraction on the target feature data based on the corresponding knowledge extraction logic to generate a knowledge graph of the target credit entity.
[0158] As an optional embodiment, before generating the knowledge graph of the target credit entity, the apparatus 800 may further include: a first construction module 850 configured to construct a domain ontology for describing the association information of the credit entity; a definition module 860 configured to define, through the Web Ontology Language, the category to which the domain ontology belongs and the identifiers belonging to the category, where the identifiers include entity identifiers, attribute identifiers, and relationship identifiers, and the entity identifiers, attribute identifiers, and relationship identifiers are stored in the graph database; and a second construction module 870 configured to construct, based on the entity identifiers, attribute identifiers, and relationship identifiers, an entity description framework for describing the association information of the domain ontology, where the entity description framework is used to generate the knowledge graph.
[0159] As an optional embodiment, the data structure includes a structured data structure, and the generation sub-module may include: a first extraction unit configured to call a knowledge extraction middleware of the graph database to perform knowledge extraction on the target feature data to obtain target field information, where the target field information includes first entity information, first attribute information, and first relationship information; and a first generation unit configured to use the target field information to fill the entity description framework to generate a knowledge graph of the target credit entity, where the knowledge graph includes the first entity information corresponding to the entity identifier, the first attribute information corresponding to the attribute identifier, and the first relationship information corresponding to the relationship identifier.
[0160] As an alternative embodiment, the data structure includes an unstructured data structure, and the generation sub-module may include: an annotation unit for annotating target feature data to obtain annotation sequence information, where the annotation sequence information includes second entity information and second attribute information; a second extraction unit for extracting the target feature data based on an extraction method of weak supervision learning to obtain second relationship information, where the second relationship information is used to represent the relationship between any two credit entities in the credit entity; and a second generation unit for generating a knowledge graph of the target credit entity based on the second entity information, the second attribute information, and the second relationship information, where the knowledge graph includes second entity information corresponding to the entity identifier, second attribute information corresponding to the attribute identifier, and second relationship information corresponding to the relationship identifier.
[0161] As an alternative embodiment, the second generation unit may include: a determination subunit for determining a confidence value corresponding to the second relationship information; an acquisition subunit for acquiring a confidence threshold of the relationship information; an extraction subunit for extracting, based on the confidence threshold, third relationship information whose confidence value meets the confidence threshold from the credit entity; and a generation subunit for filling an entity description framework with the second entity information, the second attribute information, and the third relationship information to generate a knowledge graph of the target credit entity, where the knowledge graph includes third relationship information corresponding to the relationship identifier.
[0162] As an alternative embodiment, the apparatus 800 may further include: an inference module 880 for inferring the knowledge graph through inference rules of the Web Ontology Language to improve the knowledge graph; and / or a detection module 890 for performing a consistency check on the categories to which the domain ontology belongs to clean up abnormal categories.
[0163] As an alternative embodiment, the first determination module 830 may include: a first acquisition sub-module for obtaining the shortest path including the target credit entity according to a preset path direction, where the preset path direction includes an out-degree direction and an in-degree direction; a second acquisition sub-module for obtaining a community division result of the knowledge graph through a preset community discovery algorithm, where there is an association relationship between credit entities in the same community; and a second determination sub-module for determining, based on the shortest path including the target credit entity and / or the community division result of the knowledge graph, an associated credit entity having an association relationship with the target credit entity from non-target credit entities.
[0164] As an alternative embodiment, the device 800 may further include: a first acquisition module 8100, configured to acquire a risk prediction model; a second acquisition module 8110, configured to first input the target feature data of the target credit subject into the risk prediction model to obtain a first risk prediction result of the target credit subject; and a third acquisition module 8120, configured to input the target feature data of the associated credit subject into the risk prediction model to obtain a second risk prediction result of the associated credit subject.
[0165] As an alternative embodiment, the device 800 may further include: a second acquisition module 8130, configured to acquire training sample data, where the training sample data includes the feature data of credit subjects with normal credit and the feature data of credit subjects with low credit; and a training module 8140, configured to train the training sample data to obtain a risk prediction model.
[0166] As an alternative embodiment, the device 800 may further include: an update module 8150, configured to update the risk prediction model based on the credit risk prediction result of the target credit subject.
[0167] As an alternative embodiment, the second determination module 840 may include: a third determination sub-module, configured to determine that the credit risk prediction result of the target credit subject is a high risk when the first risk prediction result indicates that the credit of the target credit subject is abnormal; or a fourth determination sub-module, configured to determine that the credit risk prediction result of the target credit subject is a high risk when the first risk prediction result indicates that the credit of the target credit subject is normal and the second risk prediction result indicates that there is an associated credit subject with abnormal credit among the associated credit subjects; or a fifth determination sub-module, configured to determine that the credit risk prediction result of the target credit subject is a low risk when the first risk prediction result indicates that the credit of the target credit subject is normal and the second risk prediction result indicates that there is no associated credit subject with abnormal credit among the associated credit subjects.
[0168] It should be noted that the implementation manners, the technical problems solved, the functions achieved, and the technical effects achieved by the modules in some embodiments of the credit risk prediction device are respectively the same as or similar to those of the corresponding steps in some embodiments of the credit risk prediction method, and will not be elaborated herein.
[0169] Any of a plurality of modules, sub-modules, units, and sub-units according to embodiments of the present disclosure, or at least part of the functions of any of them, can be implemented in one module. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure can be split into multiple modules for implementation. Any one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure can be at least partially implemented as a hardware circuit, such as a field-programmable gate array (FNGA), a programmable logic array (NLA), a system-on-chip, a system-on-substrate, a system-on-package, an application-specific integrated circuit (ASIC), or can be implemented by hardware or firmware in any other reasonable way of integrating or packaging circuits, or implemented in any one of the three implementation manners of software, hardware, and firmware, or in an appropriate combination of any several of them. Alternatively, one or more of the modules, sub-modules, units, and sub-units according to embodiments of the present disclosure can be at least partially implemented as a computer program module, and when the computer program module runs, it can execute the corresponding functions.
[0170] For example, the first acquisition module, the generation module, the first determination module, the second determination module, the first determination sub-module, the selection sub-module, the generation sub-module, the first construction module, the definition module, the second construction module, the first extraction unit, the first generation unit, the annotation unit, the second extraction unit, the second generation unit, the determination sub-unit, the acquisition sub-unit, the extraction sub-unit, the generation sub-unit, the inference module, the detection module, the first acquisition sub-module, the second acquisition sub-module, the second determination sub-module, the first acquisition module, the second acquisition module, the third acquisition module, the second acquisition module, the training module, the update module, the third determination sub-module, the fourth determination sub-module, and the fifth determination sub-module can be combined and implemented in one module, or any one of them can be split into multiple modules. Alternatively, at least part of the functions of one or more of these modules can be combined with at least part of the functions of other modules and implemented in one module.
[0171] Figure 9 A schematic diagram of a computer-readable storage medium product suitable for implementing the credit risk prediction method described above according to embodiments of the present disclosure is schematically shown.
[0172] In some possible implementation manners, various aspects of the present invention can also be implemented in the form of a program product, which includes program code. When the program product runs on a device, the program code is used to cause the device to perform the foregoing operations (or steps) in the credit risk prediction method according to various exemplary embodiments of the present invention described in the "Exemplary Method" section of this specification. For example, an electronic device can perform operations S210 to S240 as shown in Figure 2 and operations as shown in Figure 4Operations S410 to S4120 shown in, such as Figure 5 Operations S510 to S5100 shown in, and such as Figure 6 Operations S610 to S640 shown in.
[0173] The program product can adopt any combination of one or more readable media. The readable media can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (ENROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0174] Such as Figure 9 As shown, a program product 900 for predicting credit risk according to an embodiment of the present invention can adopt a portable compact disk read-only memory (CD-ROM), include program code, and can run on a device such as a personal computer. The program product of the present invention is not limited thereto. In this document, the readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0175] The readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the readable program code. Such a propagated data signal can take various forms, including - but not limited to - an electromagnetic signal, an optical signal, or any suitable combination of the above. The readable signal medium can also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the readable medium can be transmitted by any appropriate medium, including - but not limited to - wireless, wired, optical fiber, RF, etc., or any suitable combination of the above.
[0176] The program code for performing the operations of the present invention can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as "C", languages or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (e.g., by connecting through the Internet using an Internet service provider).
[0177] Figure 10 FIG. schematically shows a block diagram of an electronic device suitable for implementing the credit risk prediction method described above according to an embodiment of the present disclosure.
[0178] Figure 10 The electronic device shown is merely an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0179] As Figure 10 shown, the electronic device 1000 according to an embodiment of the present disclosure includes a processor 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage section 1008 into the random access memory (RAM) 1003. The processor 1001 can include, for example, a general-purpose microprocessor (e.g., CNU), an instruction set processor, and / or a related chipset, and / or a dedicated microprocessor (e.g., an application-specific integrated circuit (ASIC)), etc. The processor 1001 can also include on-board memory for caching purposes. The processor 1001 can include a single processing unit or multiple processing units for performing different actions of the method flow according to an embodiment of the present disclosure.
[0180] In the RAM 1003, various programs and data required for the operation of the electronic device 1000 are stored. The processor 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. The processor 1001 performs various operations of the method flow according to an embodiment of the present disclosure by executing the programs in the ROM 1002 and / or the RAM 1003. It should be noted that the program can also be stored in one or more memories other than the ROM 1002 and the RAM 1003. The processor 1001 can also perform according to an embodiment of the present disclosure by executing the programs stored in the one or more memories. Figure 2 、 Figure 4 、 Figure 5and Figure 6 the operations shown
[0181] According to an embodiment of the present disclosure, the electronic device 1000 may further include an input / output (I / O) interface 1005, and the input / output (I / O) interface 1005 is also connected to the bus 1004. The electronic device 1000 may further include one or more of the following components connected to the I / O interface 1005: an input portion 1006 including a keyboard, a mouse, etc.; an output portion 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage portion 1008 including a hard disk, etc.; and a communication portion 1009 including a network interface card such as an LAA card, a modem, etc. The communication portion 1009 performs communication processing via a network such as the Internet. The drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed so that a computer program read therefrom is installed into the storage portion 1008 as needed.
[0182] According to an embodiment of the present disclosure, the method flow according to the embodiment of the present disclosure may be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable storage medium, and the computer program includes program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication portion 1009, and / or installed from the removable medium 1011. When the computer program is executed by the processor 1001, the above functions defined in the system of the embodiment of the present disclosure are executed. According to an embodiment of the present disclosure, the above-described system, device, apparatus, module, unit, etc. may be implemented by computer program modules.
[0183] The present disclosure also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiment; or may exist separately without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the above one or more programs are executed, the operations of the credit risk prediction method according to the embodiment of the present disclosure are implemented.
[0184] According to embodiments of the present disclosure, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, it may include but is not limited to: portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. For example, according to embodiments of the present disclosure, the computer-readable storage medium may include the ROM 1002 and / or the RAM 1003 described above and / or one or more memories other than the ROM 1002 and the RAM 1003.
[0185] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. In some alternative implementations, the functions marked in the block may occur in a different order from that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as the combination of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0186] Those skilled in the art can understand that the features recited in various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly recited in the present disclosure. In particular, without departing from the spirit and teachings of the present disclosure, the features recited in various embodiments and / or claims of the present disclosure can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present disclosure.
[0187] The above describes the embodiments of the present disclosure. However, these embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure. Although the embodiments are described separately above, this does not mean that the measures in each embodiment cannot be used advantageously in combination. The scope of the present disclosure is defined by the appended claims and their equivalents. Without departing from the scope of the present disclosure, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present disclosure.
Claims
1. A method for predicting credit risk, comprising: Obtaining target feature data of a credit entity, where the credit entity includes a target credit entity and a non-target credit entity, and the target feature data is used to characterize the credit risk of the credit entity; Performing knowledge extraction on the target feature data to generate a knowledge graph of the target credit entity, where the knowledge graph is used to characterize the entities, attributes of each credit entity in the credit entity, and the relationships between the credit entities, and the target feature data includes enterprise operation attribute feature data, historical risk feature data, and external relationship feature data; Based on the knowledge graph, determining associated credit entities having an association relationship with the target credit entity from the non-target credit entities includes: Obtaining the shortest path including the target credit entity according to a preset path direction, where the preset path direction includes an out-degree direction and an in-degree direction, and the shortest path includes low-credit entities associated with the target credit entity; Based on degree metrics, path metrics, betweenness centrality metrics, edge betweenness centrality metrics, and network community partition metrics of nodes in the knowledge graph, obtaining a community partition result of the knowledge graph through a preset community discovery algorithm, where credit entities in the same community have an association relationship; Based on the shortest path including the target credit entity and / or the community partition result of the knowledge graph, determining associated credit entities having an association relationship with the target credit entity from the non-target credit entities, where the association relationship characterizes a credit risk relationship; Based on the first risk prediction result of the target credit entity and the second risk prediction result of the associated credit entities, determining a credit risk prediction result of the target credit entity.
2. The method according to claim 1, wherein The performing knowledge extraction on the target feature data to generate a knowledge graph of the target credit entity includes: Determining the data structure of the target feature data, where different data structures correspond to different knowledge extraction logics; According to the data structure, selecting the corresponding knowledge extraction logic; Based on the corresponding knowledge extraction logic, performing knowledge extraction on the target feature data to generate a knowledge graph of the target credit entity.
3. The method according to claim 2, wherein Before generating the knowledge graph of the target credit entity, the method further includes: Constructing a domain ontology for describing the association information of the credit entity; Defining the category to which the domain ontology belongs and the identifiers belonging to the category through the Web Ontology Language, where the identifiers include entity identifiers, attribute identifiers, and relationship identifiers, and the entity identifiers, attribute identifiers, and relationship identifiers are stored in a graph database; Based on the entity identifiers, attribute identifiers, and relationship identifiers, constructing an entity description framework for describing the association information of the domain ontology, where the entity description framework is used to generate a knowledge graph.
4. The method according to claim 3, wherein The data structure includes a structured data structure, and the performing knowledge extraction on the target feature data based on the corresponding knowledge extraction logic to generate a knowledge graph of the target credit entity includes: Invoke the knowledge extraction middleware of the graph database to perform knowledge extraction on the target feature data to obtain target field information, where the target field information includes first entity information, first attribute information, and first relationship information; Use the target field information to fill the entity description framework to generate the knowledge graph of the target credit subject, where the knowledge graph includes the first entity information corresponding to the entity identifier, the first attribute information corresponding to the attribute identifier, and the first relationship information corresponding to the relationship identifier.
5. The method according to claim 3, wherein, The data structure includes an unstructured data structure. The generating the knowledge graph of the target credit subject by performing knowledge extraction on the target feature data based on the corresponding knowledge extraction logic includes: Annotate the target feature data to obtain annotation sequence information, where the annotation sequence information includes second entity information and second attribute information; Extract the target feature data based on the extraction method of weak supervision learning to obtain second relationship information, where the second relationship information is used to characterize the relationship between any two credit subjects in the credit subject; Generate the knowledge graph of the target credit subject based on the second entity information, the second attribute information, and the second relationship information, where the knowledge graph includes the second entity information corresponding to the entity identifier, the second attribute information corresponding to the attribute identifier, and the second relationship information corresponding to the relationship identifier.
6. The method according to claim 5, wherein, The generating the knowledge graph of the target credit subject based on the second entity information, the second attribute information, and the second relationship information includes: Determine the confidence value corresponding to the second relationship information; Obtain the confidence threshold of the relationship information; Based on the confidence threshold, extract the third relationship information from the credit subject whose confidence value meets the confidence threshold; Use the second entity information, the second attribute information, and the third relationship information to fill the entity description framework to generate the knowledge graph of the target credit subject, where the knowledge graph includes the third relationship information corresponding to the relationship identifier.
7. The method according to claim 3, wherein The method further includes at least one of the following: Reason about the knowledge graph through the inference rules of the Web Ontology Language to improve the knowledge graph; Perform consistency detection on the categories to which the domain ontology belongs to clean up abnormal categories.
8. The method according to claim 1, wherein The method further includes: Obtain a risk prediction model; Input the target feature data of the target credit subject into the risk prediction model to obtain the first risk prediction result of the target credit subject; Input the target feature data of the associated credit subject into the risk prediction model to obtain the second risk prediction result of the associated credit subject.
9. The method according to claim 8, wherein The method further includes: Obtain training sample data, where the training sample data includes the feature data of credit subjects with normal credit and the feature data of credit subjects with low credit; Train the training sample data to obtain the risk prediction model.
10. The method according to claim 8, wherein, The method further includes: Update the risk prediction model based on the credit risk prediction result of the target credit subject.
11. The method according to claim 1, wherein, Determining the credit risk prediction result of the target credit entity based on the first risk prediction result of the target credit entity and the second risk prediction result of the associated credit entity includes: When the first risk prediction result indicates that the credit of the target credit entity is abnormal, determining that the credit risk prediction result of the target credit entity is high risk; or When the first risk prediction result indicates that the credit of the target credit entity is normal and the second risk prediction result indicates that there is an associated credit entity with abnormal credit among the associated credit entities, determining that the credit risk prediction result of the target credit entity is high risk; or When the first risk prediction result indicates that the credit of the target credit entity is normal and the second risk prediction result indicates that there is no associated credit entity with abnormal credit among the associated credit entities, determining that the credit risk prediction result of the target credit entity is low risk.
12. A prediction device for credit risk, comprising: An acquisition module, configured to acquire target feature data of a credit entity, where the credit entity includes a target credit entity and non-target credit entities, and the target feature data is used to characterize the credit risk of the credit entity; A generation module, configured to perform knowledge extraction on the target feature data to generate a knowledge graph of the target credit entity, where the knowledge graph is used to characterize the entities, attributes of each credit entity in the credit entity, and the relationships between the credit entities, and the target feature data includes enterprise operation attribute feature data, historical risk feature data, and external relationship feature data; A first determination module, configured to determine, based on the knowledge graph, an associated credit entity having an association relationship with the target credit entity from the non-target credit entities, where the first determination module includes a first acquisition sub-module, a second acquisition sub-module, and a second determination sub-module; The first acquisition sub-module is configured to obtain the shortest path including the target credit entity according to a preset path direction, where the preset path direction includes an out-degree direction and an in-degree direction, and the shortest path includes low-credit entities associated with the target credit entity; The second acquisition sub-module is configured to obtain a community division result of the knowledge graph through a preset community discovery algorithm based on degree indicators, path indicators, betweenness centrality indicators, edge betweenness centrality indicators, and network community division indicators of nodes in the knowledge graph, where credit entities in the same community have an association relationship; The second determination sub-module is configured to determine, based on the shortest path including the target credit entity and / or the community division result of the knowledge graph, an associated credit entity having an association relationship with the target credit entity from the non-target credit entities, and the association relationship represents a credit risk relationship; A second determination module, configured to determine the credit risk prediction result of the target credit entity based on the first risk prediction result of the target credit entity and the second risk prediction result of the associated credit entity.
13. An electronic device, comprising: One or more processors; And A memory, configured to store one or more programs, Wherein, when the one or more programs are executed by the one or more processors, the one or more processors are caused to execute the method according to any one of claims 1 to 11.
14. A computer-readable storage medium storing computer-executable instructions that, when executed, cause a processor to execute the method according to any one of claims 1 to 11.
15. A computer program product comprising a computer program that, when executed by a processor, executes the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Large-scale enterprise credit risk prediction method and system, storage medium and electronic equipment
CN110930249A