Risk Identification Method, Device, Electronic Device and Storage Medium

By extracting, clustering and weighting of the associated data of the enterprise, and combining with the risk knowledge graph, the problem of insufficient accuracy in enterprise risk identification is solved, and more accurate risk identification results are achieved.

CN115187066BActive Publication Date: 2025-05-27PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210813786.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-11
Publication Date
2025-05-27
Estimated Expiration
2042-07-11

AI Technical Summary

Technical Problem

In the prior art, the accuracy of enterprise risk identification results is insufficient, especially when facing big data and multi-dimensional data, it is difficult to accurately identify enterprise risk factors.

Method used

Risk factor extraction is performed on the associated data of the objects to be identified, clustering is performed using the similarity between candidate risk factors, selecting selection weights, filtering out the target risk factor, and using the preset risk knowledge graph for risk feature extraction, and finally determining the risk identification result.

Benefits of technology

The dimension of risk factors is reduced, the accuracy of risk identification results is improved, and the accuracy and reliability of risk identification are ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115187066B_ABST
    Figure CN115187066B_ABST
Patent Text Reader

Abstract

The present application relates to the field of artificial intelligence technology, and specifically discloses a risk identification method, apparatus, computer device, and computer-readable storage medium. The risk identification method of the present application extracts risk factors from the associated data of the object to be identified to obtain a plurality of candidate risk factors, performs clustering processing on the plurality of candidate risk factors according to the similarity between the plurality of candidate risk factors to obtain a clustering result, and then calculates the selection weights of the candidate risk factors according to the clustering result, and uses the candidate risk factors whose selection weights meet the preset conditions as target risk factors, thereby reducing the dimension of the risk factors. Then, risk feature extraction is performed on the target risk factors with reduced dimension according to a preset risk knowledge graph to obtain object risk features, and the risk identification result of the object to be identified is determined according to the object risk features, so that the obtained risk identification result is more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and more particularly, to a risk identification method, apparatus, computer device, and computer-readable storage medium. Background Art

[0002] Enterprises need to disclose enterprise information regularly. Enterprise information disclosure refers to that enterprises should actively disclose operation and financial information at the request of management departments for the reference of stakeholders. Therefore, in order to ensure the authenticity of information disclosure and avoid enterprises deliberately providing false information to mislead users, it is necessary to conduct risk analysis on enterprises. However, due to the characteristics of the relevant data of enterprises, such as many data dimensions and large data volume, the risk analysis of enterprises is not accurate enough.

[0003] Therefore, how to improve the accuracy of the identification results obtained by the existing risk identification for enterprises is a technical problem to be solved. Summary of the Invention

[0004] To solve the above technical problems, embodiments of the present application provide a risk identification method, apparatus, computer device, and computer-readable storage medium to improve the accuracy of risk identification.

[0005] In a first aspect, the present application provides a risk identification method, including: extracting risk factors from the associated data of the object to be identified to obtain a plurality of candidate risk factors; clustering the plurality of candidate risk factors according to the similarity between the plurality of candidate risk factors to obtain a clustering result; calculating a selection weight of the candidate risk factors according to the clustering result, and using the candidate risk factors whose selection weights meet the preset conditions as target risk factors; where the selection weight is used to indicate the contribution degree of the candidate risk factors to determining the risk identification result of the object to be identified; extracting risk characteristics of the target risk factors according to a preset risk knowledge graph to obtain object risk characteristics; and determining the risk identification result of the object to be identified according to the object risk characteristics.

[0006] According to a preferred embodiment of the present invention, clustering the plurality of candidate risk factors according to the similarity between the plurality of candidate risk factors to obtain a clustering result includes: splitting the plurality of candidate risk factors according to the data generation time corresponding to each of the plurality of candidate risk factors to obtain a plurality of candidate risk factor sequences; and clustering the candidate risk factor sequences according to the similarity between the respective candidate risk factor sequences to obtain a clustering result.

[0007] According to a preferred embodiment of the present invention, before clustering the candidate risk factor sequences according to the similarity between each pair of candidate risk factor sequences to obtain a clustering result, it further includes: constructing a distance matrix based on the lengths of each candidate risk factor sequence, where each position in the distance matrix represents the distance between each pair of candidate risk factor sequences; calculating a first cumulative distance from the starting position to the target position in the distance matrix, and calculating a second cumulative distance from the ending position to the target position in the distance matrix; calculating the minimum distance between each pair of candidate risk factor sequences based on the first cumulative distance and the second cumulative distance, and determining the similarity between each pair of candidate risk factor sequences according to the minimum distance.

[0008] According to a preferred embodiment of the present invention, the clustering result includes multiple clustering sets; calculating the selection weight of the candidate risk factors according to the clustering result, and taking the candidate risk factors whose selection weights meet the preset conditions as target risk factors, including: determining the clustering center vector corresponding to each clustering set; calculating the probability that the candidate risk factors belong to each clustering set according to the clustering center vector, so as to generate a weak label matrix of the candidate risk factors according to the probability; calculating the selection weight of the candidate risk factors according to the feature selection matrix and the weak label matrix of the candidate risk factors, and taking the candidate risk factors whose selection weights meet the preset conditions as target risk factors; wherein, the feature selection matrix is obtained through deep learning training based on the sample risk factors and sample risk recognition results in the training samples.

[0009] According to a preferred embodiment of the present invention, extracting risk characteristics of the object from the target risk factors according to a preset risk knowledge graph, including: determining the risk entity corresponding to the target risk factor; extracting a sub-graph matching the risk entity from the risk knowledge graph; encoding each node in the sub-graph to obtain node features; fusing the node features of each node to obtain the object risk characteristics.

[0010] According to a preferred embodiment of the present invention, determining the risk recognition result of the object to be recognized according to the object risk characteristics, including: obtaining the risk data of the associated object having an association relationship with the object to be recognized; performing risk conduction calculation on the risk data according to the category of the association relationship to obtain the risk conduction characteristics of the associated object relative to the object to be recognized; determining the risk recognition result of the object to be recognized according to the object risk characteristics and the risk conduction characteristics.

[0011] According to a preferred embodiment of the present invention, risk conduction calculation is performed on risk data according to the category of the association relationship to obtain the risk conduction characteristics of the associated object relative to the object to be identified, including: calculating the risk association degree between the risk data and the object to be identified; and determining the weight coefficient corresponding to the associated object according to the category of the association relationship; performing weighted calculation on the risk association degree according to the weight coefficient to obtain the risk conduction characteristics of the associated object relative to the object to be identified.

[0012] In a second aspect, the present application provides a risk identification device, including: a risk factor extraction module configured to extract risk factors from the association data of the object to be identified to obtain a plurality of candidate risk factors; a clustering module configured to perform clustering processing on the plurality of candidate risk factors according to the similarity between the plurality of candidate risk factors to obtain a clustering result; a target risk factor selection module configured to calculate the selection weight of the candidate risk factors according to the clustering result, and use the candidate risk factors whose selection weight meets the preset conditions as target risk factors; wherein the selection weight is used to indicate the contribution degree of the candidate risk factors to determining the risk identification result of the object to be identified; a risk feature extraction module configured to extract risk features from the target risk factors according to a preset risk knowledge graph to obtain object risk features; a risk identification module configured to determine the risk identification result of the object to be identified according to the object risk features.

[0013] In a third aspect, the present application provides a computer device, which includes a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the steps of the above risk identification method when executing the computer program.

[0014] In a fourth aspect, the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the steps of the above risk identification method.

[0015] The risk identification method, device, computer device and computer-readable storage medium disclosed in the embodiments of the present application extract risk factors from the association data of the object to be identified to obtain a plurality of candidate risk factors, perform clustering processing on the plurality of candidate risk factors according to the similarity between the plurality of candidate risk factors to obtain a clustering result, and then calculate the selection weight of the candidate risk factors according to the clustering result, and use the candidate risk factors whose selection weight meets the preset conditions as target risk factors, thereby reducing the dimension of the risk factors. Then, risk features are extracted from the target risk factors with reduced dimension according to a preset risk knowledge graph to obtain object risk features, and the risk identification result of the object to be identified is determined according to the object risk features, so that the obtained risk identification result is more accurate. Description of the Drawings

[0016] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:

[0017] Figure 1 is a schematic diagram of the application environment of the risk identification method provided by an exemplary embodiment of this application;

[0018] Figure 2 is a flowchart of the risk identification method provided by an exemplary embodiment of this application;

[0019] Figure 3 is a schematic diagram of generating a risk knowledge graph provided by an exemplary embodiment of this application;

[0020] Figure 4 is a flowchart of the risk identification method provided by another exemplary embodiment of this application;

[0021] Figure 5 is a schematic diagram of clustering processing on a candidate risk factor sequence provided by an exemplary embodiment of this application;

[0022] Figure 6 is a flowchart of a risk identification provided by another exemplary embodiment of this application;

[0023] Figure 7 is a flowchart of a risk identification provided by another exemplary embodiment of this application;

[0024] Figure 8 is a schematic diagram of obtaining associated data of an object to be identified provided by an exemplary embodiment of this application;

[0025] Figure 9 is a schematic block diagram of a risk identification device provided by an exemplary embodiment of this application;

[0026] Figure 10 is a schematic block diagram of a computer device provided by an exemplary embodiment of this application. Detailed Embodiments

[0027] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. On the contrary, they are only examples of devices and methods consistent with some aspects of this application as detailed in the appended claims.

[0028] The block diagrams shown in the drawings are only functional entities and do not necessarily correspond to physically independent entities. That is, these functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.

[0029] The flowcharts shown in the drawings are only illustrative and do not necessarily include all the content and operations / steps, nor do they have to be executed in the described order. For example, some operations / steps can be decomposed, while some operations / steps can be combined or partially combined, so the actual execution order may change according to the actual situation.

[0030] It should also be noted that: "a plurality of" mentioned in this application means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0031] Figure 1 Fig. shows a schematic diagram of a system architecture of the operating environment of an exemplary embodiment of the present application. Refer to Figure 1 As shown, the system may include a terminal 110 and a server 120. The terminal 110 and the server 120 are communicatively connected via a network, and the network may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc.

[0032] The terminal 110 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a vehicle-mounted terminal, etc., but is not limited thereto. Those skilled in the art can know that the number of the above terminals can be more or less. For example, the above terminal may be only one, or the above terminals may be dozens or hundreds, or more. In this case, the implementation environment of the above image processing method further includes other terminals. The embodiments of the present application do not limit the number and device types of the terminals.

[0033] The server 120 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The server 120 is used to provide background services for the application programs running on the terminal 110.

[0034] Optionally, the above-mentioned wireless network or wired network uses standard communication technologies and / or protocols. The network is usually the Internet, but can also be any network, including but not limited to any combination of a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a private network or a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent the data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec), etc. can be used to encrypt all or some of the links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the above data communication technologies.

[0035] Optionally, the server 120 undertakes the main risk identification work, and the terminal 110 undertakes the secondary risk identification work; or, the server 120 undertakes the secondary risk identification work, and the terminal 110 undertakes the main risk identification work; or, the server 120 or the terminal 110 can separately undertake the risk identification work.

[0036] The following will describe in detail some embodiments of the present application with reference to the accompanying drawings. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0037] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of a risk identification method provided for an embodiment of the present application. This risk identification method can be applied to Figure 1 the implementation environment shown, and is specifically executed by the server 120 in this implementation environment. It should be understood that this method can also be applicable to other exemplary implementation environments and be specifically executed by devices in other implementation environments. This embodiment does not limit the implementation environment applicable to this method.

[0038] As shown in Figure 2As shown in the figure, in an exemplary embodiment, the method at least includes steps S210 to S250, which are introduced in detail as follows:

[0039] Step S210, extracting risk factors from the associated data of the object to be recognized to obtain multiple candidate risk factors.

[0040] It should be noted that the associated data of the object to be recognized refers to the data related to the object to be recognized. For example, when the object to be recognized is an enterprise, the associated data can be the registered capital of the enterprise, financial information, various news public opinions, legal judgment documents, industrial and commercial information, etc. A risk factor refers to a factor that causes risks to the object to be recognized. For example, when the object to be recognized is an enterprise, the risk factors can be the transaction records of the enterprise, overdue records, equity changes, etc.

[0041] Exemplarily, it can be to periodically obtain and record the associated data of the object to be recognized, so as to extract risk factors for the object to be recognized according to the associated data, and then detect the risks of the object to be recognized in real time; it can also be to obtain and record the associated data of the object to be recognized when a preset trigger event is detected, so as to extract risk factors for the object to be recognized according to the associated data. For example, when the object to be recognized is an enterprise, when it is detected that the enterprise discloses enterprise information, extract risk factors from the registered capital, financial information, various news public opinions, enterprise evaluations, legal judgment documents, industrial and commercial information, etc. of the enterprise within a preset time period.

[0042] In some implementation manners, in combination with the above description, the data types of the associated data of the object to be recognized include structured data and unstructured data. In different scenarios, the data types of the obtained associated data may be different. In order to process these associated data, this application proposes a variety of data processing rules for structuring the associated data to obtain corresponding structured data, which is convenient for subsequent analysis of the structured data. Therefore, after the server obtains the associated data of the object to be recognized, it can first determine the data type of the associated data, and then select the corresponding data processing rule, so as to structure the obtained associated data according to the selected data processing rule to obtain the corresponding structured data.

[0043] By structuring the associated data of the object to be recognized, unstructured data such as news public opinions and enterprise evaluations can be considered when analyzing the risks of the object to be recognized, enriching the data dimension of risk analysis and improving the accuracy of risk analysis.

[0044] Optionally, it can be that the server stores a preset risk factor set, which contains pre-extracted risk factors, and keywords in the associated data of the object to be recognized are matched through the risk factor set to obtain multiple candidate risk factors.

[0045] For example, when the object to be recognized is an enterprise, the server crawls the historical associated data of each enterprise from the web page. The historical associated data includes historical risk events and enterprise data associated with the historical risk events. Then, natural language processing (NLP) is used to process the historical associated data of each enterprise crawled, such as performing lexical analysis, sentiment analysis, semantic analysis, etc., to obtain multiple risk factors to be stored included in the enterprise data associated with the historical risk events. The risk factors to be stored can include different types, such as financial factors, event factors, etc.

[0046] Then, an association analysis is performed on the historical risk events and the enterprise data associated with the historical risk events to obtain the association strength of each risk factor to be stored, and the risk factors to be stored with an association strength greater than or equal to the preset association strength threshold are stored in the risk factor set.

[0047] Taking the risk factor to be stored as an event factor as an example, the calculation formula for the association strength of the risk factor to be stored can be as follows:

[0048]

[0049] where Score k,i (t) represents the score of the k-th type of event of the i-th enterprise in the t-th quarter, Pro t (Event k |D t ) represents the ratio of the occurrence frequency of Event k to the occurrence frequency of all events in D t days, Power(Event k ) represents the impact of event Event k , which depends on the type k. w(D t ) represents the weight of event Event k on the D t ′-th day after the occurrence of the associated risk.

[0050] For example, when determining whether the enterprise has risks, it can be that when the enterprise meets one of the following three situations, it is determined that the enterprise has risks, and the enterprise is marked as 1 at this time point, otherwise it is 0:

[0051] 1. The company has a bond default situation;

[0052] 2. The company has a credit rating downgrade compared to the previous reporting period;

[0053] 3. The company has a major financial risk event such as a performance loss or bankruptcy.

[0054] It can be understood that the stronger the correlation strength of the risk factor to be stored, the stronger the influence of the risk factor to be stored on the historical risk event. Therefore, the risk factors included in the risk factor set are more accurate, and the candidate risk factors obtained by extracting the risk factors from the associated data of the object to be identified according to the risk factor set are also more accurate.

[0055] Step S220: Cluster the multiple candidate risk factors according to the similarity between the multiple candidate risk factors to obtain a clustering result.

[0056] In practical applications, the dimension of the candidate risk factors included in the associated data of the object to be identified is relatively high. For example, when the object to be identified is an enterprise, the candidate risk factors included in the obtained associated data may be market risk, product risk, operation risk, investment risk, foreign exchange risk, personnel risk, institutional risk, etc., so that there are dozens or even hundreds of dimensions of data input for the enterprise to be predicted at each time point. High-dimensional data will have a negative impact on subsequent risk identification and model training, reducing the accuracy of risk identification.

[0057] Based on this, the embodiment of the present application determines similar candidate risk factors according to the similarity between the multiple candidate risk factors, so as to cluster the multiple candidate risk factors to obtain a clustering result.

[0058] Exemplarily, the multiple candidate risk factors can be clustered according to the semantic similarity of each candidate risk factor. For example, according to the word vectors corresponding to each candidate risk factor, the multiple candidate risk factors are semantically clustered to obtain multiple clustering sets, so that the candidate risk factors are divided into multiple clustering sets according to the semantic clustering method. The candidate risk factors in the same clustering set express relatively similar semantics. For example, the candidate risk factors "financial statements" and "consumption records" will be divided into the same clustering set to represent the semantics related to the enterprise economy.

[0059] The present application does not specifically limit the semantic clustering method, such as using models such as K-means clustering model, K-center clustering model, Density-Based Spatial Clustering of Applications with Noise (DBSCAN) for semantic clustering.

[0060] Step S230: Calculate the selection weight of the candidate risk factor according to the clustering result, and use the candidate risk factor whose selection weight meets the preset condition as the target risk factor; wherein, the selection weight is used to indicate the contribution degree of the candidate risk factor to the risk identification result of the object to be identified.

[0061] After clustering the candidate risk factors, although the dimension of the risk factors is reduced, not all the candidate risk factors obtained based on the associated data are necessarily useful for risk identification. Therefore, there must be a large amount of redundant data in the candidate risk factors, which affects the prediction effect of risk identification. Therefore, the embodiments of the present application screen the candidate risk factors to remove redundant data and further reduce the dimension of the risk factors, improving the accuracy of subsequent risk identification.

[0062] It should be noted that the candidate risk factors and the target risk factors involved in the embodiments of the present application are essentially risk factor data. Only different names are used to distinguish different stages in the data screening process of the risk factors, so as to accurately understand the process of screening the target risk factors suitable for risk identification from a large number of candidate risk factors in the embodiments of the present application.

[0063] The embodiments of the present application calculate the selection weights of the candidate risk factors according to the clustering results, take the candidate risk factors whose selection weights meet the preset conditions as the target risk factors, and then screen the candidate risk factors to remove the candidate risk factors with a low contribution degree to the risk identification result of the object to be identified, and retain the candidate risk factors with a high contribution degree to the risk identification result of the object to be identified.

[0064] Among them, the higher the contribution degree of the candidate risk factor to the risk identification result of the object to be identified, the higher the possibility that the information contained in the candidate risk factor leads to risk, that is, the higher the selection weight of the candidate risk factor; the lower the contribution degree of the candidate risk factor to the risk identification result of the object to be identified, the lower the possibility that the information contained in the candidate risk factor leads to risk, that is, the lower the selection weight of the candidate risk factor.

[0065] Exemplarily, a feature selection algorithm or a machine learning algorithm can be used to calculate the selection weights of the candidate risk factors. In the embodiments of the present application, there may be correlations among the candidate risk factors in each obtained clustering set. A feature selection algorithm can be used, and a machine learning algorithm can also be combined to further screen the candidate risk factors in the clustering set to obtain the target risk factors with a high contribution degree to the risk identification result of the object to be identified, and at the same time ensure the dimension number of the screened target risk factors to avoid reducing the accuracy of risk identification due to too low a dimension. Among them, regarding the implementation method of screening the candidate risk factors to obtain the target risk factors, reference can be made to the following embodiments for description, but it is not limited to the implementation methods described in the present application, and the present application will not elaborate here.

[0066] Step S240, extract risk features from the target risk factors according to the preset risk knowledge graph to obtain object risk features.

[0067] It should be noted that the preset risk knowledge graph is obtained based on the associated data of all objects. For example, when the object to be recognized is an enterprise, the associated data of all enterprises is obtained, and the corresponding risk knowledge graph is generated based on these associated data.

[0068] In some embodiments, before generating the risk knowledge graph, it further includes preprocessing the associated data.

[0069] Exemplarily, the associated data crawled from news web pages by a crawler tool includes a large amount of advertisements, page head and tail information, etc. Therefore, denoising processing is required. In the denoising processing, the algorithms that can be used include at least one of the following:

[0070] 1. Character filtering. For example, other characters can be removed by the numbers of Chinese characters and common punctuation marks in the American Standard Code for Information Interchange (ASCII) code.

[0071] 2. Hyper Text Markup Language (HTML) field matching. By identifying symbols indicating titles, contents, etc. in HTML text, such as <title>< / title> <content> Symbols such as are used to extract key text content.

[0072] 3. Cross-check the web page results. By comparing the crawling results of different pages of the same website, remove the duplicate content (such as navigation bars, advertisements, logos, etc.) matched by the web page.

[0073] After denoising the associated data, it also includes performing word segmentation and stop word removal on the denoised associated data. For example, the sequence labeling method can be used to perform word segmentation on the denoised associated data, and data cleaning is performed on the results of the word segmentation process. By performing data cleaning on the obtained associated data, the situation of errors in subsequent processing caused by the defects existing in the associated data itself can be avoided.

[0074] Then, perform text vectorization on the associated data after word segmentation and stop word removal. For example, it can be to fuse the ALBERT (A Lite Bidirectional Encoder Representations from Transformers) model and the TinyBERT (Tiny Bidirectional Encoder Representations from Transformers) model, and use the distillation technology to compress the training load of ALBERT to achieve high-speed and efficient text vectorization. It can be understood that the specific algorithm used for text vectorization of the associated data can be selected according to the actual situation, and this application does not make specific limitations on this.

[0075] Furthermore, generate a risk knowledge graph based on the preprocessed associated data.

[0076] Please refer to Figure 3 , Figure 3 which is a schematic diagram of generating a risk knowledge graph in the scenario of enterprise risk identification, as Figure 3 shown:

[0077] Perform text topic classification on the associated data of each enterprise. It can be to perform text topic classification on the associated data of each enterprise by combining the Attention mechanism and the Bi-LSTM (Bi-directional Long Short-Term Memory) text classification model. Since the cell state of the continuously input Bi-LSTM will gradually lose the information of the previous input, but sometimes the important information is actually in the front, the Attention mechanism is used to assign a large attention weight to the important information for strengthening and a small weight to the unimportant information for weakening, thereby improving the accuracy of text topic classification.

[0078] Perform sentiment analysis on the classified associated data. Sentiment analysis can use the Multi-Glance Mechanism (MGM), which is mainly responsible for multi-faceted semantic extraction within the target domain to simulate the reading behavior habit, that is, when reading a text, obtaining a general meaning by skimming a passage of text, and then re-reading the text based on the obtained rough information to extract more important key content, thereby improving the accuracy of sentiment analysis.

[0079] Perform entity extraction on the associated data after sentiment analysis. In the specific implementation of enterprise entity extraction, a multi-classification model is adopted. Each word or term can belong to multiple entity categories simultaneously. The classifier can adopt the softmax classification method or the multi-layer single-classification logistic regression classification method. The loss function used during the training of the classifier can be the Binary CrossEntropy (BCE) or the Kullback-Leibler Divergence (KL-divergence). This application does not limit this.

[0080] Perform entity disambiguation on the extracted entities. Entity ambiguity means that the same entity reference can refer to different entities in different contexts. In the embodiments of this application, entity disambiguation can adopt methods such as entity disambiguation based on clustering and entity disambiguation based on entity linking. This application does not limit this.

[0081] Perform relationship recognition on the disambiguated entities. Exemplarily, enterprise relationships can include raw material production, equity relationship, upstream and downstream of technology services, upstream and downstream of sales channels, investment relationship, technology competition, direct supplier, direct service object, etc.

[0082] Construct triples based on the recognized relationships and entities. For example, the triple can be "entity - event - sentiment".

[0083] Construct a risk knowledge graph based on the triples. Through all the triple information obtained, a triple set is obtained, and a corresponding risk knowledge graph is obtained based on the triple set. The risk knowledge graph includes nodes and edges. Among them, the nodes are the corresponding entity information, and the edges are used to connect two nodes, which refer to the directed lines connecting nodes in the knowledge graph and are used to represent the relationships between different nodes.

[0084] Optionally, the associated data of each enterprise can be obtained periodically to update the risk knowledge graph, thereby ensuring the effectiveness of the risk knowledge graph and avoiding inaccurate risk identification caused by information lag.

[0085] Further, entity extraction is performed on the target risk factor pair according to the risk knowledge graph to obtain the entities included in the target risk factor. For example, the server can use an entity recognition tool to obtain the entities included in the target risk factor. The entity recognition tool is obtained based on entity recognition technology, which can be TexSmart (a text understanding tool and service), or other entity recognition tools. This application does not limit this.

[0086] After the server obtains the entities included in the target risk factor, it can use entity linking technology to link the entities included in the target risk factor to the corresponding entities in the pre-established risk knowledge graph. It should be noted that the corresponding entities in the pre-established risk knowledge graph do not need to be exactly the same as the entities in the target risk factor. For example, "retail store" and "supermarket" can be said to be the same entity. The server can obtain the entities and relationships around the corresponding entity in the pre-established risk knowledge graph, so as to obtain the target knowledge graph associated with the target risk factor. For example, the entity included in the target risk factor is "supermarket", and the target knowledge graph obtained by the server includes: triple <supermarket, turnover, x yuan>, triple <supermarket, business hours, 9:00 am to 9:00 pm>, etc. Therefore, by extracting risk features from the target knowledge graph corresponding to the target risk factor, object risk features are obtained.

[0087] Step S250, determine the risk recognition result of the object to be recognized according to the object risk feature.

[0088] Exemplarily, a risk recognition model can be called, and the object risk feature is input into the risk recognition model for risk recognition to obtain the risk recognition result output by the risk recognition model.

[0089] Among them, the training method of the preset enterprise risk level evaluation model can include: collecting the historical data of the sample object and the historical risks corresponding to the historical data, using the historical data as the input, and using the historical risks corresponding to the historical data as the target output result, and performing model deep learning training on the preset neural network basic model to obtain the risk recognition model.

[0090] The risk identification method provided by this application extracts risk factors from the associated data of the object to be identified, obtains multiple candidate risk factors, performs clustering processing on the multiple candidate risk factors according to the similarity between the multiple candidate risk factors, obtains a clustering result, and then calculates the selection weights of the candidate risk factors according to the clustering result, and uses the candidate risk factors whose selection weights meet the preset conditions as target risk factors, thereby reducing the dimension of the risk factors. Then, according to the preset risk knowledge graph, risk feature extraction is performed on the target risk factors with reduced dimension to obtain object risk features, and the risk identification result of the object to be identified is determined according to the object risk features, making the obtained risk identification result more accurate.

[0091] Please refer to Figure 4 , Figure 4 which is a flowchart of a risk identification shown in another exemplary embodiment. As Figure 4 shown, in an exemplary embodiment, in step S220, performing clustering processing on the multiple candidate risk factors according to the similarity between the multiple candidate risk factors to obtain a clustering result may include the following steps:

[0092] Step S221, splitting the multiple candidate risk factors according to the data generation time corresponding to each of the multiple candidate risk factors to obtain multiple candidate risk factor sequences.

[0093] Among them, all the data related to the candidate risk factors in the associated data of the object to be identified is used as panel data. Panel data has two dimensions, cross-section and time series. It is the repeated measurement data of individuals at different time points on the cross-section. From the cross-section perspective, panel data is the cross-sectional observation values composed of several individuals at a certain time point, and from the longitudinal section perspective, each individual is a time series.

[0094] By splitting the panel data corresponding to the candidate risk factors according to the data generation time, T cross-sectional data are obtained, and T candidate risk factor sequences are obtained, where T is the number of time points corresponding to the data generation time.

[0095] Among them, the candidate risk factor sequence learns a one-dimensional representation through the following formula:

[0096] min‖L‖ * +λ‖S‖ 1

[0097] s.t.X=L+S

[0098] Among them, X is a candidate risk factor sequence corresponding to a certain time point, L is the low-dimensional representation, S is the noise, and the low-dimensional representations of all candidate risk factor sequences form a new data set as the subsequent input for clustering processing.

[0099] Step S222: Cluster the candidate risk factor sequences according to the similarity between each pair of candidate risk factor sequences to obtain a clustering result.

[0100] In some embodiments, clustering the candidate risk factor sequences according to the similarity between each pair of candidate risk factor sequences to obtain a clustering result includes: constructing a distance matrix based on the lengths of each candidate risk factor sequence, where each position in the distance matrix represents the distance between each pair of candidate risk factor sequences; calculating a first cumulative distance from a starting position to a target position in the distance matrix, and calculating a second cumulative distance from an ending position to the target position in the distance matrix; calculating the minimum distance between each pair of candidate risk factor sequences according to the first cumulative distance and the second cumulative distance, and determining the similarity between each pair of candidate risk factor sequences according to the minimum distance.

[0101] Please refer to Figure 5 , Figure 5 which is a schematic diagram for clustering the candidate risk factor sequences. As Figure 5 shown, obtain a candidate risk factor sequence X corresponding to a certain time point, and obtain the one-dimensional representation of X. Then, obtain a distance matrix according to the one-dimensional representations of each candidate risk factor sequence. Each position in the distance matrix represents the distance between a point on one candidate risk factor sequence and a point on another candidate risk factor sequence, and this distance can be the Euclidean distance.

[0102] The starting position in the distance matrix is the position corresponding to the first point on one candidate risk factor sequence and the first points on other candidate risk factor sequences in the distance matrix. The ending position in the distance matrix is the position corresponding to the last point on one candidate risk factor sequence and the last points on other candidate risk factor sequences in the distance matrix. The target position in the distance matrix can be a position other than the starting position and the ending position in the distance matrix.

[0103] Calculate the cumulative distances from the starting position to multiple first candidate positions associated with the target position in the distance matrix respectively, where the first candidate positions are located between the starting position and the target position. For example, the distance accumulation calculation can be performed position by position from three directions on the matrix. Then, according to the cumulative distances from the starting position to each first candidate position and the distance values represented by each first candidate position, calculate multiple first candidate cumulative distances from the starting position to the target position. Then, take the minimum value among the multiple first candidate cumulative distances as the first cumulative distance.

[0104] The process of calculating the second cumulative distance from the ending position to the target position is similar to the process of calculating the first cumulative distance from the starting position to the target position, and this application will not elaborate here.

[0105] Further, calculate the minimum distance between each candidate risk factor sequence according to the first cumulative distance and the second cumulative distance, and determine the similarity between each candidate risk factor sequence according to the minimum distance. For example, the distance value represented by the target position, the first cumulative distance, and the second cumulative distance can be summed to obtain the minimum cumulative distance corresponding to the target position, and this minimum cumulative distance can be used as the similarity between the candidate risk factor sequences corresponding to the target position.

[0106] Please refer to Figure 6 , Figure 6 which is a flowchart of risk identification shown in another exemplary embodiment. As Figure 6 shown, in an exemplary embodiment, the clustering result includes multiple clustering sets. In step S230, calculate the selection weight of the candidate risk factors according to the clustering result, and use the candidate risk factors whose selection weight meets the preset conditions as the target risk factors, which may include the following steps:

[0107] Step S231, determine the clustering center vector corresponding to each clustering set.

[0108] The clustering center vector refers to a special sample in cluster analysis, which can be used to represent a certain category. Other data in the clustering set can determine whether they belong to this category by calculating the distance from it.

[0109] The clustering center vector can be a vector determined in advance based on the feature vectors corresponding to each candidate risk factor included in the clustering set. Generally, the clustering center vector is used to characterize the center point of the feature vector clustering composed of multiple feature vectors (that is, the feature vectors corresponding to each candidate risk factor included in the clustering set). Therefore, the clustering center vector can accurately characterize the average features of the candidate risk factors included in the clustering set.

[0110] As an example, the elements at the same position in each feature vector can be averaged as the value of the corresponding position element in the clustering center vector. Or, the median of the elements at the same position in each feature vector can be taken as the value of the corresponding position element in the clustering center vector. It should be understood that the clustering center vector can also be obtained by other methods, which will not be listed one by one here.

[0111] Step S232, calculate the probability that the candidate risk factor belongs to each clustering set according to the clustering center vector, so as to generate a weak label matrix of the candidate risk factor according to the probability.

[0112] Among them, the probability that the candidate risk factor belongs to each clustering set is the distance between the candidate risk factor and the clustering center vector corresponding to each clustering set.

[0113] Step S233: Calculate the selection weights of the candidate risk factors based on the feature selection matrix and the weak label matrix of the candidate risk factors, and use the candidate risk factors whose selection weights meet the preset conditions as the target risk factors. Among them, the feature selection matrix is obtained through deep learning training based on the sample risk factors and sample risk identification results in the training samples.

[0114] Among them, the candidate risk factors whose selection weights meet the preset conditions can be the candidate risk factors whose weights are greater than or equal to the selection weight threshold, or the candidate risk factors can be sorted according to the selection weights, and the candidate risk factors with the top preset number in the sequence are used as the candidate risk factors that meet the preset conditions.

[0115] Exemplarily, the selection of the target risk factors can refer to the following formula:

[0116]

[0117]

[0118] Among them, o j is the cluster center vector of the j-th cluster set in the low-dimensional representation space, n is the number of candidate risk factors, c is the number of cluster sets, h ij represents the possibility that the candidate risk factor x i belongs to the j-th class, and the weak label matrix H of the candidate risk factor x i is obtained. I is the identity matrix, represents the square of the largest singular value of the matrix, represents the sum of the squares of the matrix elements. P is the feature selection matrix, and P T is the transpose operation of the matrix P, and P T X is the required target risk factor.

[0119] Among them, the feature selection matrix P and the parameters α, β, λ are obtained through deep learning training based on the sample risk factors and sample risk identification results in the training samples. For example, the sample risk factors in the training samples are used as the input, and the sample target risk factors corresponding to the sample risk identification results are used as the output. Then, the feature selection matrix P and the parameters α, β, λ are adjusted according to the gap between the actually output target risk factors and the sample target risk factors. When the gap between the actually output target risk factors and the sample target risk factors is less than the threshold, the trained feature selection matrix P and the parameters α, β, λ are obtained.

[0120] By selecting the clustered candidate risk factors, target risk factors with reduced dimensions and high discriminant information are obtained, improving the accuracy of subsequent risk identification.

[0121] In some embodiments, during the deep learning training process, a distributed machine learning scheduling framework is adopted. For example, during the training process of machine learning model parameters, learning tasks include GPU (Graphics Processing Unit) versions and CPU (Central Processing Unit) versions. Before each iteration operation in the training process, first obtain the number m of available CPU devices and the number n of GPU devices, and then determine the ratio of the running time on the CPU and GPU according to the historical allocation statistical data of the learning tasks (this ratio is equivalent to the execution efficiency ratio of the CPU and GPU). According to the ratio, the learning tasks can be decomposed into p and q tasks. Then submit the GPU tasks to the GPU computing resources and the CPU tasks to the CPU computing resources. Finally, ensure the synchronous execution between the allocated CPU tasks and GPU tasks, that is, there is no lag between the CPU tasks and GPU tasks, so as to improve the speed of deep learning training.

[0122] In some embodiments, risk feature extraction is performed on the target risk factor according to a preset risk knowledge graph to obtain the object risk feature, including: determining the risk entity corresponding to the risk factor; extracting the sub-graph matching the risk entity from the risk knowledge graph; encoding each node in the sub-graph to obtain the node feature; and fusing the node features of each node to obtain the object risk feature.

[0123] After obtaining the target risk factor, the target risk factor can be matched according to the risk knowledge graph to obtain the corresponding sub-graph. The entity corresponding to each sub-graph node in the sub-graph is an entity with a matching degree greater than the matching degree threshold with the target risk factor.

[0124] Exemplarily, calculate the matching degree between the target risk factor and the entities corresponding to each graph node in the risk knowledge graph. When the matching degree is larger, it indicates that the corresponding entity is more similar to the target risk factor; when the matching degree is smaller, it indicates that the corresponding entity is more different from the target risk factor. Then, select the graph nodes with a matching degree greater than the matching degree threshold as the target nodes to obtain the sub-graph according to the target nodes.

[0125] Then, encode each node in the sub-graph to obtain the node feature of each node. For example, encode according to the node content and node position of each node to obtain the node content feature and node position feature, and splice the node content feature and node position feature to obtain the node feature of each node.

[0126] Furthermore, fuse the node features of each node to obtain the object risk feature.

[0127] Please refer to Figure 7 , Figure 7 is a flowchart of a risk identification shown in another exemplary embodiment. As Figure 7 shown, in an exemplary embodiment, the clustering result includes multiple clustering sets. Determining the risk identification result of the object to be identified according to the object risk characteristics in step S250 may include the following steps:

[0128] Step S251, obtaining the risk data corresponding to the associated object having an association relationship with the object to be identified.

[0129] In the embodiment of the present application, a relationship network is constructed based on the upstream and downstream relationships between each object, and the relationship in this network is the potential path for the conduction of risk occurrence.

[0130] Therefore, the associated object having an association relationship with the object to be identified is obtained through the relationship network, and then the risk data corresponding to each associated object is obtained. For example, when the object to be identified is an enterprise, the risk data of the associated enterprise having an association relationship with the enterprise to be identified is obtained.

[0131] Step S252, performing risk conduction calculation on the risk data according to the category of the association relationship to obtain the risk conduction characteristics of the associated object relative to the object to be identified.

[0132] Exemplarily, the risk conduction calculation of the risk data can refer to the following formula:

[0133]

[0134]

[0135]

[0136] where a is the risk degree of enterprise V in the relationship network i Pow(E c ) is the weight based on the data-driven relationship Str(E i,j ) is the risk association degree between enterprise V i and enterprise V j , α is a hyperparameter, which can be determined by the cross-validation method, and its default value is 0.5, represents that after enterprise V i has a risk, it is transmitted to enterprise V j through the association relationship C. If there is only an association relationship C between enterprise V i and enterprise V j but no risk conduction occurs, then

[0137] where Str(E i,j ) The larger the value, the greater the feasibility of risk conduction through this edge.

[0138] In some embodiments, risk conduction calculation is performed on risk data according to the category of the association relationship to obtain the risk conduction characteristics of the associated object relative to the object to be identified, including: calculating the risk association degree between the risk data and the object to be identified; and determining the weight coefficient corresponding to the associated object according to the category of the association relationship; performing weighted calculation on the risk association degree according to the weight coefficient to obtain the risk conduction characteristics of the associated object relative to the object to be identified.

[0139] It can be understood that the probability of risk conduction between the associated object and the object to be identified is different due to the different association relationships between them. Therefore, different weight coefficients are set for different association relationships, and then weighted calculation is performed on the risk association degree according to the weight coefficient to obtain the risk conduction degree of the associated object relative to the object to be identified, and this risk conduction degree is used as the risk conduction characteristic.

[0140] Step S253, determining the risk identification result of the object to be identified according to the object risk characteristics and the risk conduction characteristics.

[0141] Since there may be a risk conduction phenomenon among multiple objects, therefore, by combining the object risk characteristics of the object to be identified itself and the risk conduction characteristics of other objects associated with the object to be identified, the risk identification result of the object to be identified is obtained, making the obtained risk identification result more accurate.

[0142] Taking the risk identification scenario for enterprises as an example, the risk identification process is described as follows:

[0143] Exemplarily, the associated data of the object to be identified is obtained based on the retrieval information. As Figure 8 shown, in practical applications, the user can input retrieval information in the search interface. Among them, the retrieval information reflects the user's retrieval intention, and the specific form of the retrieval information can be text, image, etc. For example, the retrieval information obtained by the server can be the text "Enterprise A", or an image containing the trademark of "Enterprise A". The search interface can be the interface entered through the search entry provided by the enterprise risk analysis software, the interface where the search bar is located in information software such as video and news, etc. The object to be identified is obtained based on the retrieval term input by the user. For example, the server stores the enterprise names of multiple enterprises, and matches these enterprise names with the retrieval term, and takes the successfully matched enterprise name as the object to be identified. Then, the server can search for the associated data corresponding to the object to be identified, and the associated data can be text, audio, video, picture, etc., and the text can be data structures such as documents, news, web pages, etc.

[0144] Determining associated data based on the object to be recognized can be divided into two cases, which are described separately below.

[0145] Case 1: If the associated data determined based on the object to be recognized is in text form, then the text in any one of the retrieval results in text form is used as the associated data.

[0146] Case 2: If the associated data determined based on the object to be recognized is in non-text form, such as video, audio, picture, etc., then the retrieval results in non-text form are converted into their corresponding text forms. For example, the audio in the video is extracted, and the audio is converted into the corresponding text based on semantics, and the converted text is used as the associated data.

[0147] Then, risk factors are extracted from the obtained associated data of the object to be recognized to obtain multiple candidate risk factors, and the candidate risk factors are clustered to obtain a clustering result. According to the clustering result, the selection weights of the candidate risk factors are calculated, and the candidate risk factors with selection weights greater than or equal to the selection weight threshold are used as target risk factors. Furthermore, risk feature extraction is performed on the target risk factors according to the risk knowledge graph to obtain object risk features. At the same time, the risk data corresponding to the associated object having an associated relationship with the object to be recognized is obtained, and risk conduction calculation is performed according to the risk data to obtain the risk conduction feature of the associated object relative to the object to be recognized. Finally, the risk recognition result of the object to be recognized is determined according to the object risk feature and the risk conduction feature.

[0148] The risk recognition method provided by this application extracts risk factors from the associated data of the object to be recognized to obtain multiple candidate risk factors, and clusters the multiple candidate risk factors according to the similarity between the multiple candidate risk factors to obtain a clustering result. Then, the selection weights of the candidate risk factors are calculated according to the clustering result, and the candidate risk factors whose selection weights meet the preset conditions are used as target risk factors, thereby reducing the dimension of the risk factors. Then, risk feature extraction is performed on the target risk factors with reduced dimension according to the preset risk knowledge graph to obtain object risk features, and the risk recognition result of the object to be recognized is determined according to the object risk features, making the obtained risk recognition result more accurate.

[0149] Please refer to Figure 9 , Figure 9 which is a schematic block diagram of a risk recognition device 900 provided by an embodiment of this application. The risk recognition device 900 can be configured in a server or a terminal and is used to execute the foregoing risk recognition method.

[0150] As Figure 9 As shown in the figure, the risk identification device 900 includes: a risk factor extraction module 910, a clustering module 920, a target risk factor selection module 930, a risk feature extraction module 940, and a risk identification module 950.

[0151] The risk factor extraction module 910 is configured to extract risk factors from the associated data of the object to be identified, and obtain a plurality of candidate risk factors;

[0152] The clustering module 920 is configured to perform clustering processing on the plurality of candidate risk factors according to the similarity between the plurality of candidate risk factors, and obtain a clustering result;

[0153] The target risk factor selection module 930 is configured to calculate the selection weight of the candidate risk factors according to the clustering result, and use the candidate risk factors whose selection weight meets the preset conditions as the target risk factors; wherein, the selection weight is used to indicate the contribution degree of the candidate risk factors to determining the risk identification result of the object to be identified;

[0154] The risk feature extraction module 940 is configured to extract risk features from the target risk factors according to a preset risk knowledge graph, and obtain object risk features;

[0155] The risk identification module 950 is configured to determine the risk identification result of the object to be identified according to the object risk features.

[0156] In some embodiments, based on the foregoing solution, the clustering module 920 includes a splitting unit and a clustering unit.

[0157] The splitting unit is configured to split the plurality of candidate risk factors according to the data generation time corresponding to each of the plurality of candidate risk factors, and obtain a plurality of candidate risk factor sequences;

[0158] The clustering unit is configured to perform clustering processing on the candidate risk factor sequences according to the similarity between the candidate risk factor sequences, and obtain a clustering result.

[0159] In some embodiments, based on the foregoing solution, the clustering unit includes a distance matrix construction unit, a distance calculation unit, and a similarity determination unit.

[0160] The distance matrix construction unit is configured to construct a distance matrix according to the lengths of the candidate risk factor sequences, and each position in the distance matrix represents the distance between the candidate risk factor sequences;

[0161] The distance calculation unit is configured to calculate a first cumulative distance between a starting position and a target position in the distance matrix, and calculate a second cumulative distance between an ending position and the target position in the distance matrix;

[0162] A similarity determination unit, configured to calculate the minimum distance between each candidate risk factor sequence according to the first cumulative distance and the second cumulative distance, and determine the similarity between each candidate risk factor sequence according to the minimum distance.

[0163] In some embodiments, based on the foregoing solution, the clustering result includes a plurality of clustering sets; the target risk factor selection module 930 includes a clustering center vector determination unit, a weak label matrix generation unit, and a selection weight determination unit.

[0164] The clustering center vector determination unit is configured to determine the clustering center vector corresponding to each clustering set;

[0165] The weak label matrix generation unit is configured to calculate the probability that a candidate risk factor belongs to each clustering set according to the clustering center vector, so as to generate a weak label matrix of the candidate risk factor according to the probability;

[0166] The selection weight determination unit is configured to calculate the selection weight of the candidate risk factor according to the feature selection matrix and the weak label matrix of the candidate risk factor, and use the candidate risk factor whose selection weight meets the preset conditions as the target risk factor; wherein, the feature selection matrix is obtained through deep learning training according to the sample risk factors and sample risk recognition results in the training samples.

[0167] In some embodiments, based on the foregoing solution, the risk feature extraction module 940 includes a risk entity determination unit, a sub-graph extraction unit, an encoding unit, and a fusion unit.

[0168] The risk entity determination unit is configured to determine the risk entity corresponding to the target risk factor;

[0169] The sub-graph extraction unit is configured to extract a sub-graph matching the risk entity from the risk knowledge graph;

[0170] The encoding unit is configured to encode each node in the sub-graph to obtain node features;

[0171] The fusion unit is configured to fuse the node features of each node to obtain the object risk feature.

[0172] In some embodiments, based on the foregoing solution, the risk identification module 950 includes an association acquisition unit, a risk conduction feature acquisition unit, and a comprehensive identification unit.

[0173] The association acquisition unit is configured to acquire the risk data of the associated object having an association relationship with the object to be identified;

[0174] The risk conduction feature acquisition unit is configured to perform risk conduction calculation on the risk data according to the category of the association relationship to obtain the risk conduction feature of the associated object relative to the object to be identified;

[0175] The comprehensive recognition unit is configured to determine the risk recognition result of the object to be recognized according to the object risk characteristics and the risk conduction characteristics.

[0176] In some embodiments, based on the foregoing solution, the risk conduction characteristic acquisition unit includes a data determination unit and a weighted calculation unit.

[0177] The data determination unit is configured to calculate the risk correlation degree between the risk data and the object to be recognized; and determine the weight coefficient corresponding to the associated object according to the category of the association relationship;

[0178] The weighted calculation unit is configured to perform weighted calculation on the risk correlation degree according to the weight coefficient to obtain the risk conduction characteristic of the associated object relative to the object to be recognized.

[0179] It should be noted that the risk recognition device provided in the above embodiment and the risk recognition method provided in the above embodiment belong to the same concept. The specific manners in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In practical applications, the risk recognition device provided in the above embodiment can, as needed, allocate the above functions to different functional modules, that is, divide the internal structure of the device into different functional modules to complete all or part of the functions described above. This is not limited here.

[0180] The methods and devices of the present application can be used in many general-purpose or special-purpose computing system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, small computers, large computers, distributed computing environments including any of the above systems or devices, and so on.

[0181] Figure 10 The structural schematic diagram of the computer system of the electronic device suitable for implementing the embodiments of the present application is shown.

[0182] It should be noted that Figure 10 The computer system 1000 of the electronic device shown is only an example and should not impose any limitation on the functions and usage scope of the embodiments of the present application.

[0183] As Figure 10 shown, the electronic device 1000 is presented in the form of a general-purpose computing device. The components of the electronic device 1000 may include, but are not limited to: at least one of the above processing units 1010, at least one of the above storage units 1020, a bus 1030 connecting different system components (including the storage unit 1020 and the processing unit 1010), and a display unit 1040.

[0184] Among them, the storage unit stores program code, which can be executed by the processing unit 1010, so that the processing unit 1010 executes the steps according to various exemplary embodiments of the present disclosure described in the "Exemplary Method" section of this specification.

[0185] The storage unit 1020 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 1021 and / or a cache storage unit 1022, and may further include a read-only storage unit (ROM) 1023.

[0186] The storage unit 1020 may also include a program / utilities 1024 having a set (at least one) of program modules 1025. Such program modules 1025 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0187] The bus 1030 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0188] The electronic device 1000 may also communicate with one or more external devices 1070 (such as a keyboard, a pointing device, a Bluetooth device, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 1000, and / or communicate with any device that enables the electronic device 1000 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 1050. In addition, the electronic device 1000 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 1060. As shown in the figure, the network adapter 1060 communicates with other modules of the electronic device 1000 through the bus 1030. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 1000, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0189] In particular, according to an embodiment of the present application, the processes described above with reference to the flowchart can be implemented as computer software programs. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a computer program for executing the method shown in the flowchart. When the computer program is executed by the processing unit 1010, various functions defined in the system of the present application are executed.

[0190] It should be noted that the computer-readable medium shown in the embodiments of the present application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0191] In the present application, a computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In the present application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable computer program. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The computer program contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0192] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. Among them, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code, and the above-mentioned module, program segment, or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in an order different from that marked in the accompanying drawings.

[0193] The units described in the embodiments of the present application can be implemented in software or in hardware, and the described units can also be set in a processor. Among them, the names of these units do not constitute a limitation to the unit itself in some cases.

[0194] Another aspect of the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the risk identification method as described above is implemented. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist alone without being assembled into the electronic device.

[0195] Another aspect of the present application also provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the risk identification method provided in the above various embodiments.

[0196] The above content is only a preferred exemplary embodiment of the present application and is not used to limit the implementation of the present application. Those of ordinary skill in the art can easily make corresponding changes or modifications according to the main idea and spirit of the present application. Therefore, the protection scope of the present application should be subject to the protection scope required by the claims. < / content>

Claims

1. A risk identification method, characterized in that, it includes: Crawling historical associated data of each enterprise from a web page, where the historical associated data includes historical risk events and enterprise data associated with the historical risk events; Using natural language processing to process the historical associated data of each enterprise obtained by crawling, to obtain multiple risk factors to be stored included in the enterprise data associated with the historical risk events, and the risk factors to be stored include financial factors and event factors; Performing association analysis on the historical risk events and the enterprise data associated with the historical risk events to obtain the association strength of each of the risk factors to be stored; Storing the risk factors to be stored with an association strength greater than or equal to a preset association strength threshold into a risk factor set; Extracting risk factors from the associated data of the object to be identified through the risk factor set to obtain multiple candidate risk factors; Splitting the multiple candidate risk factors according to the data generation time corresponding to each of the multiple candidate risk factors to obtain multiple candidate risk factor sequences; Constructing a distance matrix according to the lengths of each of the candidate risk factor sequences, where each position in the distance matrix represents the distance between each of the candidate risk factor sequences; Calculating a first cumulative distance between a starting position in the distance matrix and a target position in the distance matrix, and calculating a second cumulative distance between an ending position in the distance matrix and the target position; Calculating the minimum distance between each of the candidate risk factor sequences according to the first cumulative distance and the second cumulative distance, and determining the similarity between each of the candidate risk factor sequences according to the minimum distance; Performing clustering processing on the candidate risk factor sequences according to the similarity between each of the candidate risk factor sequences to obtain a clustering result; Calculating the selection weight of the candidate risk factor according to the clustering result, and taking the candidate risk factor whose selection weight meets the preset condition as the target risk factor; wherein, the selection weight is used to indicate the contribution degree of the candidate risk factor to determining the risk identification result of the object to be identified; Extracting risk characteristics of the target risk factor according to a preset risk knowledge graph to obtain object risk characteristics; Determining the risk identification result of the object to be identified according to the object risk characteristics.

2. The method according to claim 1, characterized in that, the clustering result includes multiple clustering sets; the calculating the selection weight of the candidate risk factor according to the clustering result and taking the candidate risk factor whose selection weight meets the preset condition as the target risk factor includes: Determining the clustering center vector corresponding to each of the clustering sets; Calculating the probability that the candidate risk factor belongs to each of the clustering sets according to the clustering center vector, so as to generate a weak label matrix of the candidate risk factor according to the probability; Calculate the selection weights of the candidate risk factors according to the feature selection matrix and the weak label matrix of the candidate risk factors, and use the candidate risk factors whose selection weights meet the preset conditions as the target risk factors; wherein, the feature selection matrix is obtained through deep learning training based on the sample risk factors and sample risk recognition results in the training samples.

3. The method according to claim 1, wherein, the risk feature extraction of the target risk factor according to the preset risk knowledge graph to obtain the object risk feature includes: Determine the risk entity corresponding to the target risk factor; Extract the sub-graph matching the risk entity from the risk knowledge graph; Encode each node in the sub-graph to obtain node features; Fuse the node features of each node to obtain the object risk feature.

4. The method according to claim 1, wherein, the determination of the risk recognition result of the object to be recognized according to the object risk feature includes; Obtain the risk data of the associated object having an association relationship with the object to be recognized; Perform risk conduction calculation on the risk data according to the category of the association relationship to obtain the risk conduction feature of the associated object relative to the object to be recognized; Determine the risk recognition result of the object to be recognized according to the object risk feature and the risk conduction feature.

5. The method according to claim 4, wherein, the risk conduction calculation of the risk data according to the category of the association relationship to obtain the risk conduction feature of the associated object relative to the object to be recognized includes: Calculate the risk association degree between the risk data and the object to be recognized; and determine the weight coefficient corresponding to the associated object according to the category of the association relationship; Perform weighted calculation on the risk association degree according to the weight coefficient to obtain the risk conduction feature of the associated object relative to the object to be recognized.

6. A risk recognition device, wherein, the device includes: A risk factor extraction module, configured to crawl the historical association data of each enterprise from the web page, where the historical association data includes historical risk events and enterprise data associated with the historical risk events; use natural language processing to process the historical association data of each enterprise obtained by crawling to obtain multiple risk factors to be stored included in the enterprise data associated with the historical risk events, and the risk factors to be stored include financial factors and event factors; perform association analysis on the historical risk events and the enterprise data associated with the historical risk events to obtain the association strength of each risk factor to be stored; store the risk factors to be stored whose association strength is greater than or equal to the preset association strength threshold into the risk factor set; extract risk factors from the association data of the object to be recognized through the risk factor set to obtain multiple candidate risk factors; The clustering module is configured to split the multiple candidate risk factors according to the data generation time corresponding to each of the multiple candidate risk factors, to obtain multiple candidate risk factor sequences; construct a distance matrix according to the lengths of the candidate risk factor sequences, where each position in the distance matrix represents the distance between each of the candidate risk factor sequences; calculate a first cumulative distance between a starting position and a target position in the distance matrix, and calculate a second cumulative distance between an ending position and the target position in the distance matrix; calculate the minimum distance between each of the candidate risk factor sequences according to the first cumulative distance and the second cumulative distance, and determine the similarity between each of the candidate risk factor sequences according to the minimum distance; perform clustering processing on the candidate risk factor sequences according to the similarity between each of the candidate risk factor sequences, to obtain a clustering result; The target risk factor selection module is configured to calculate a selection weight of the candidate risk factor according to the clustering result, and use the candidate risk factor whose selection weight meets a preset condition as a target risk factor; wherein, the selection weight is used to indicate the contribution degree of the candidate risk factor to determining the risk identification result of the object to be identified; The risk feature extraction module is configured to extract risk features from the target risk factor according to a preset risk knowledge graph, to obtain object risk features; The risk identification module is configured to determine the risk identification result of the object to be identified according to the object risk features.

7. A computer device, wherein, the computer device includes a memory and a processor; the memory is used to store a computer program; the processor is configured to execute the computer program and, when executing the computer program, implement the risk identification method according to any one of claims 1 to 5.

8. A computer-readable storage medium, wherein, the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the processor is caused to implement the risk identification method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Service data risk processing method, device and equipment and storage medium

    CN111709661A

  • Risk identification method and device, equipment and storage medium

    CN113989019A