Risk assessment method and device for enterprise
By obtaining and preprocessing enterprise data, determining label tasks and calculating risk values, the problem of traditional regulatory methods being difficult to assess enterprise risks in a timely manner is solved, and efficient risk assessment and early warning is achieved.
Patent Information
- Application Number
- CN202411930422.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-06
AI Technical Summary
Traditional regulatory methods are difficult to detect and evaluate corporate risks in a comprehensive and timely manner, and cannot adapt to the rapid market changes.
By obtaining the initial enterprise data of the target enterprise, pre-processing and generating the target enterprise data, determining the enterprise data label, calculating the risk value based on the label task, and generating risk assessment results.
It has realized the transformation of the complex characteristics of the enterprise into quantifiable indicators, improved the efficiency of enterprise risk assessment, effectively identified potential risks and provided timely warnings and suggestions to decision makers.
Smart Images

Figure CN119940914A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of risk assessment for enterprises, and in particular to a risk assessment method for enterprises, a risk assessment device for enterprises, an electronic device and a computer-readable storage medium. Background Art
[0002] With the vigorous development of the market economy, the number of enterprises has exploded, and business models have become increasingly diverse and complex. As the main force of market supervision, the Industrial and Commercial Bureau faces unprecedented regulatory challenges. Traditional regulatory methods, such as regular on-site inspections and single financial indicator analysis, have been unable to adapt to the rapidly changing market situation and have difficulty in comprehensively and timely discovering and assessing corporate risks. Summary of the invention
[0003] The embodiments of the present invention provide a risk assessment method, device, electronic device and computer-readable storage medium for an enterprise to overcome the above problems or at least partially solve the above problems.
[0004] The embodiment of the present invention discloses a risk assessment method for an enterprise, comprising:
[0005] Obtaining initial enterprise data for target enterprises;
[0006] Performing a preprocessing operation on the initial enterprise data to generate target enterprise data;
[0007] Determine an enterprise data tag for the target enterprise, and determine a tag task based on the enterprise data tag;
[0008] Calculate the risk value for the target enterprise based on the label task;
[0009] A risk assessment result for the target enterprise is generated using the risk value.
[0010] Optionally, the step of obtaining initial enterprise data for the target enterprise includes:
[0011] Determine the target data source;
[0012] Initial enterprise data is acquired from a plurality of different target data sources based on real-time data stream processing technology.
[0013] Optionally, the step of performing a preprocessing operation on the initial enterprise data to generate target enterprise data includes:
[0014] Performing a data cleaning operation on the initial enterprise data, and obtaining text entities from the initial enterprise data after the data cleaning operation; the text entities include enterprise entities and personnel entities;
[0015] Generate a relationship graph in which the user expresses the association relationship between the enterprise entity and the person entity;
[0016] Determine the merged and fused dimension of the target enterprise, and determine the target attributes of the plurality of text entities through the relationship graph based on the merged and fused dimension;
[0017] A knowledge graph is constructed based on the target attributes, the relationship graph and the text entities, and the knowledge graph is determined as target enterprise data.
[0018] Optionally, before the step of constructing a knowledge graph based on the target attribute, the relationship graph and the text entity, and determining the knowledge graph as the target enterprise data, the step further includes:
[0019] Obtaining the occurrence duration, occurrence length, and data source information of the text entity;
[0020] The credibility of the target attribute is calculated based on the appearance duration, the appearance duration and the data source information; the credibility is used to delete repeated text entities and the association relationships.
[0021] Optionally, the step of determining the enterprise data tag for the target enterprise includes:
[0022] Determine the input gate, forget gate, and output gate;
[0023] Generate a neural network deep learning model based on the input gate, the forget gate and the output gate;
[0024] The enterprise data label for the target enterprise is determined through the neural network deep learning model.
[0025] Optionally, the step of calculating the risk value for the target enterprise based on the label task includes:
[0026] Obtaining a task score value of the label task and an indicator weight for the label task;
[0027] The task score values are weightedly summed based on the indicator weights to calculate a risk value for the target enterprise.
[0028] Optionally, the risk assessment result includes visual display information and risk warning information.
[0029] The embodiment of the present invention also discloses a risk assessment device for an enterprise, comprising:
[0030] An initial enterprise data acquisition module is used to acquire initial enterprise data for a target enterprise;
[0031] A target enterprise data generating module, used for performing a preprocessing operation on the initial enterprise data to generate target enterprise data;
[0032] A label task determination module, used to determine an enterprise data label for the target enterprise, and determine a label task based on the enterprise data label;
[0033] A risk value calculation module, used to calculate the risk value for the target enterprise based on the label task;
[0034] The risk assessment result generating module is used to generate a risk assessment result for the target enterprise according to the risk value.
[0035] The embodiment of the present invention further discloses an electronic device, comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other via the communication bus;
[0036] The memory is used to store computer programs;
[0037] The processor is used to implement the method described in the embodiment of the present invention when executing the program stored in the memory.
[0038] The embodiment of the present invention further discloses a computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, enables the processors to execute the method described in the embodiment of the present invention.
[0039] The embodiments of the present invention include the following advantages:
[0040] The embodiment of the present invention obtains initial enterprise data for a target enterprise; performs preprocessing operations on the initial enterprise data to generate target enterprise data; determines an enterprise data label for the target enterprise, and determines a label task based on the enterprise data label; calculates a risk value for the target enterprise based on the label task; and generates a risk assessment result for the target enterprise through the risk value; thereby achieving the goal of converting the complex characteristics of an enterprise into quantifiable indicators by defining a series of labels, and designing corresponding calculation tasks for each label, thereby quantifying the enterprise characteristics, providing actionable indicators for risk assessment, and improving the efficiency of risk assessment for the enterprise.
[0041] Through the above steps, the embodiment of the present invention establishes a complete enterprise risk assessment system. The system can effectively identify the potential risks of the enterprise and provide timely warnings and suggestions to decision makers, thereby helping the enterprise to avoid risks and achieve sustainable development.
[0042] The advantages of this system are:
[0043] Systematic: Covering all aspects of enterprise risk assessment.
[0044] Scientificity: Based on data analysis and model calculation, it has high accuracy.
[0045] Practicality: It provides actionable risk assessment results and provides guidance for decision makers. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] Figure 1 is a flowchart of a risk assessment method for an enterprise provided in an embodiment of the present invention;
[0047] Figure 2 It is a flow chart of a risk assessment method for an enterprise provided in an embodiment of the present invention;
[0048] Figure 3 is a structural block diagram of a risk assessment device for an enterprise provided in an embodiment of the present invention;
[0049] Figure 4 is a hardware structure block diagram of an electronic device provided in an embodiment of the present invention;
[0050] Figure 5 is a schematic diagram of a computer-readable medium provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0051] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0052] With the rapid development of the market economy, the number of enterprises has increased sharply, and the types of enterprises and operating models have become increasingly diversified. As the main body of market supervision, the Industrial and Commercial Bureau faces huge regulatory challenges.
[0053] Traditional supervision methods often rely on regular on-site inspections and single financial indicators, which makes it difficult to comprehensively and timely discover and assess corporate risks. Therefore, developing an enterprise risk assessment and early warning system based on a labeling system is of great significance for improving the supervision efficiency of market supervision departments and ensuring market order.
[0054] In order to enable those skilled in the art to better understand the embodiments of the present invention, some technical names involved in the embodiments of the present invention are explained below.
[0055] 1. Data label: Data label is a form of data organization. It can be manually identified or calculated according to certain logical rules (models) in a specific business scenario or context, and can relatively accurately reflect the description of a certain feature of the labeled entity.
[0056] 2. BERT (Bidirectional Encodex Representations from Transformers): BERT is a language representation model. BERT stands for Bidirectional Encoder Representations from Transformers. BERT aims to pre-train deep bidirectional representations by jointly conditioning on the left and right contexts in all layers.
[0057] 3. Word2Vec: Word2Vec is a popular natural language processing (NLP) tool that converts each word in the vocabulary into a unique high-dimensional space vector so that these word vectors can mathematically represent their semantic relationships.
[0058] 4. RNN (Recurrent Neural Network): A recurrent neural network is a type of recurrent neural network that takes sequence data as input, recursively in the direction of sequence evolution, and all nodes are connected in a chain. The main feature of RNN is that it has a recurrent connection characteristic, which can input the state of the previous moment into the current moment, thereby realizing the modeling of time series data.
[0059] 5. LDA (local-density approximation): Local density approximation is an approximation used in one of the exchange-correlation energy functionals in density functional theory.
[0060] Reference Figure 1 , shows a flowchart of a risk assessment method for an enterprise provided in an embodiment of the present invention, which may specifically include the following steps:
[0061] Step 101, obtaining initial enterprise data for a target enterprise;
[0062] Step 102, performing a preprocessing operation on the initial enterprise data to generate target enterprise data;
[0063] Step 103, determining an enterprise data tag for the target enterprise, and determining a tag task based on the enterprise data tag;
[0064] Step 104, calculating a risk value for the target enterprise based on the label task;
[0065] Step 105: Generate a risk assessment result for the target enterprise using the risk value.
[0066] The embodiment of the present invention can obtain initial enterprise data for a target enterprise;
[0067] Purpose: To collect basic information of the target enterprise and provide a data basis for subsequent analysis.
[0068] Example description: The target company's registration information, financial data, operating conditions, legal records, etc. can be obtained from multiple channels such as the business registration system, tax system, judicial decision system, etc.
[0069] Beneficial effects: Provide comprehensive and accurate original data for subsequent risk assessment.
[0070] In the embodiment of the present invention, a pre-processing operation may be performed on the initial enterprise data to generate target enterprise data;
[0071] Purpose: To clean, convert, integrate and process the raw data to make it meet the requirements of subsequent analysis.
[0072] Example description:
[0073] Data cleaning: remove duplicate data, abnormal data, and fill in missing values.
[0074] Data conversion: Convert data into a unified format, for example, unify the date format into YYYY-MM-DD.
[0075] Data integration: Bringing together data from different data sources.
[0076] Beneficial effects: Improve the quality of data and provide a reliable data basis for subsequent analysis.
[0077] In the embodiment of the present invention, an enterprise data tag for the target enterprise can be determined, and a tag task can be determined based on the enterprise data tag;
[0078] Purpose: To define a series of labels for target enterprises to describe their different characteristics and conduct subsequent risk assessment based on these labels.
[0079] Example description:
[0080] Data labels: financial risk, operational risk, legal risk, reputation risk, etc.
[0081] Tag task: Based on each tag, design corresponding calculation rules or models to determine the enterprise's score on the tag.
[0082] Beneficial effect: Convert complex data problems into classification problems, facilitating subsequent risk assessment.
[0083] In the embodiment of the present invention, the risk value for the target enterprise can be calculated based on the label task;
[0084] Purpose: According to the defined labeling tasks, calculate the score of the target enterprise on each label, and combine these scores to calculate the comprehensive risk value of the enterprise.
[0085] Example description:
[0086] Risk value calculation: The risk value can be calculated using methods such as weighted average method and fuzzy comprehensive evaluation method.
[0087] Weight setting: The weights of different tags are set according to their importance.
[0088] Beneficial effects: Quantify the risk level of the enterprise and provide a basis for decision-making.
[0089] In the embodiment of the present invention, a risk assessment result for the target enterprise can be generated through the risk value;
[0090] Purpose: To convert the calculated risk value into an understandable risk assessment result.
[0091] Example description:
[0092] Risk level classification: Divide the risk value into different levels such as high, medium and low.
[0093] Risk report generation: Generate a detailed risk assessment report, including risk level, risk cause, risk recommendations, etc.
[0094] Benefits: Provide decision makers with clear and intuitive risk assessment results, helping them make more informed decisions.
[0095] The embodiment of the present invention obtains initial enterprise data for a target enterprise; performs preprocessing operations on the initial enterprise data to generate target enterprise data; determines an enterprise data label for the target enterprise, and determines a label task based on the enterprise data label; calculates a risk value for the target enterprise based on the label task; and generates a risk assessment result for the target enterprise through the risk value; thereby achieving the goal of converting the complex characteristics of an enterprise into quantifiable indicators by defining a series of labels, and designing corresponding calculation tasks for each label, thereby quantifying the enterprise characteristics, providing actionable indicators for risk assessment, and improving the efficiency of risk assessment for the enterprise.
[0096] Through the above steps, the embodiment of the present invention establishes a complete enterprise risk assessment system. The system can effectively identify the potential risks of the enterprise and provide timely warnings and suggestions to decision makers, thereby helping the enterprise to avoid risks and achieve sustainable development.
[0097] The advantages of this system are:
[0098] Systematic: Covering all aspects of enterprise risk assessment.
[0099] Scientificity: Based on data analysis and model calculation, it has high accuracy.
[0100] Practicality: It provides actionable risk assessment results and provides guidance for decision makers.
[0101] Based on the above embodiment, a variant embodiment of the above embodiment is proposed. It should be noted that in order to make the description concise, only the differences from the above embodiment are described in the variant embodiment.
[0102] In an optional embodiment of the present invention, the step of obtaining initial enterprise data for the target enterprise includes:
[0103] Determine the target data source;
[0104] Initial enterprise data is acquired from a plurality of different target data sources based on real-time data stream processing technology.
[0105] Exemplarily, the embodiment of the present invention may obtain the initial number of enterprises in the following manner.
[0106] Determination and integration of data sources:
[0107] Multi-source data integration: Establish a unified data warehouse to integrate data from different sources to form a complete enterprise data view.
[0108] Real-time and reliability of data sources: For data with high real-time requirements, such as public opinion data, real-time data stream processing technologies such as Kafka and Flink can be used.
[0109] Selection of data acquisition technology:
[0110] Crawler technology: For web page data, you can use Python's Scrapy framework to crawl.
[0111] API interface: For data sources that provide API interfaces, you can directly call the API to obtain data.
[0112] Database connection: For structured data, you can use database connectors to access it.
[0113] Data extraction tools: You can use ETL tools (such as Kettle and Talend) to extract and transform data.
[0114] Data cleaning and preprocessing:
[0115] Data cleaning: remove duplicate data, abnormal data, missing values, etc.
[0116] Data standardization: Unifying data from different sources into the same format and encoding.
[0117] Data transformation: Convert unstructured data into structured data for subsequent analysis.
[0118] The strategy of obtaining initial enterprise data from multiple different data sources has the following significant advantages:
[0119] 1. Improved data comprehensiveness and accuracy:
[0120] Multi-angle observation: By integrating information from different data sources, the target company can be comprehensively observed from multiple angles, reducing the bias or omissions that may exist in a single data source.
[0121] Data verification and complementation: Data from multiple data sources can be verified with each other to improve data reliability. If a data source is wrong or missing, other data sources can be used as a supplement to improve data integrity.
[0122] 2. Enhanced real-time performance:
[0123] Timely response: Based on real-time data stream processing technology, the latest enterprise data can be obtained in a timely manner, so as to monitor and analyze the enterprise status in real time and improve the timeliness of decision-making.
[0124] 3. Rich data diversity:
[0125] Multi-type data: Different data sources provide various data types, including structured data, unstructured data, etc., which can meet different types of analysis needs.
[0126] Deep insights: By integrating different types of data, we can dig out deeper corporate information and discover hidden relationships and patterns.
[0127] 4. Adapt to complex business scenarios:
[0128] Flexible response to changes: Different data sources can adapt to different business scenarios and needs and have strong flexibility.
[0129] Enhance system robustness: Multiple data sources can disperse risks and improve system stability.
[0130] 5. Support more complex analysis models:
[0131] Rich features: Multiple data sources provide rich features, which can build more complex analysis models and improve model accuracy.
[0132] Deeply explore value: By deeply mining multi-source data, new business opportunities and potential risks can be discovered.
[0133] 6. Improve decision-making support capabilities:
[0134] Data-driven decision-making: Based on comprehensive, accurate and real-time enterprise data, it can provide a more reliable basis for decision-making.
[0135] Optimize business processes: Through data analysis, you can optimize business processes and improve operational efficiency.
[0136] In an optional embodiment of the present invention, the step of performing a preprocessing operation on the initial enterprise data to generate target enterprise data includes:
[0137] Performing a data cleaning operation on the initial enterprise data, and obtaining text entities from the initial enterprise data after the data cleaning operation; the text entities include enterprise entities and personnel entities;
[0138] Generate a relationship graph in which the user expresses the association relationship between the enterprise entity and the person entity;
[0139] Determine the merged and fused dimension of the target enterprise, and determine the target attributes of the plurality of text entities through the relationship graph based on the merged and fused dimension;
[0140] A knowledge graph is constructed based on the target attributes, the relationship graph and the text entities, and the knowledge graph is determined as target enterprise data.
[0141] Optionally, before the step of constructing a knowledge graph based on the target attribute, the relationship graph and the text entity, and determining the knowledge graph as the target enterprise data, the step further includes:
[0142] Obtaining the occurrence duration, occurrence length, and data source information of the text entity;
[0143] Calculating the credibility of the target attribute based on the appearance duration, the appearance duration, and the data source information;
[0144] The credibility is used to delete repeated text entities and association relationships.
[0145] In a specific implementation, the embodiment of the present invention can generate target enterprise data in the following manner.
[0146] 1. Collect corporate data from different sources, such as business registration information, legal litigation information, public opinion data, etc.
[0147] The raw data is cleaned, deduplicated, and format converted to ensure data quality and consistency.
[0148] Data cleaning and preprocessing:
[0149] Data cleaning: In addition to deduplication, grid conversion, and normalization, you can also consider outlier detection and missing value filling.
[0150] Text data cleaning: Preprocess the text data by performing word segmentation, removing stop words, and part-of-speech tagging to prepare for subsequent entity recognition and relationship extraction.
[0151] Standard library construction: The establishment of a standard library helps to improve data consistency. It is recommended to establish a unified vocabulary and entity type library.
[0152] Extract key fields such as company name, unified social credit code, legal representative, etc.
[0153] 2. Text entity recognition and relationship extraction:
[0154] Corporate entity identification: Identify different corporate entities based on unique identifiers such as the unified social credit code.
[0155] Personnel entity identification: Identify individuals related to the enterprise, such as legal representatives, senior executives, etc.
[0156] Relationship extraction: Identify the relationships between enterprises and between enterprises and individuals, such as investment relationships, employment relationships, etc., to build a relationship map.
[0157] 3. Entity attribute fusion:
[0158] Determine the merge dimension: Select the attribute used to determine whether two entities are the same entity, such as the unified social credit code.
[0159] Attribute merging: Integrate the attributes of the same entity from different data sources to obtain a target attribute, and calculate the credibility of each target attribute.
[0160] Optionally, in the case of multiple values for an attribute, retain all values and record their origin and trustworthiness.
[0161] Select optimal value: Select the optimal value for each attribute based on factors such as credibility.
[0162] Uses of Credibility:
[0163] Solve the problem of data inconsistency: For the same entity, there may be different attribute values in different data sources. Through credibility calculation, the most reliable attribute value can be determined.
[0164] Improve the quality of knowledge graphs: By evaluating the credibility of entities and relationships, we can filter out noise data and improve the accuracy of knowledge graphs.
[0165] Support reasoning and decision-making: Information with high credibility can serve as a more reliable basis for subsequent reasoning and decision-making.
[0166] 4. Build a knowledge graph:
[0167] Create nodes: Use the identified business entities and person entities as nodes in the graph.
[0168] Create edges: Treat the relationships between entities as edges in the graph.
[0169] Add attributes: Add target attributes to nodes and edges, such as company name, registered capital, tenure time, etc.
[0170] Optionally, specific implementation methods of constructing a knowledge graph include:
[0171] 4.1 Detailed data extraction: Extract key fields from different data sources to form a detailed table.
[0172] 4.2 Traceability data consolidation:
[0173] Count the number of times, time, source, and other information that each attribute appears in different data sources.
[0174] The credibility of each attribute is calculated, and attributes with high credibility are more likely to represent the true attributes of the entity.
[0175] Handle multi-valued attributes, preserve all values and record provenance and trustworthiness.
[0176] 4.3 Object attribute analysis:
[0177] Select an optimal value for each attribute based on factors such as credibility.
[0178] If the credibility of multiple values is similar, a voting mechanism or expert judgment can be used.
[0179] Credibility calculation example:
[0180] Assume that there are three data sources, each providing the phone number of company A:
[0181] Data source 1: 133xxxx1234
[0182] Data source 2: 133xxxx4321
[0183] Data source 3: 133xxxx1234
[0184] The trustworthiness of each phone number can be calculated based on the following factors:
[0185] Number of occurrences: 133xxxx1234 appeared twice, and 133xxxx4321 appeared once.
[0186] Data source credibility: Assume that data source 1 and data source 2 have the same credibility, while data source 3 has a higher credibility.
[0187] Time dimension: If time information is available, the most recent value can be considered more credible.
[0188] Through calculation, it can be concluded that 133xxxx1234 is more credible, so it is used as the final phone number.
[0189] The specific calculation methods include:
[0190] The scoring function is a function of variables such as (number of occurrences, time of occurrence), and various forms of functions (linear, nonlinear; including time variables or only considering the number of occurrences) can be preset. This method uses a nonlinear function.
[0191] The calculation formula for the credibility score k of each source of attribute value is as follows:
[0192] ω1·log3(x1+1)+ω2·log5(x2+1)+ω3·x3.
[0193] Where ω1 represents the weight of the number of days, ω2 represents the weight of the number of times, and ω3 represents the weight of the data source. ω1+ω2+ω3=1. x1 represents the number of days, x2 represents the number of times, and x3 represents the coefficient of the data source.
[0194] The formula for calculating the scores of attribute values of multiple data sources is as follows:
[0195] kxd=∑k n
[0196] in,
[0197] kxd: represents the comprehensive credibility score of a certain attribute in multiple data sources.
[0198] ∑k n : represents the credibility score k of all data sources n Perform the summation.
[0199] k n : Indicates the credibility score of the nth data source for this attribute.
[0200] Through the above steps, we can integrate the enterprise information scattered in different data sources and build a complete enterprise knowledge graph. This knowledge graph can be used for various analyses and applications, such as risk assessment, market research, customer relationship management, etc.
[0201] For example, suppose you want to build a knowledge graph about technology companies.
[0202] 1. Data preparation and preprocessing:
[0203] Data sources: Data on technology companies are obtained from public information from the Industrial and Commercial Registration Bureau, corporate information query applications, Crunchbase and other platforms.
[0204] Data cleaning: remove duplicate data, erroneous data, and unify data formats.
[0205] Extract key fields: extract fields such as company name, founding time, financing round, investors, core products, and technology fields.
[0206] 2. Entity recognition and relationship extraction:
[0207] Business entity identification: Identify different technology companies based on company name and unified social credit code.
[0208] Personnel entity recognition: Identify key figures such as the company’s founder, CEO, CTO, etc.
[0209] Extract association relationships and create a relationship graph:
[0210] Investment Relationship: Identify the investment relationship between the investor and the investee company.
[0211] Competitive relationships: Identify competitors in the same field.
[0212] Partnerships: Identify partnerships between companies, such as technical collaboration, joint R&D, etc.
[0213] 3. Entity attribute fusion:
[0214] Merging dimension: Select the unified social credit code as the main merging dimension, while considering information such as company name and establishment time.
[0215] Attribute merging: Integrate information about the same company from different data sources. For example, merge financing round information from multiple data sources and calculate the credibility of each financing information.
[0216] Handling multi-valued attributes: If a company has multiple financing rounds, keep all financing round information and record the corresponding investors and time.
[0217] 4. Build a knowledge graph:
[0218] Create nodes: Identify companies and people as nodes in the graph.
[0219] Create edges: Use relationships such as investment, competition, and cooperation as edges in the graph.
[0220] Add attributes: Add attribute information to nodes and edges, such as company size, valuation, technology field, etc.
[0221] Finally, we may get a knowledge graph:
[0222] In this knowledge graph, the following queries can be performed:
[0223] Find competitors of a company: Find its competitors by looking up the competitive relationship with the company.
[0224] Analyze investment trends in a certain field: Understand investors’ preferences and investment hotspots by analyzing the financing situation of companies in this field.
[0225] Discover potential cooperation opportunities: Find potential partners by analyzing the cooperation relationships between companies.
[0226] In an optional embodiment of the present invention, the step of determining the enterprise data tag for the target enterprise includes:
[0227] Determine the input gate, forget gate, and output gate;
[0228] Generate a neural network deep learning model based on the input gate, the forget gate and the output gate;
[0229] The enterprise data label for the target enterprise is determined through the neural network deep learning model.
[0230] Exemplarily, the enterprise comprehensive risk assessment process based on the labeling system can be carried out in the following manner.
[0231] Data preparation and integration phase: Collect enterprise-related data from multiple data sources (standard libraries, archives, relational databases, etc.), and clean, deduplicate, and normalize them to ensure data quality and consistency.
[0232] Label system construction stage: Based on business needs and data characteristics, define a set of labels that can fully describe the characteristics of the enterprise to form a complete label system. For example, it can include dimensions such as financial status, operating status, legal risks, and public opinion risks.
[0233] Label algorithm library establishment phase: Develop algorithms for extracting features from target enterprise data and assigning labels. These algorithms include traditional rule matching, statistical analysis, and advanced machine learning algorithms. Optionally, sequence data analysis can be performed through LSTM (Long Short-Term Memory, neural network deep learning) models.
[0234] Among them, the LSTM theorem mainly includes the following three key components:
[0235] (1) Forget Gate: determines which information in the memory unit at the previous moment should be forgotten. Its calculation formula is:
[0236] f t =σ(W f ·[h t-1 , x t ]+b f )
[0237] Among them, f t is the output of the forget gate, σ is the sigmoid activation function, W f and b f are the weight matrix and bias term respectively, and [h_{t-1}, x_t] is the concatenated vector of the previous hidden state and the current input.
[0238] (2) Input Gate: determines which input information should be stored in the memory unit. It includes two parts:
[0239] Update candidate value calculation:
[0240] Input gate output: i t =σ(W i ·[h t-1 , x t ]+b t )
[0241] (3) Memory unit update and output gate (Output Gate), memory unit C t The update of is determined by the forget gate and the input gate:
[0242]
[0243] Among them, ⊙ represents element-wise multiplication; then, the final hidden state h is calculated through the output gate t .
[0244] o t =σ(W o ·[h t-1 , x t ]+b o )
[0245] h t =O t ⊙tanh(C t )
[0246] Output Gate O t Controls how the memory cell information affects the current hidden state.
[0247] Label task construction phase: define specific calculation logic and data sources for each label. For example, for the "limit high consumption" label, it is achieved by matching the enterprise's unified social credit code with the list of restricted high consumption.
[0248] Label calculation and output stage: Use the constructed algorithm model to analyze enterprise data and label each enterprise accordingly. Specifically, first pre-process the original data and extract features; then use the trained model to predict new data and obtain label results; finally, store the label results in the label library.
[0249] Scheduled task setting stage: According to the data update frequency and business needs, set the execution cycle of the label task, such as once a day, week, or month. Through the scheduled task, ensure the timely update of the label data.
[0250] In an optional embodiment of the present invention, the step of calculating the risk value for the target enterprise based on the label task includes:
[0251] Obtaining a task score value of the label task and an indicator weight for the label task;
[0252] The task score values are weightedly summed based on the indicator weights to calculate a risk value for the target enterprise.
[0253] In a specific implementation, the embodiment of the present invention can obtain the task score value of the tag task. For example, if there are n tags, the task score value of the nth tag task can be a n .
[0254] Determination of indicator weights: Different label tasks have different impacts on enterprise risks, and different weights should be set. The weights can be determined based on expert experience, data analysis, or machine learning methods.
[0255] Risk score calculation:
[0256] Weighted summation: The most common method is to multiply the score of each label by the corresponding weight and then sum them up.
[0257] Non-linear functions: For some risk factors, non-linear functions can be used to amplify or reduce their impact.
[0258] Risk level classification: Risk level classification can be adjusted based on business needs and historical data.
[0259] Optimization of risk models:
[0260] Introducing the time dimension: Consider the timeliness of different tags and give higher weight to recent events.
[0261] Consider interaction effects: There may be interactions between different labels, and the combined effects between labels need to be considered.
[0262] Introducing external information: External information, such as industry risks, macroeconomic conditions, etc., can be introduced to improve the risk assessment model.
[0263] Exemplarily, the risk value for the target enterprise may be calculated by the following formula.
[0264] fxd=∑an
[0265] fxd: risk score of the enterprise;
[0266] an: The score of the nth label.
[0267] Discussion on the formula fxd = ∑an:
[0268] This formula appears to be a simplified risk score calculation formula where:
[0269] fxd: Risk score for your business
[0270] an: the score of the nth label
[0271] In practical applications, different tags have different importance and should be given different weights.
[0272] The relationship between tags is not considered: Different tags may be associated with each other, and their combined effects need to be considered.
[0273] The time dimension is not taken into account: events occurring at different times have different impacts on risk.
[0274] The improved formula can be as follows:
[0275] Qiu d=∑(wn*sn)
[0276] in:
[0277] wn: weight of the nth label task;
[0278] sn: the score of the nth label task.
[0279] Going further, we can consider introducing nonlinear functions:
[0280] fxd=f(∑(wn*sn))
[0281] Wherein, f is a nonlinear function, such as sigmoid function, exponential function, etc.
[0282] The risk assessment results include visual display information and risk warning information.
[0283] Optionally, the risk assessment result includes visual display information and risk warning information.
[0284] Visual display information:
[0285] Enterprise portrait: Display the risk distribution of a single enterprise through visual charts (such as radar charts, word clouds, etc.), and intuitively present the risk characteristics of the enterprise.
[0286] Risk distribution: Use bar charts, pie charts, etc. to display the number of companies with different risk levels to understand the overall risk distribution.
[0287] Trend analysis: Use a line chart to show the changing trend of the enterprise risk score over time so as to detect risk changes in a timely manner.
[0288] Association analysis: Use a network diagram to show the associations between enterprises and the paths by which risks spread between enterprises.
[0289] 2. Risk warning information:
[0290] Threshold setting: Reasonably set thresholds for different risk levels based on business needs and risk tolerance.
[0291] Warning methods: In addition to emails, text messages, and system notifications, you can also consider instant messaging tools such as WeChat and DingTalk.
[0292] Warning content: The warning information should include detailed information such as the company name, risk level, and risk cause.
[0293] Warning frequency: Set different warning frequencies according to risk level and importance.
[0294] In order to enable those skilled in the art to better understand the embodiment of the present invention, an example is used below to illustrate the embodiment of the present invention.
[0295] refer to Figure 2 , Figure 2 It is a flow chart of a risk assessment method for an enterprise provided in an embodiment of the present invention;
[0296] S1. Data collection. Obtain basic data, including basic enterprise information, enterprise legal litigation information, business analysis, business status, intellectual property rights, listing information, and collect enterprise network public opinion data in real time.
[0297] S2. Data processing. Perform data cleaning on basic data, including data cleaning, deduplication, grid conversion, and normalization to form a standard library. Perform data extraction and integration management on data. Form enterprise and personnel portraits, build relationship maps, and analyze the relationship between personnel, personnel and enterprises, and enterprises.
[0298] S3. Label task construction. According to label rules, label tasks are constructed. Label tasks are debugged. According to business scenarios and data update conditions, the calculation cycle and triggering method of label tasks are set. Labels are output.
[0299] S4. Risk value calculation. Summarize and calculate the data in the tag library. Calculate the risk score of each enterprise and divide it into high, low and medium risk score ranges according to the threshold distribution.
[0300] S5. Visual display and risk warning. You can search for companies by tags to display their portraits; you can find the tags you have applied to them by searching for them; a visual chart will show the statistical distribution of companies that have won bids for each tag; when the risk score exceeds the preset threshold, the warning mechanism will be automatically triggered.
[0301] In the specific implementation, you can first build the enterprise's labeling system and determine the risk factors involved in the risk assessment of the enterprise.
[0302] Based on the risk factors and business scenarios involved in the risk assessment of the enterprise, a label system is constructed, including: basic label information, label value information, and label rule information.
[0303] Among them, each label provides an angle for describing and evaluating the enterprise (that is, the enterprise description and evaluation include multiple angles, each angle corresponds to a label; each label includes: basic label information, label value information, label rule information), and each label defines clear evaluation criteria and quantitative indicators (label value and label rule information) to ensure the accuracy and consistency of the evaluation.
[0304] For example, corporate debt ratio = total liabilities / total assets × 100%. The appropriate level of asset-liability ratio is 40% to 60%. If the corporate debt ratio exceeds the appropriate level, the corporate debt risk will be labeled as high. A comprehensive assessment of multiple labels can determine the risk level of the enterprise. A score is set for each label based on the risk impact on the enterprise.
[0305] In addition, the risk level of an enterprise can be determined through a comprehensive assessment of multiple labels, and a score can be set for each label based on the risk impact on the enterprise.
[0306] S1. Collect basic data.
[0307] Obtain basic data of the enterprise, including: basic information of the enterprise, legal information of the enterprise, business analysis, business status, intellectual property rights, listing information, public opinion collection and other data.
[0308] Among them, basic data can be obtained from a variety of sources, such as web pages, databases, documents, social media, etc. Here, data types include structured data (such as databases), semi-structured data (such as XML, JSON files) and unstructured data (such as text, images).
[0309] S2. Process the collected data (data cleaning, text data recognition, graph construction, entity attribute fusion, and credibility score calculation).
[0310] 1. Perform data cleaning on basic data, including data cleaning, deduplication, grid conversion, and normalization to form a standard library for data extraction and fusion processing;
[0311] 2. Use natural language processing technology to identify entities in text data, such as names of people, places, organizations, etc. Match the identified entities with existing entities in the knowledge base. If no match is found, you may need to create a new entity entry.
[0312] 3. First, based on the processed data above, form enterprise portraits and personnel portraits; then, by analyzing the relationships between personnel, personnel and enterprises, and enterprises and enterprises, build a relationship map; finally, extract, integrate, analyze and merge entities and relationships of the data to form a knowledge map.
[0313] 4. Determine the dimensions for merging and integrating each enterprise entity, and integrate the attributes of the entity (the extraction and integration process of personnel entities and the relationships between personnel and enterprises, and between enterprises is the same as that of enterprise entities).
[0314] Here, the corporate entity refers to the enterprise's unified social credit code as the enterprise's unique identifier, and the enterprise's basic information, corporate legal litigation information, business analysis, operating conditions, intellectual property rights, listing information, collected public opinion and other data are the enterprise's attribute information.
[0315] The following steps are used to determine the merged dimensions for each business entity:
[0316] 4.1. Detailed data extraction. Through data extraction, data from different sources are extracted, and relevant fields such as unified social credit code and basic enterprise information are extracted into corresponding fields of the defined enterprise archive table to form a detailed table;
[0317] 4.2. Traceability data merging. First, count the number of times each attribute data appears, the number of days, and the credibility of the data source (the credibility of the data source is generally determined based on the way the data is obtained and the credibility of the expert experience); then, calculate the credibility score of each attribute of the entity based on the source agreement of the element, and merge the attributes. If the attribute extracted from a certain source has multiple values, the field is saved as a multi-valued form. For example, the unit phone field: 133xxxx1234{source 1, number of times, earliest time, latest time}, 133xxxx4321{source 1, number of times, earliest time, latest time}.
[0318] 4.3. Object attribute analysis. Each field retains the most likely value according to a certain strategy and all source information. For example, the unit phone field: 133xxxx1234 (maximum weighted number of times). The "source field" of the entire object data is multi-valued, and all sources involved in the analysis of the data (multi-valued); the "number field" of the entire object data is single-valued, and the sum of the number of times of all sources involved in the analysis of the data (multi-valued).
[0319] 5. Calculate the credibility score of each attribute of the entity based on the number of data occurrences, days, and credibility of the data source. This can solve the inconsistency of the same entity in different data sources, remove duplicate entities and relationships, and keep the knowledge graph refined.
[0320] The specific calculation methods include:
[0321] The scoring function is a function of variables such as (number of occurrences, time of occurrence), and various forms of functions (linear, nonlinear; including time variables or only considering the number of occurrences) can be preset. This method uses a nonlinear function.
[0322] The calculation formula for the credibility score k of each source of attribute value is as follows:
[0323] ω1·log3(x1+1)+ω2·log5(x2+1)+ω3·x3.
[0324] Where ω1 represents the weight of the number of days, ω2 represents the weight of the number of times, and ω3 represents the weight of the data source. ω1+ω2+ω3=1. x1 represents the number of days, x2 represents the number of times, and x3 represents the coefficient of the data source.
[0325] The formula for calculating the scores of attribute values of multiple data sources is as follows:
[0326] kxd=∑k n
[0327] in,
[0328] kxd: represents the comprehensive credibility score of a certain attribute in multiple data sources.
[0329] ∑k n : represents the credibility score k of all data sources n Perform the summation.
[0330] k n : Indicates the credibility score of the nth data source for this attribute.
[0331] S3. Establish a relevant labeling algorithm library based on the processed standard library data, archival data and relational data.
[0332] Construct enterprise labeling tasks. Among them, the subject of each label is the enterprise, the labeling results are output to the label library, and the data is uniquely identified as the unique identifier of the data in the enterprise subject library, that is, the enterprise unified social credit code. The enterprise labeling task is constructed through the labeling system.
[0333] For example, in constructing a label to restrict high-consumption enterprises,
[0334] First, create a new data source connection and configure the enterprise database connection information to the data source management;
[0335] Then, enter an operator in the label task, drag the data source of the high-consumption restriction table into the task canvas, analyze the high-consumption restriction data in the enterprise legal litigation information from the enterprise data, and extract the unified social credit code of all enterprises with high-consumption restrictions;
[0336] Finally, it is fully matched with the unified social credit code in the enterprise entity information table. If it matches, the enterprise entity will be labeled with a high-consumption restriction label and the label data will be output to the label library.
[0337] In addition, machine learning algorithms are used to evaluate public opinion data and legal litigation data, including a company’s credit rating, market competition, potential risks, and product quality.
[0338] 1. Use natural language processing (NLP) technology: pre-process online public opinion data to remove meaningless symbols and stop words, and perform stem extraction or word form restoration;
[0339] 2. Apply BERT and Word2Vec language models for embedding and convert text data into numerical features;
[0340] 3. Build a deep learning-based classification model RNN to determine sentiment tendency. Specifically, build an LSTM-based RNN structure, where the input layer receives word embedding vectors and captures the sentiment information of the text sequence through multi-layer LSTM units; at the last time step, use the fully connected layer to map the final hidden state to the sentiment category space (such as positive, negative, and neutral), and use the softmax function to output the probability distribution of each category; the LSTM theorem mainly includes the following three key components:
[0341] (1) Forget Gate: determines which information in the memory unit at the previous moment should be forgotten. Its calculation formula is:
[0342] f t =σ(W f ·[h t-1 , x t ]+b f ).
[0343] Among them, f t is the output of the forget gate, σ is the sigmoid activation function, W f and b f are the weight matrix and bias term respectively, and [h_{t-1}, x_t] is the concatenated vector of the previous hidden state and the current input.
[0344] (2) Input Gate: determines which input information should be stored in the memory unit. It includes two parts:
[0345] Update candidate value calculation:
[0346] Input gate output: i t =σ(W i ·[h t-1 , x t ]+b t )
[0347] (3) Memory unit update and output gate (Output Gate), memory unit C t The update of is determined by the forget gate and the input gate:
[0348]
[0349] Among them, ⊙ represents element-wise multiplication; then, the final hidden state h is calculated through the output gate t .
[0350] o t =σ(W o ·[h t-1 , xt ]+b o )
[0351] h t =o t ⊙tanh(C t )
[0352] Output gate o t Controls how the memory cell information affects the current hidden state.
[0353] 4. Use LDA algorithm to identify topics.
[0354] 5. According to the business scenario and data update situation, set the calculation cycle and triggering method of the label task, and perform scheduled and real-time label calculation for the label task.
[0355] Among them, the calculation cycle setting: the purpose of the calculation cycle is to determine the time point for the label task job to be executed so that the data can be processed at a predetermined time interval. The data acquisition cycle is set to every day; the trigger mode setting: the label task related to the basic information of the enterprise is set as a scheduled task, which is executed every night without affecting the efficiency of business data query during the day. The data source is set as a real-time task for public opinion monitoring tasks.
[0356] S4. Calculation of risk value.
[0357] Summarize and calculate the data in the tag library. Calculate the risk score of each enterprise and divide it into high, low and medium risk score ranges according to the threshold distribution.
[0358] Set an indicator score a for each tag n The scores of all the tags for each enterprise are added up and calculated. The total score is the enterprise's risk score.
[0359] fxd=∑a n
[0360] S5. Visual display and risk assessment warning.
[0361] The results are displayed through search query and chart analysis; different warning thresholds are set according to the risk score. When the risk score exceeds the threshold, the warning mechanism is automatically triggered. The warning information is pushed to the relevant personnel of the market supervision department in a timely manner through email, SMS, system notification, etc., so that regulatory measures can be taken in a timely manner.
[0362] Through the above methods, a complete labeling system is built, combined with advanced technologies such as machine learning, to achieve a comprehensive, intelligent and timely assessment of corporate risks, providing strong support for market supervision.
[0363] The specific manifestations are as follows:
[0364] Comprehensiveness:
[0365] Multi-dimensional assessment: By constructing corporate portraits, personnel portraits and knowledge graphs, we conduct a comprehensive assessment of the company from multiple angles, covering the company's financial status, operating conditions, legal risks, reputation risks and other aspects.
[0366] Fine-grained labeling: Through a fine-grained labeling system, enterprise risks can be classified and described more accurately, thereby improving the accuracy of risk assessment.
[0367] Timeliness:
[0368] Real-time monitoring: By setting up real-time tag calculation tasks, changes in corporate risks can be discovered in a timely manner to avoid the expansion of risks.
[0369] Early warning mechanism: Establish a risk early warning mechanism that can issue timely warnings before risks occur, providing regulators with sufficient response time.
[0370] accuracy:
[0371] Machine learning algorithms: Advanced machine learning algorithms, such as BERT, Word2Vec, RNN, etc., are used to analyze and mine massive data, improving the accuracy of risk assessment.
[0372] Label model: By building a reasonable label model, the risk characteristics of an enterprise can be identified more accurately.
[0373] Intelligent:
[0374] Automation: By automating labeling tasks, manual intervention is reduced and work efficiency is improved.
[0375] Intelligence: Using machine learning algorithms, intelligent risk assessment is achieved, which can detect potential risks from massive data.
[0376] Overall advantages include:
[0377] Improve regulatory efficiency: Through automation and intelligent means, the efficiency of market supervision has been improved and the cost of supervision has been reduced.
[0378] Improve regulatory accuracy: Through multi-dimensional and fine-grained risk assessment, the targetedness and effectiveness of supervision are improved.
[0379] Reduce risks: By timely discovering and responding to potential risks, the company's operating risks are reduced and market order is maintained.
[0380] Providing support for decision-making: It provides data support for the decision-making of regulatory authorities and helps to formulate more scientific and effective regulatory policies.
[0381] The embodiments of the present invention realize intelligent and automated assessment of enterprise risks by constructing a complete labeling system and applying advanced machine learning technology, providing a powerful tool for market supervision and having important theoretical significance and application value.
[0382] Construction of the labeling system: Combining enterprise portraits, personnel portraits and knowledge graphs to build a more comprehensive and fine-grained labeling system.
[0383] Application of machine learning algorithms: The use of advanced machine learning algorithms improves the accuracy and efficiency of risk assessment.
[0384] Establishment of a risk early warning mechanism: timely identification and response to potential risks, providing regulatory authorities with sufficient time to respond.
[0385] It should be noted that, for the sake of simplicity, the method embodiments are described as a series of action combinations, but those skilled in the art should be aware that the embodiments of the present invention are not limited by the order of the actions described, because according to the embodiments of the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present invention.
[0386] Reference Figure 3 , shows a structural block diagram of a risk assessment device for an enterprise provided in an embodiment of the present invention, which may specifically include the following modules:
[0387] An initial enterprise data acquisition module 301 is used to acquire initial enterprise data for a target enterprise;
[0388] The target enterprise data generating module 302 is used to perform a preprocessing operation on the initial enterprise data to generate target enterprise data;
[0389] A label task determination module 303 is used to determine an enterprise data label for the target enterprise, and determine a label task based on the enterprise data label;
[0390] A risk value calculation module 304 is used to calculate the risk value for the target enterprise based on the label task;
[0391] The risk assessment result generating module 305 is used to generate a risk assessment result for the target enterprise according to the risk value.
[0392] As for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0393] In addition, an embodiment of the present invention further provides an electronic device, such as Figure 4 As shown, it includes a processor 401, a communication interface 402, a memory 403 and a communication bus 404, wherein the processor 401, the communication interface 402, and the memory 403 communicate with each other through the communication bus 404.
[0394] Memory 403, used for storing computer programs;
[0395] The processor 401 is used to implement the risk assessment method for an enterprise described in any one of the above embodiments when executing the program stored in the memory 403:
[0396] The communication bus mentioned in the above terminal can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, only one thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0397] The communication interface is used for communication between the above terminal and other devices.
[0398] The memory may include a random access memory (RAM) or a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.
[0399] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0400] like Figure 5As shown, in another embodiment provided by the present invention, a computer-readable storage medium 501 is also provided, in which instructions are stored. When the computer-readable storage medium 501 is run on a computer, the computer executes the risk assessment method for the enterprise described in the above embodiment.
[0401] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the enlightenment of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are within the protection of the present invention.
[0402] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in the embodiments of the present invention can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0403] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0404] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0405] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0406] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0407] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical disks.
[0408] The above is only a specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. Therefore, the protection scope of the present invention should be based on the protection scope of the claims.
Claims
1. A risk assessment method for an enterprise, characterized in that: include: Obtaining initial enterprise data for target enterprises; Performing a preprocessing operation on the initial enterprise data to generate target enterprise data; Determine an enterprise data tag for the target enterprise, and determine a tag task based on the enterprise data tag; Calculate the risk value for the target enterprise based on the label task; A risk assessment result for the target enterprise is generated using the risk value.
2. The method according to claim 1, characterized in that: The step of obtaining initial enterprise data for the target enterprise includes: Determine the target data source; Initial enterprise data is acquired from a plurality of different target data sources based on real-time data stream processing technology.
3. The method according to claim 2, characterized in that The step of performing a preprocessing operation on the initial enterprise data to generate target enterprise data comprises: Performing a data cleaning operation on the initial enterprise data, and obtaining text entities from the initial enterprise data after the data cleaning operation; the text entities include enterprise entities and personnel entities; Generate a relationship graph in which the user expresses the association relationship between the enterprise entity and the person entity; Determine the merged and fused dimension of the target enterprise, and determine the target attributes of the plurality of text entities through the relationship graph based on the merged and fused dimension; A knowledge graph is constructed based on the target attributes, the relationship graph and the text entities, and the knowledge graph is determined as target enterprise data.
4. The method according to claim 3, characterized in that Before the step of constructing a knowledge graph based on the target attribute, the relationship graph and the text entity, and determining the knowledge graph as the target enterprise data, the step further includes: Obtaining the occurrence duration, occurrence length, and data source information of the text entity; The credibility of the target attribute is calculated based on the appearance duration, the appearance duration and the data source information; the credibility is used to delete repeated text entities and the association relationships.
5. The method according to claim 4, characterized in that The step of determining the enterprise data tag for the target enterprise comprises: Determine the input gate, forget gate, and output gate; Generate a neural network deep learning model based on the input gate, the forget gate and the output gate; The enterprise data label for the target enterprise is determined through the neural network deep learning model.
6. The method according to claim 1, characterized in that The step of calculating the risk value for the target enterprise based on the label task includes: Obtaining a task score value of the label task and an indicator weight for the label task; The task score values are weightedly summed based on the indicator weights to calculate a risk value for the target enterprise.
7. The method according to claim 1, characterized in that The risk assessment results include visual display information and risk warning information.
8. A risk assessment device for an enterprise, characterized in that: include: An initial enterprise data acquisition module is used to acquire initial enterprise data for a target enterprise; A target enterprise data generating module, used for performing a preprocessing operation on the initial enterprise data to generate target enterprise data; A label task determination module, used to determine an enterprise data label for the target enterprise, and determine a label task based on the enterprise data label; A risk value calculation module, used to calculate the risk value for the target enterprise based on the label task; The risk assessment result generating module is used to generate a risk assessment result for the target enterprise according to the risk value.
9. An electronic device, characterized in that: It includes a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; The memory is used to store computer programs; The processor is used to implement the method according to any one of claims 1 to 7 when executing the program stored in the memory.
10. A computer-readable storage medium having instructions stored thereon, which, when executed by one or more processors, cause the processors to perform the method according to any one of claims 1 to 7.