Risk assessment method and device, nonvolatile storage medium and electronic equipment
By using pre-trained language models and deep learning networks to extract features from textual and structured data of micro and small enterprises and constructing knowledge graphs, the problem of single data dimensions and insufficient capture of related risks in pre-loan risk assessment of micro and small enterprises is solved, and comprehensive and accurate risk assessment is achieved.
Patent Information
- Application Number
- CN202510854273.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-24
- Publication Date
- 2025-10-28
AI Technical Summary
Pre-loan risk assessment for small and micro enterprises relies on single-dimensional structured financial data, resulting in incomplete risk assessment, lack of capture of correlated risks, and poor model interpretability.
Pre-trained language models and deep learning networks are used to extract features from text data, construct a knowledge graph, and combine a pre-trained graph network and a risk assessment model to fuse text feature vectors and node embedding feature vectors to generate risk scores for loan companies.
It enables a comprehensive and accurate assessment of pre-loan risks for micro and small enterprises, captures the risk transmission effect between enterprises, and improves the interpretability and accuracy of the assessment model.
Smart Images

Figure CN120852034A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a risk assessment method and apparatus, a non-volatile storage medium, and an electronic device. Background Technology
[0002] In the fintech sector, pre-loan risk assessment for micro, small, and medium-sized enterprises (MSEs) is a crucial step in the healthy development of financial services and support for the real economy. However, relevant assessment methods primarily rely on structured financial data, such as financial statements and credit ratings. This approach has significant limitations, especially when dealing with MSEs. MSEs typically lack standardized financial reporting, and information regarding their business operations, market environment, and cash flow often exists in unstructured text formats (such as corporate credit reports, operating logs, and judicial announcements), which are traditionally difficult to utilize effectively. Furthermore, the complex relationships between MSEs, such as upstream and downstream supply chains, mutual guarantees, and cross-shareholdings, can significantly impact their credit risk, but relevant assessment systems often neglect the dynamic analysis of these interconnected information.
[0003] Specifically, the problems with related technologies include: 1. Limited data dimensions: Over-reliance on financial indicators, while the financial data of micro and small enterprises is incomplete and non-standardized, making it difficult to fully reflect their true risk status. This results in limited input information for the assessment model, failing to capture the full picture of enterprise operations. 2. Lack of correlation analysis: For complex inter-enterprise relationships, such as guarantee chains and investment networks, existing technologies lack effective analytical methods, making it difficult to quantify the impact of these relationships on the credit risk of micro and small enterprises, thus failing to accurately predict the risk transmission effect between enterprises. 3. Poor model interpretability: Although some assessment methods based on statistical models or simple machine learning algorithms can predict risk, their decision-making processes and the importance of their features are often difficult to understand and interpret. This not only affects financial institutions' trust in the models but also limits their application in actual risk management and decision-making.
[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention
[0005] This application provides a risk assessment method and apparatus, a non-volatile storage medium, and an electronic device to at least solve the technical problems of incomplete risk assessment and insufficient capture of associated risks caused by the reliance on single-dimensional structured financial data in relevant pre-loan risk assessment methods for micro and small enterprises.
[0006] According to one aspect of this application, a risk assessment method is provided, comprising: acquiring a first dataset, wherein the first dataset includes data related to loan enterprises; performing word segmentation on the text data in the first dataset to obtain a word sequence, and processing the word sequence using a pre-trained language model to obtain an output sequence, wherein the output sequence includes an embedding representation of each word in the word sequence; performing feature extraction processing on the output sequence using a pre-trained text convolutional neural network and a pre-trained bidirectional long short-term memory network to obtain a text feature vector; determining the knowledge graph corresponding to the loan enterprise based on the first dataset, wherein the knowledge graph includes at least: triple data, node attribute data, and edge attribute data; processing the triple data, node attribute data, and edge attribute data using a pre-trained graph network to obtain a node embedding feature vector including risk information; and processing the text feature vector and the node embedding feature vector using a pre-trained risk assessment model to obtain a risk score for the loan enterprise.
[0007] Optionally, the word sequence is processed using a pre-trained language model to obtain an output sequence, including: adding classification markers and delimiter markers to the beginning and end of the word sequence to obtain a tokenized sequence, and mapping the tokenized sequence to the corresponding symbolized sequence; the symbolized sequence is processed using a pre-trained language model to obtain an output sequence, wherein the output sequence includes at least: a character vector matrix, and the character vector matrix includes: the hidden layer dimension corresponding to each symbol in the symbolized sequence.
[0008] Optionally, a pre-trained text convolutional neural network and a pre-trained bidirectional long short-term memory network are used to perform feature extraction processing on the output sequence to obtain a text feature vector, including: extracting the first vector corresponding to the label position of the classification label in the word vector matrix, and determining the first vector as the text global feature vector; using multiple convolutional kernels of different sizes in the pre-trained text convolutional neural network to extract local semantic features in the word vector matrix to obtain the first feature vector; using the pre-trained bidirectional long short-term memory network to capture long-distance dependencies in the word vector matrix to obtain multiple temporal features, and performing mean pooling processing on the multiple temporal features to obtain the second feature vector; concatenating the text global feature vector, the first feature vector, and the second feature vector to obtain the text feature vector.
[0009] Optionally, a pre-trained risk assessment model is used to process the text feature vector and node embedding feature vector to obtain the risk score of the loan enterprise. This includes: imputing missing values and normalizing the tabular data in the first dataset to obtain the target tabular data; determining the month-on-month revenue growth rate within m months based on the revenue of the current month and the revenue of the previous m months, and determining the time-series features based on the month-on-month revenue growth rate within m months, where m is a positive integer; determining the interest coverage ratio based on the profit before interest and taxes and the interest expense, and determining the ratio features based on the interest coverage ratio; determining the derived features based on the time-series features and the ratio features; extracting features from the target tabular data and adding derived features to the feature extraction results to obtain the tabular features; and processing the text feature vector, node embedding feature vector, and tabular features using the pre-trained risk assessment model to obtain the risk score of the loan enterprise.
[0010] Optionally, a pre-trained risk assessment model is used to process text feature vectors, node embedding feature vectors, and table features to obtain a risk score for the lending company. This includes: processing table features using a fully connected layer to obtain a warm-start node embedding vector; processing the warm-start node embedding vector using a pre-trained graph network to obtain a graph structure feature vector; fusing the graph structure feature vector with the text feature vector to obtain a first target feature vector; fusing the first target feature vector with the node embedding feature vector to obtain a second target feature vector; and processing the second target feature vector using the pre-trained risk assessment model to obtain a risk score for the lending company.
[0011] Optionally, the language model and text convolutional neural network are trained using the following method: Evaluation data is acquired and divided into training and testing sets; the text data in the training set is segmented to obtain a first word sequence, and the language model is used to process the first word sequence to obtain a first output sequence; the language model and text convolutional neural network are trained using the first output sequence. The text convolutional neural network and bidirectional long short-term memory network are trained using the following method: the text convolutional neural network and bidirectional long short-term memory network are trained using the first word sequence to obtain a first text vector output by the text convolutional neural network and bidirectional long short-term memory network; the first output sequence is merged into the first text vector as an intermediate feature to obtain first feature data; the corresponding label of the first feature data is used as the first target variable, and the first target variable is combined with historical target variables to fine-tune the parameters of the text convolutional neural network to obtain the parameters of the first target variable, wherein the historical target variable is the risk category label already labeled in the text data of the training set, used to supervise model training.
[0012] Optionally, the graph network is trained using the following method: A first knowledge graph is determined based on the training set; the graph network is used to process the first triplet data, first node attribute data, and first edge attribute data in the first knowledge graph to obtain a first node embedding feature vector including risk information; the first feature data and the first node embedding feature vector are fused to obtain second feature data; the corresponding label of the second feature data is used as a second target variable, and the second target variable is combined with historical target variables to fine-tune the parameters of the graph network, obtaining the second target variable parameters; the risk assessment model is trained using the following method: the risk assessment model is used to process the first text vector and the first node embedding feature vector to obtain a first output result; the second feature data and the first output result are fused to obtain third feature data; the corresponding label of the third feature data is used as a third target variable, and the third target variable is combined with historical target variables to fine-tune the parameters of the risk assessment model, obtaining the third target variable parameters.
[0013] According to another aspect of this application, a risk assessment apparatus is also provided, comprising: an acquisition module for acquiring a first dataset, wherein the first dataset includes data related to loan enterprises; a first processing module for performing word segmentation on the text data in the first dataset to obtain a word sequence, and processing the word sequence using a pre-trained language model to obtain an output sequence, wherein the output sequence includes the embedding representation of each word in the word sequence; a second processing module for performing feature extraction processing on the output sequence using a pre-trained text convolutional neural network and a pre-trained bidirectional long short-term memory network to obtain a text feature vector; a determination module for determining the knowledge graph corresponding to the loan enterprise based on the first dataset, wherein the knowledge graph includes at least: triple data, node attribute data, and edge attribute data; a third processing module for processing the triple data, node attribute data, and edge attribute data using a pre-trained graph network to obtain a node embedding feature vector including risk information; and a generation module for processing the text feature vector and the node embedding feature vector using a pre-trained risk assessment model to obtain a risk score for the loan enterprise.
[0014] According to another aspect of this application, a non-volatile storage medium is also provided, the storage medium including a stored program, wherein the program, when running, controls the device where the storage medium is located to execute the above-mentioned risk assessment method.
[0015] According to another aspect of this application, an electronic device is also provided, comprising: a memory and a processor, the processor being configured to run a program stored in the memory, wherein the program executes the risk assessment method described above during runtime.
[0016] According to another aspect of this application, a computer program is also provided, wherein the computer program, when executed by a processor, implements the above-described risk assessment method.
[0017] According to another aspect of this application, a computer program product is also provided, comprising a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium stores a computer program, which, when executed by a processor, implements the above-described risk assessment method.
[0018] In this application, a method is employed to obtain a first dataset, which includes data related to loan enterprises; the text data in the first dataset is segmented to obtain word sequences, and a pre-trained language model is used to process the word sequences to obtain an output sequence, wherein the output sequence includes the embedding representation of each word in the word sequence; a pre-trained text convolutional neural network and a pre-trained bidirectional long short-term memory network are used to extract features from the output sequence to obtain text feature vectors; based on the first dataset, a knowledge graph corresponding to the loan enterprise is determined, wherein the knowledge graph includes at least: triple data, node attribute data, and edge attribute data; a pre-trained graph network is used to process the triple data, node attribute data, and edge attribute data to obtain node embedding feature vectors including risk information; a pre-trained risk assessment model is used to process the text feature vectors and node embedding feature vectors to obtain a risk score for the loan enterprise. This method achieves a comprehensive and accurate technical effect for pre-loan risk assessment, thereby solving the technical problem of incomplete risk assessment and insufficient capture of associated risks caused by the reliance on single-dimensional structured financial data in relevant micro and small enterprise pre-loan risk assessment methods. Attached Figure Description
[0019] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0020] Figure 1 This is a flowchart of a risk assessment method according to an embodiment of this application;
[0021] Figure 2 This is a flowchart of another risk assessment method according to an embodiment of this application;
[0022] Figure 3 This is a flowchart of a hierarchical framework for model training according to an embodiment of this application;
[0023] Figure 4 This is a structural diagram of a risk assessment device according to an embodiment of this application;
[0024] Figure 5 This is a hardware structure block diagram of a computer terminal for a risk assessment method according to an embodiment of this application. Detailed Implementation
[0025] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.
[0026] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0027] According to an embodiment of this application, a method embodiment for risk assessment is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0028] Figure 1 This is a flowchart of a risk assessment method according to an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps:
[0029] Step S101: Obtain the first dataset, which includes data related to loan companies.
[0030] In step S101 above, the first dataset contains all data related to the loan company. This data can be structured, such as financial statements and tax records, or it can be unstructured, such as textual descriptions in corporate credit reports and summaries of judicial announcements.
[0031] Step S101 involves comprehensively collecting data related to loan recipients to provide a foundation for subsequent analysis and risk assessment. By integrating structured and unstructured data, a more comprehensive corporate profile can be constructed, going beyond just a financial description.
[0032] Step S102: The text data in the first dataset is segmented to obtain a word sequence, and the word sequence is processed using a pre-trained language model to obtain an output sequence, wherein the output sequence includes the embedding representation of each word in the word sequence.
[0033] In step S102 above, the text data in the first dataset is first processed through word segmentation. For Chinese text, the jieba word segmentation tool is used, combined with a custom financial terminology dictionary, to ensure the correct segmentation of financial professional terms. For English text, it is processed according to standard English word segmentation rules. After word segmentation, a pre-trained language model (such as BERT) is used to transform the word sequence into an embedding representation, i.e., to generate an embedding vector for each word. The BERT model can understand the meaning of each word based on context, which helps to capture semantic information in the text. For example, when processing a corporate credit report, key phrases in the report such as "deteriorating financial situation" and "multiple legal disputes" will be transformed into embedding vectors. These vectors not only contain the semantic information of the words themselves, but also reflect their meaning in a specific context.
[0034] Step S102 transforms unstructured text data into structured word sequences through word segmentation, and then uses a pre-trained language model to convert each word into an embedding vector, which can improve the analyzability of text data and the depth of machine understanding.
[0035] Step S103: Use a pre-trained text convolutional neural network and a pre-trained bidirectional long short-term memory network to perform feature extraction processing on the output sequence to obtain the text feature vector.
[0036] In step S103 above, a pre-trained Text Convolutional Neural Network (TextCNN) and a Bidirectional Long Short-Term Memory Network (BiLSTM) are used to extract features from the output sequence generated in step S102. TextCNN captures local and global semantic features through convolutional kernels of different sizes, while BiLSTM can effectively handle temporal dependencies in text. The combination of the two can comprehensively extract semantic information and time-series features from text data. These features are finally integrated into a text feature vector for subsequent analysis. For example, for a report that describes a company's operating status in detail, TextCNN captures positive semantic features such as "revenue growth" and "effective cost control," while BiLSTM focuses on time-series changes describing the company's financial health, such as revenue fluctuations over several consecutive months.
[0037] Step S103 utilizes pre-trained TextCNN and BiLSTM to perform deep feature extraction on the output sequence processed by the language model, aiming to capture deeper semantic features and temporal dependency information from the text. The combination of the two models enables the model to effectively identify and quantify risk signals in text data, ultimately obtaining a text feature vector that comprehensively reflects the risk of the text.
[0038] Step S104: Based on the first dataset, determine the knowledge graph corresponding to the loan enterprise, wherein the knowledge graph includes at least: triple data, node attribute data and edge attribute data.
[0039] In step S104 above, a knowledge graph containing loan companies is constructed based on the enterprise relationship information in the first dataset. The knowledge graph consists of nodes (representing entities, such as enterprises, individuals, and industries) and edges (representing relationships between entities, such as investment and guarantees). Edges also carry attribute data (such as transaction amount and guarantee period). The attribute data of nodes and edges further enriches the information content of the knowledge graph. For example, node attributes may include enterprise age and credit rating, while edge attributes may include transaction frequency and guarantee amount.
[0040] Step S104 constructs a knowledge graph based on enterprise-related data, which can represent the complex relationships and attribute information between enterprises in the form of a graph structure, facilitating subsequent correlation risk analysis. By defining the attributes of nodes and edges, the knowledge graph can more comprehensively reflect the correlation of enterprise risks, including but not limited to upstream and downstream relationships in the supply chain, guarantee chain relationships, shareholder relationships, etc., which are implicit correlation risks that are difficult to explore in depth using traditional methods.
[0041] Step S105: Process the triplet data, node attribute data, and edge attribute data using a pre-trained graph network to obtain node embedding feature vectors that include risk information.
[0042] In step S105 above, a pre-trained graph network (such as GraphSAGE) is used to extract features from the knowledge graph constructed in step S104. The graph network enhances the feature representation of the target node by aggregating information from neighboring nodes. The node embedding feature vector not only reflects the node's own attributes but also considers its position in the graph and the attribute information of the edges connected to it, thus obtaining a node embedding feature vector containing risk information.
[0043] Step S105 utilizes a pre-trained graph network model to analyze the knowledge graph, which can quantify the transmission effect of risks among enterprises. By learning the attribute information of nodes and edges, node embedding feature vectors that reflect the risk status of enterprises are obtained. These vectors not only contain the risk information of individual enterprises but also consider their position and risk transmission path in the complex relationship network, thereby more accurately assessing the true risk level of enterprises.
[0044] Step S106: Use the pre-trained risk assessment model to process the text feature vector and node embedding feature vector to obtain the risk score of the loan enterprise.
[0045] In step S106 above, a pre-trained risk assessment model is used to fuse the text feature vector generated in step S103 and the node embedding feature vector generated in step S105 to generate a risk score for the lending company. The risk assessment model can be a deep neural network. This model integrates data from different modalities, captures complex risk factors through multi-layer nonlinear transformations, and ultimately outputs a numerical risk score that intuitively reflects the credit risk level of the lending company.
[0046] In step S106, the pre-trained risk assessment model receives the text feature vectors and node embedding feature vectors from steps S103 and S105. It uses multi-hidden-layer nonlinear transformations in deep learning to capture the complex interactions between risk factors. Step S106 integrates the previously extracted feature information to generate a comprehensive risk score, which is directly used for loan decisions. The risk score not only considers the company's own financial and operational status but also incorporates quantitative analysis of inter-company risk correlations, providing a more comprehensive and multi-dimensional perspective on risk assessment.
[0047] In summary, steps S101 to S106 address the core pain points in pre-loan risk assessment for micro and small enterprises, such as data fragmentation, difficulty in capturing associated risks, and insufficient model interpretability. By integrating knowledge graphs and multimodal deep learning models, a comprehensive assessment method covering data preprocessing, hierarchical modeling, feature fusion, and scorecard generation is constructed, achieving a three-dimensional risk control approach from "static qualification analysis of a single enterprise" to "dynamic risk extrapolation across the entire network."
[0048] The following are Figure 1 The steps shown are illustrated and explained by way of example.
[0049] According to some optional embodiments of this application, the word sequence is processed using a pre-trained language model to obtain an output sequence, which can be achieved by adding a classification label and a delimiter label to the beginning and end of the word sequence respectively to obtain a labeled sequence, and mapping the labeled sequence to the corresponding symbolic sequence; the symbolic sequence is processed using a pre-trained language model to obtain an output sequence, wherein the output sequence includes at least: a character vector matrix, and the character vector matrix includes: the hidden layer dimension corresponding to each symbol in the symbolic sequence.
[0050] In this embodiment, the classification marker ([CLS]) is placed at the beginning of the text sequence to indicate the start of the entire text. The pre-trained language model pays special attention to the [CLS] marker when processing text because its output is used to represent the global features of the entire text sequence. The delimiter marker ([SEP]) is used to separate different text sequences. When processing a single text sequence, the [SEP] marker is placed at the end of the sequence to indicate the end of the text. A tokenized sequence refers to a text sequence with specific markers ([CLS] and [SEP]) added before and after the original text sequence; it is one of the standard formats input to a pre-trained language model. Specifically, the word sequence is preprocessed by inserting the "classification marker" [CLS] at the beginning of the word sequence and the "delimiter marker" [SEP] at the end of the word sequence, resulting in a tokenized word sequence. For example, the word sequence "financial situation deteriorated and equity frozen" becomes "[CLS] financial situation deteriorated and equity frozen [SEP]" after tokenization.
[0051] The process of converting words in text into numerical form is called tokenization. Each word corresponds to a unique numerical ID, and these numerical IDs constitute a symbolic sequence. Mapping refers to converting a sequence of words into a corresponding sequence of numerical IDs, which can be achieved through the vocabulary of a pre-trained language model. When processing text, the pre-trained language model maps each symbol (word) to a fixed-length vector, and the set of these vectors is called the word vector matrix. The word vector matrix can reflect the semantic information of words and their meaning in a specific context.
[0052] This process involves mapping the tokenized sequence to its corresponding symbolic sequence. Specifically, it uses the vocabulary built into the pre-trained language model to convert each word in the sequence into its corresponding ID. For example, the word "financial situation" might be mapped to the numeric ID "1234," while "deterioration" might be mapped to the ID "56." In this way, the entire word sequence is transformed into a series of IDs, forming a symbolic sequence.
[0053] A pre-trained language model is used to process the symbolized sequence to obtain the output sequence, which includes a word vector matrix. The pre-trained model contains multiple hidden layers. Each word (or symbol) is transformed into a vector containing the dimensions of the hidden layers through these layers, and these vectors constitute the word vector matrix. For example, if the hidden layer dimension of the model is set to 768 (a common setting for BERT models), then each word will be converted into a vector of length 768, and the final word vector matrix will contain the semantic representations of all words in the symbolized sequence.
[0054] Furthermore, feature extraction processing of the output sequence is performed using a pre-trained text convolutional neural network and a pre-trained bidirectional long short-term memory network to obtain text feature vectors. This can be achieved through the following method: extract the first vector corresponding to the label position of the classification label in the word vector matrix and determine the first vector as the global feature vector of the text; extract local semantic features in the word vector matrix using multiple convolutional kernels of different sizes in the pre-trained text convolutional neural network to obtain the first feature vector; capture long-distance dependencies in the word vector matrix using the pre-trained bidirectional long short-term memory network to obtain multiple temporal features, and perform mean pooling processing on the multiple temporal features to obtain the second feature vector; concatenate the global feature vector of the text, the first feature vector, and the second feature vector to obtain the text feature vector.
[0055] In the above steps, when processing the relevant text data for pre-loan risk assessment of micro and small enterprises, the word vector matrix obtained by using the pre-trained language model is the key to analyzing the implicit risk information in the text.
[0056] First, the first vector corresponding to the classification label ([CLS]) at the beginning of the tokenized sequence is extracted from the word vector matrix and defined as the text global feature vector. This vector condenses the main information of the entire text, especially those contents that can reflect the overall risk situation of the enterprise, such as the description of the operating status and the financial health status.
[0057] Next, a pre-trained text convolutional neural network is used, employing multiple built-in convolutional kernels of varying sizes to slide across the word vector matrix, thereby capturing local semantic features in the text. These kernels can recognize keyword phrases of varying lengths, such as "declining profits" and "increased litigation risk," which are important clues for assessing corporate risk. By scanning the word vector matrix with the convolutional kernels, a series of local feature representations can be extracted. Then, using max pooling or average pooling methods, these local features are integrated to form a first feature vector, which represents the sum of all local semantic features in the text.
[0058] Simultaneously, a pre-trained bidirectional long short-term memory (BiLSTM) network captures long-distance dependencies between words in the word vector matrix, a step crucial for understanding the overall semantic structure of the text. BiLSTM not only considers the order of words but also captures connections between distant words, such as changes in financial status described at the beginning and end of the year. After processing the word vector matrix, BiLSTM generates a series of temporal features reflecting the temporal dependencies and sequential relationships of words in the text. To integrate these temporal features, they are averaged and pooled to obtain a second feature vector. This second feature vector integrates the long-term and short-term semantic dependencies of words in the text, reflecting the temporal structural characteristics of the text.
[0059] Finally, the previously obtained global text feature vector, first feature vector (i.e., local semantic feature vector), and second feature vector (i.e., temporal dependency feature vector) are concatenated to form a complete text feature vector. This vector contains global information, local semantics, and temporal dependencies of the text, making it an ideal input for subsequent risk assessment models and helping them to fully understand the risk signals implicit in the text.
[0060] According to some alternative embodiments of this application, the risk score of the lending company is obtained by processing the text feature vector and node embedding feature vector using a pre-trained risk assessment model. This can be achieved through the following methods: Missing value imputation and normalization are performed on the tabular data in the first dataset to obtain target tabular data; the month-on-month revenue growth rate within m months is determined based on the revenue of the current month and the revenue of the previous m months, and the time-series feature is determined based on the month-on-month revenue growth rate within m months, where m is a positive integer; the interest coverage ratio is determined based on the profit before interest and taxes and the interest expense, and the ratio feature is determined based on the interest coverage ratio; derived features are determined based on the time-series feature and the ratio feature; features are extracted from the target tabular data, and derived features are added to the feature extraction results to obtain tabular features; the risk score of the lending company is obtained by processing the text feature vector, node embedding feature vector, and tabular features using the pre-trained risk assessment model.
[0061] In this embodiment, for the tabular data in the first dataset, data preprocessing techniques are first used to impute missing values and normalize the data to obtain the target tabular data. Impute missing values uses a machine learning model (such as XGBoost) to predict and fill in missing financial indicators, ensuring data integrity. Normalization scales numerical features to the same scale, preventing extreme values from having an excessive impact on model training and ensuring that all indicators can compete fairly in the model. The resulting target tabular data, such as the debt-to-equity ratio, cash flow indicators, and revenue volatility, are standardized, eliminating the influence of unit of measurement.
[0062] Subsequently, based on the revenue data for the current month and the revenue data for the previous m months (where m is a positive integer, e.g., 3 months), the month-on-month revenue growth rate over those m months is calculated to reveal the recent trend in the company's operating revenue, helping to understand the stability and growth potential of the company's operations. For example, if a company's revenue for the past three months is 1 million yuan, 1.1 million yuan, and 1.2 million yuan respectively, then the calculated month-on-month growth rates are 10% and 9.09%, which will be added to the target data as part of the time-series characteristics. Simultaneously, the interest coverage ratio is calculated based on the company's earnings before interest and taxes (EBIT) and interest expenses. A higher EBIT and lower interest expenses result in a higher interest coverage ratio, and vice versa; this value reflects the company's ability to bear financial debt.
[0063] Next, based on the calculated time-series and ratio characteristics, further derived characteristics are determined. For example, a new characteristic, the "Financial Health Index," can be constructed, which combines the interest coverage ratio and the quarter-on-quarter revenue growth rate to provide a more comprehensive evaluation of the health of a company's financial condition.
[0064] Secondly, feature extraction is performed on the processed target table data to capture risk patterns in the financial data. The previously calculated derived features are added to the feature extraction results to obtain a complete table feature vector. This step ensures that the model can simultaneously consider structured financial data and advanced risk indicators derived from that data.
[0065] Finally, text feature vectors (capturing semantic risks from corporate credit reports and operational logs), node embedding feature vectors (quantifying risk transmission through inter-corporate investments, guarantees, and other relationships), and table features (integrating standardized financial indicators and derived risk features) are input into a pre-trained risk assessment model. This risk assessment model is a multi-hidden-layer neural network capable of comprehensively analyzing these three types of features, capturing their complex interactions, and ultimately outputting a numerical risk score that intuitively reflects the credit risk level of the lending company.
[0066] Furthermore, the risk score of the lending company is obtained by processing the text feature vector, node embedding feature vector, and table features using a pre-trained risk assessment model. This can be achieved as follows: The table features are processed using a fully connected layer to obtain a warm-start node embedding vector; the warm-start node embedding vector is processed using a pre-trained graph network to obtain a graph structure feature vector; the graph structure feature vector is fused with the text feature vector to obtain a first target feature vector; the first target feature vector is fused with the node embedding feature vector to obtain a second target feature vector; and the second target feature vector is processed using the pre-trained risk assessment model to obtain the risk score of the lending company.
[0067] In the above steps, the preprocessed tabular features are first processed using a fully connected layer. The fully connected layer can convert structured information such as financial indicators and operating data into hot-start node embedding vectors with high-dimensional spatial representation. This conversion process maps standardized financial data to a feature space that the pre-trained graph network can understand, thereby achieving the fusion of graph features and financial information.
[0068] Then, a pre-trained graph network model is used to process the hot-start node embedding vector. Graph networks can capture and quantify the impact of complex relationships between enterprises on risk, such as upstream and downstream relationships in the supply chain and guarantee chains. Through graph network processing, a graph structure feature vector is obtained, which contains information about the risks associated with enterprises, such as the degree of impact of neighboring enterprises' defaults on the target enterprise.
[0069] Next, the graph structure feature vector is fused with the previously extracted text feature vector to obtain the first target feature vector. The fused first target feature vector more comprehensively reflects the enterprise's risk status, including both the enterprise's textual description risk and the impact of inter-enterprise correlation risks.
[0070] Next, the first target feature vector is further fused with the node embedding feature vector to obtain the second target feature vector. The node embedding feature vector is constructed based on a knowledge graph and includes not only the node's own attribute risks, such as the company's establishment year, paid-in capital, and shareholder risk scores, but also the edge features between the company and external entities, such as transaction frequency and guarantee amount. The generation of the second target feature vector ensures that the model can simultaneously consider the company's internal financial and operational status, the implicit relationship risks between companies, and the company's adaptability to the environment, forming a comprehensive risk representation that encompasses internal and external risk information.
[0071] Finally, using a pre-trained risk assessment model, the second target feature vector is taken as input, and the interaction of complex risk factors is captured through multi-layer nonlinear transformation, ultimately outputting a risk score for the loan enterprise. This risk score, based on multimodal data fusion and deep learning model training, can not only effectively assess the independent financial risk of enterprises but also capture the transmission effect of inter-enterprise related risks, thus providing a more accurate and comprehensive pre-loan risk assessment.
[0072] In some optional embodiments of this application, the language model and the text convolutional neural network are trained by the following method: obtaining evaluation data and dividing the evaluation data into a training set and a test set; performing word segmentation on the text data in the training set to obtain a first word sequence, and processing the first word sequence using the language model to obtain a first output sequence; using the first output sequence to train the language model and the text convolutional neural network respectively; the text convolutional neural network and the bidirectional long short-term memory network are trained by the following method: using the first word sequence to train the text convolutional neural network and the bidirectional long short-term memory network respectively to obtain a first text vector output by the text convolutional neural network and the bidirectional long short-term memory network; merging the first output sequence as an intermediate feature into the first text vector to obtain first feature data; using the corresponding label of the first feature data as a first target variable, and combining the first target variable with historical target variables to fine-tune the parameters of the text convolutional neural network to obtain the first target variable parameters, wherein the historical target variables are the risk category labels already labeled in the text data in the training set, used to supervise model training.
[0073] In this embodiment, a batch of data is first obtained from the assessment dataset, which covers various textual descriptions and financial information required for pre-loan risk assessment of micro and small enterprises. This assessment data is carefully divided into a training set and a test set. The training set is used for training and optimizing the model, while the test set is used to evaluate the model's generalization ability and performance.
[0074] For the text data in the training set, word segmentation is used for preprocessing to divide the text into meaningful word units, forming the first word sequence. For example, for the sentence describing the business situation of a company, "Company A recently encountered a financial crisis, and shareholder B was listed as a dishonest executor," after word segmentation, the resulting first word sequence might be "Company A," "recently," "encountered," "financial crisis," "shareholder B," "listed as," and "dishonest executor." Subsequently, the first word sequence is input into a pre-trained language model for deep processing, outputting a first output sequence with rich semantic information, namely a character vector matrix. These character vectors not only contain the meaning of the words themselves but also capture the potential semantic features of the words in the context.
[0075] Once the initial output sequence is obtained, training the language model and the text convolutional neural network can begin. First, the language model is trained. By comparing the initial output sequence with the actual text, the model gradually learns how to more accurately understand and represent the semantic information in the text. Next, the text convolutional neural network is trained. Again, using the initial output sequence and its semantic features, the neural network is trained to recognize and extract local semantic information, such as key phrases describing specific risks. This ultimately yields the first text vector, representing a concentrated representation of the local semantic features in the text.
[0076] Secondly, the training processes of the text convolutional neural network and the bidirectional long short-term memory network are complementary. First, the first word sequence is fed into the incompletely trained text convolutional neural network and bidirectional long short-term memory network for preliminary feature extraction, resulting in the first text vector. The first output sequence is then merged with the first text vector as an intermediate feature to ensure that the model can not only capture local semantic features but also understand the overall semantic structure of the text. The merged feature data (i.e., the first feature data) is combined with corresponding risk category labels, which serve as the first target variable to guide model training. The first target variable is then combined with historical target variables (i.e., labeled risk category labels) from the training set to form a richer supervisory signal. Based on these supervisory signals, the parameters of the text convolutional neural network are fine-tuned, and the model's weights are continuously iterated and optimized through backpropagation to improve the accuracy of the model in predicting text risk categories.
[0077] Taking the training of a text convolutional neural network for assessing pre-loan risk of micro and small enterprises as an example, after obtaining the training set and the first word sequence after corresponding word segmentation, the BERT model is first used to generate the first output sequence. Then, the first output sequence is combined with the output of the text convolutional neural network to obtain the first feature data. Each piece of text data in the training set is assigned a risk category label, such as "low risk," "medium risk," and "high risk," which constitute the first target variable. In the early stages of training, it may be found that the model's prediction results deviate from the true risk category. At this time, the first target variable is compared with historical target variables, and this deviation is quantified by calculating a loss function (such as cross-entropy loss). This deviation is then used to guide the fine-tuning of the model's parameters until the model's performance on the training set reaches the expected standard. This yields a text convolutional neural network model and its parameter settings that can effectively identify and quantify risk information in text.
[0078] Through the above training process, not only can the performance of language models and text convolutional neural networks be improved when processing pre-loan risk description texts for micro and small enterprises, but a model that can gain a deeper understanding of fine-grained risk information in the text can also be built, laying a solid foundation for subsequent multi-level risk assessment.
[0079] Furthermore, the graph network is trained using the following method: A first knowledge graph is determined based on the training set; the graph network is used to process the first triplet data, first node attribute data, and first edge attribute data in the first knowledge graph to obtain a first node embedding feature vector including risk information; the first feature data and the first node embedding feature vector are fused to obtain second feature data; the corresponding label of the second feature data is used as a second target variable, and the second target variable is combined with historical target variables to fine-tune the parameters of the graph network, obtaining the second target variable parameters; the risk assessment model is trained using the following method: the risk assessment model is used to process the first text vector and the first node embedding feature vector to obtain a first output result; the second feature data and the first output result are fused to obtain third feature data; the corresponding label of the third feature data is used as a third target variable, and the third target variable is combined with historical target variables to fine-tune the parameters of the risk assessment model, obtaining the third target variable parameters.
[0080] In the above steps, a first knowledge graph is constructed based on the training set data. The first knowledge graph not only includes enterprise entities and their attributes, but also details the relationships between enterprises, such as upstream and downstream supply chains, guarantee relationships, and shareholder structures. It is represented in the form of triples (entity-relationship-attribute). The first triple data is the specific instance of these relationships, such as "enterprise A-holding-enterprise B" and "enterprise C-guarantee-enterprise D, guarantee amount of 5 million". The first node attribute data and the first edge attribute data record in detail the characteristics of each node (such as the number of years the enterprise has been established and the paid-in capital) and the characteristics of the edge (such as the duration of cooperation and the transaction amount).
[0081] Next, a graph network model is used to perform in-depth processing on this information in the first knowledge graph. The graph network model can effectively analyze the complex relationships between nodes. By iteratively updating the embedding features of nodes, the first node embedding feature vector, which includes risk information, is finally obtained. These vectors not only reflect the risk characteristics of the nodes themselves, but also incorporate the potential risks of the relationships between the nodes and other nodes. For example, the node embedding feature vector of company A includes not only company A's credit score and financial indicators, but also comprehensively considers the credit status of companies B and C, which have a guarantee relationship with it.
[0082] Then, the first feature data (i.e., data incorporating text features) is fused with the embedded feature vector of the first node to obtain the second feature data. This fusion operation ensures that the model can simultaneously consider the risk information described in the text and the risks associated with inter-enterprise relationships, forming a more comprehensive risk feature representation. The corresponding labels of the second feature data (i.e., risk category labels) are used as the second target variable and combined with the historical target variable (i.e., the risk category labels already labeled in the training set) to fine-tune the parameters of the graph network model, thereby optimizing the model's ability to predict associated risks. Finally, the parameters of the second target variable are obtained, enabling the graph network to more accurately quantify the transmission effect of risks associated with inter-enterprise relationships.
[0083] Furthermore, the first text vector and the first node embedding feature vector are processed using a risk assessment model to obtain the first output result. The second feature data is then fused with the first output result to further integrate text risk features, node embedding features, and inter-enterprise relationship risks, resulting in the third feature data. Next, the corresponding labels of the third feature data (also risk category labels) are used as the third target variable. Combined with historical target variables, the model parameters are fine-tuned to optimize the model's ability to process comprehensive risk features, thereby obtaining the parameters for the third target variable. This ensures that the risk assessment model can accurately predict the pre-loan risk level of micro and small enterprises.
[0084] Figure 2 This is a flowchart of another risk assessment method according to an embodiment of this application, such as... Figure 2 As shown, risk assessment methods can also be implemented through the following steps:
[0085] First, step S1 is executed to divide the evaluation data into training and test sets and perform preprocessing. Preprocessing includes normalization of text data and tabular data to obtain training and test set data containing text data and tabular data.
[0086] Next, step S3 is performed to segment the text data. If it is not English text, Chinese words and word order information are obtained using Chinese word segmentation. If it is English text, English words and word order information are obtained using English word segmentation. Then, the word sequence is input into the BERT network to obtain each word and its corresponding word vector, thereby obtaining the character-level features of the text data.
[0087] Next, step S2 is performed, using the text data processed in step S3 to train the BERT network and the TextCNN network respectively. The loss functions for both are:
[0088]
[0089] Where S is the total number of samples in the training set, and y is the true classification label. A score vector or probability vector for category prediction.
[0090] In step S4, a convolutional neural network model (using ReLU as the activation function) and a bidirectional LSTM network are used to train the word sequences in the preprocessed text data to obtain the character-level features of the word sequences.
[0091] In step S5, the character-level features of the text data are merged into the original feature data of the subtask as intermediate features to generate feature 1 data. At the same time, the corresponding label of feature 1 data is used as the target variable of the classification task. The TextCNN network is fine-tuned by combining the historical target variables to obtain the target variable parameters.
[0092] Among them, the original feature data of the above sub-tasks are the outputs after processing by BERT+CNN / BiLSTM; the historical target variable is the risk category label already labeled in the text data, which is used to supervise the training of the model; the target variable parameters refer to the model parameters optimized through training in the TextCNN network, which are used to characterize the mapping relationship between text features and target variables (risk categories), including the weight matrix (such as convolution kernel parameters, fully connected layer parameters) and bias terms of the TextCNN network.
[0093] Further, we move on to the knowledge graph-related steps. First, we generate the K-graph for subtask 1 based on historical data. Historical data refers to structured, unstructured, and relational data related to micro and small enterprises accumulated in financial institutions or external data sources, including basic enterprise information, transaction records, risk events, etc.
[0094] The node data in the K-graph includes node type, weight, and node label data. The graph structure includes triples and attribute data. Node weights are then randomly initialized, and the root mean square propagation method is used to learn the weight values. Positive and negative samples are randomly selected. The loss function for subtask 1 is:
[0095]
[0096] Where u(x,y) represents the positive edge of (x,y), v(x,y) represents the negative edge of (x,y), W represents the weight, and σ represents the activation function.
[0097] After obtaining the weights of subtask 1, input them into the corresponding TextCNN network and graph network (step S6), and then execute step S7, which specifically includes:
[0098] The feature 1 data is used as input to the graph network to obtain the output. The output data of subtask 1 and subtask 2 are used with the loss function:
[0099]
[0100] Where y is the actual category label. is the predicted classification label, and N is the number of training set samples.
[0101] Next, proceed to step S8, where the output data of step S7 is merged into the original features of subtask 1 to generate feature 2 data. At the same time, the corresponding label of feature 2 data is used as the target variable of subtask 2. Combined with the historical target variable, the parameters of the feedforward neural network are fine-tuned to obtain the target variable parameters. Then, proceed to step S9.
[0102] The data of feature 2 is used as the input of the model of subtask 2 to obtain the output data, which is used as the original feature of subtask 3. At the same time, the corresponding label of the data of feature 3 is used as the target variable of subtask 3. The parameters of the feedforward neural network are fine-tuned in combination with the historical target variable to obtain the target variable parameters. Finally, step S10 is executed to merge the data of feature 3 as the original feature of subtask 3 into the original feature of subtask 3 to complete the pre-loan risk scoring card.
[0103] In summary, it can be understood that the input for subtask 1 includes: raw text data such as enterprise credit reports (PDF / plain text), business registration change records, and court announcement summaries. Word sequences are input into the BERT network to generate a word vector matrix (containing CLS global features). Character-level and temporal features are extracted using CNN+BiLSTM to generate a text feature vector (e.g., 1168 dimensions). The baseline model is BERT+TextCNN+fully connected layers. The training process includes: Step S2: Training the BERT and TextCNN networks using the preprocessed text data, with cross-entropy as the loss function, and optimizing the network weights. Step S5: Fusing the text character-level features with the original features of the subtask into feature 1 data, and fine-tuning the parameters of TextCNN to output the text risk classification result (label probability). The output includes: feature 1 data, i.e., the fused text feature vector, which contains semantic risk information.
[0104] For subtask 2, the input includes: original knowledge graph data, node embedding vectors, and graph structure, and the output is node embedding feature vectors. In step S8, the output of subtask 2 (node embedding features) is fused with feature 1 data (textual semantic features) from subtask 1 to generate cross-modal feature 2, which is used for graph structure modeling.
[0105] For subtask 3, it is formed by fusing the feature 1 data of subtask 1 (text classification) and the node embedding features of subtask 2 (graph network risk assessment). It includes text semantic features and graph structure associated risk features, integrates multimodal features and captures the interaction of complex risk factors through multi-layer neural networks, and outputs the final risk score.
[0106] In the above steps, the model training adopts a hierarchical framework of "data preprocessing → unimodal modeling → multimodal feature fusion → joint training" as shown in Figure 3 , and the specific process is as follows:
[0107] Input step. Multimodal raw dataset (structured / unstructured / relational data). Among them, the multimodal raw dataset includes: 1. Text data (N×T dimension, N is the number of samples, T is the text length): enterprise credit reports (PDF / plain text), industrial and commercial change records, court announcement summaries; unstructured fields: natural language descriptions such as "business anomaly" and "equity freeze". 2. Tabular data (N×F dimension, F is the number of features): financial indicators: asset-liability ratio (%), current ratio, revenue volatility in the past 12 months; behavioral data: tax credit rating (A / B / C level), invoice amount stability. 3. Knowledge graph data (E×3 dimension, E is the number of triples): triple examples: (Enterprise A, holds, Enterprise B), (Enterprise C, guarantees, Enterprise D, guarantee amount 5 million); node attributes: enterprise establishment years, paid-in capital, shareholder risk score; edge attributes: cooperation duration (months), transaction frequency (times / year).
[0108] Text data preprocessing step.
[0109] Step 1: Chinese and English word segmentation. Chinese: Use the jieba word segmentation tool, combined with a custom dictionary (such as financial terms), to generate a word list (such as "enterprise / operation / anomaly"); English: Based on the NLTK library, segment by spaces and convert to lowercase (such as "credit / freeze"); Output: word sequence w1, w2,..., w T , with word order information attached.
[0110] Step 2: BERT word vector generation. Input: The word sequence is converted into an ID sequence through the BERT tokenizer, and special tokens (CLS / SEP) are added; Output: word vector matrix H∈RT×d h , where, R is a real number matrix, T is the sequence length, d h is the hidden dimension, and the CLS vector in the word vector matrix is used as the global text feature h cls .
[0111] Step 3: CNN+BiLSTM feature enhancement. CNN layer: Use 3 types of convolutional kernels (window sizes 3 / 7 / 9) to extract local semantic features, and output where, d c is the feature dimension output by each convolutional kernel; BiLSTM layer: Capture long-range dependencies and output temporal features where, d l is the dimension of the hidden state. Take average pooling to obtain Imean Text feature fusion: splicing Obtain the text feature vector (d text =768 + 3 * 100 + 100 = 1168 (assuming the output dimension of CNN / BiLSTM is 100).
[0112] Tabular data preprocessing steps.
[0113] Step 1: Handling Missing Values. Numerical fields (e.g., revenue): use XGBoost regression model prediction to fill missing values; categorical fields (e.g., tax class): fill with the mode or map to one-hot encoding.
[0114] Step 2: Normalization and Scaling. Continuous Features: Z-Score Standardization Classification features: Label coding, such as tax level A→3, B→2, C→1.
[0115] Step 3: Derivative Feature Construction. Time-Series Feature: Month-on-Month Revenue Growth Rate (3 Months) = (Current Month Revenue - Revenue 3 Months Ago) / Revenue 3 Months Ago; Ratio Feature: Interest Coverage Ratio = Earnings Before Interest and Taxes / Interest Expense. Output of Tabular Data Processing: Standardized Table Features. (d table =F + derived feature number).
[0116] Steps for constructing a knowledge graph.
[0117] Step 1: Entity Relationship Extraction. Tools: Extracting triples based on rule engines (such as regular expressions matching keywords like "holding" and "guarantee") + Distant Supervision; for example: extracting (Enterprise A, holding, Enterprise B, 60% shareholding) from "Enterprise A holds 60% of Enterprise B's equity".
[0118] Step 2: Graph structure definition. Node types: Enterprise (E), Legal entity (P), Industry (I); Edge types: Investment (E→E), Employment (P→E), Belong to (E→I).
[0119] Step 3: Node initialization embedding. Cold start node: Random initialization vector. (d g =200); Hot start node: using table feature f table e is mapped through the fully connected layer. i =W emb ·f table +b emb Where IR is the real number field, e i W is the embedding vector.emb Let b be the weight matrix of the fully connected layer. emb This is the bias vector for the fully connected layer.
[0120] Single-modal modeling steps. Design a dedicated model to extract core risk features based on the characteristics of different data modalities.
[0121] Text classification model: semantic risk identification. Architecture: BERT + TextCNN + fully connected layers; Input: f text (From the preprocessing stage); Output: Probability of text risk labels (C represents the number of categories, such as 3 categories: low / medium / high risk text).
[0122] Financial risk modeling: Architecture: Single hidden layer neural network (comparison with baseline model); Input: f table Output: Financial Risk Score fin ∈R(0-1 interval, the higher the value, the weaker the debt repayment ability);
[0123] Feature splicing and dimensionality reduction:
[0124] Step 1: Original Feature Fusion: Stitching Among them, h i This represents the feature vectors related to the knowledge graph. Example: If d text =1168, d g =200, d table =50, then the fusion feature dimension is 1418.
[0125] Step 2: Layered dimensionality reduction: First layer fully connected: f1 = ReLU(W3·f fusion +b3)∈R 512 Second fully connected layer: f2 = ReLU(W4·f1 + b4) ∈ R 256 Output: Compact feature vector f compact ∈R 256 Where f1 and f2 are the feature vectors processed by the first and second fully connected layers, respectively.
[0126] Joint training steps (end-to-end optimization): Optimize all-link parameters using a unified loss function to balance classification and regression objectives. Binary classification task → Regression task (risk scoring) → Weighted total loss. Training configuration and tuning: Optimizer: RMSprop (learning rate 1e-3, momentum 0.9); Batch size: 64 (can be increased to 128 if memory allows); Training epochs: 50 epochs (early stopping strategy: terminate if validation set loss does not decrease for 3 consecutive epochs). Finally, generate a scorecard.
[0127] The above steps transform heterogeneous data—text, tables, and graphs—into standardized features (such as BERT word vectors and graph node embeddings) through data preprocessing. Semantic features, financial indicator features, and associated risk features are extracted using unimodal modeling. Cross-modal feature fusion is then achieved through hierarchical dimensionality reduction and attention mechanisms. Finally, joint training optimizes the multi-task loss function to generate a risk scoring card that can map to traditional scoring ranges. This framework not only breaks through the reliance of traditional risk control on structured data but also achieves dynamic capture of associated risks through the combination of knowledge graphs and deep learning. Furthermore, interpretable design and engineering adaptation improve the model's efficiency for business deployment, providing a fully intelligent solution for pre-loan risk assessment of micro and small enterprises, from data processing to model deployment.
[0128] It should be noted that the relevant technologies mainly rely on structured financial data, making it difficult to effectively utilize the large amount of unstructured text (such as credit reports and business logs) and implicit relationships (such as guarantee chains and investment networks) that exist in micro and small enterprises, resulting in a single dimension of risk assessment.
[0129] This application innovatively constructs a multimodal data processing system of "text + table + graph": by combining the BERT pre-trained model with CNN / BiLSTM network, it realizes the extraction of semantic features of Chinese word sequences after word segmentation (such as identifying high-risk semantics from the text "continuous losses and frozen equity"), and uses the TextCNN network to capture character-level features of English text, which solves the problem that traditional methods such as TF-IDF cannot deeply understand semantic dependencies.
[0130] Meanwhile, based on enterprise business registration data and transaction records, a knowledge graph containing triples (entity-relationship-attribute) is constructed. This transforms inter-enterprise relationships such as investment and guarantees into graph-structured data. Through quantitative modeling of node attributes (such as credit scores) and edge weights (such as transaction amounts), it fills the gap in the analysis of "relationship risk" in traditional methods. This combination of technologies enables the model to cover more than 90% of the risk information for micro and small enterprises, and it can still effectively assess enterprises with missing financial data through text and graph data.
[0131] Many related technologies employ a single model to process mixed data, making it difficult to account for the differences in characteristics of different modalities, resulting in insufficient risk feature extraction. This application designs a three-tiered modeling framework of "text classification → risk assessment → risk prediction," with customized models for each sub-task: The text classification task uses BERT as the baseline model, combined with a feature extraction module and a fully connected layer, to accurately identify risk categories in the text (such as "normal operation" and "abnormal warning"); the risk assessment task constructs a single-hidden-layer graph network model based on a knowledge graph, optimizes node weights through the root mean square propagation algorithm, and quantifies the risk transmission effect between enterprises (such as the impact of neighboring enterprises' defaults on the target enterprise); the risk prediction task utilizes a multi-hidden-layer neural network to integrate the text semantic features, graph structure features, and traditional financial indicators output from previous sub-tasks, capturing the interaction of complex risk factors through multi-layer nonlinear transformations (such as the risk superposition of highly indebted enterprises in dense guarantee chains). Furthermore, through a layer-by-layer merging mechanism of Feature 1 (text feature fusion), Feature 2 (graph feature fusion), and Feature 3 (full feature integration), a progressive abstraction from low-dimensional single-modal features to high-dimensional comprehensive risk is achieved, which improves the model's prediction accuracy (AUC-ROC) by 10%-15% compared to traditional Logistic regression, effectively distinguishing the score distribution of "high-risk enterprises" from "normal operating enterprises".
[0132] Current technologies for analyzing inter-enterprise risk rely heavily on manual rules (such as counting the number of guarantees or shareholder overlap), making it difficult to uncover hidden risk transmission paths and leading to frequent problems like multiple borrowing and guarantee chain risk clusters. This application constructs an enterprise relationship network using a knowledge graph, upgrading the traditional two-dimensional "enterprise-indicator" analysis to a three-dimensional "entity-relationship-attribute" model. First, a K-graph containing node types, weights, and labels is generated based on historical data. By randomly initializing node weights and using the root mean square propagation algorithm to learn the ability to distinguish between positive and negative edges, the graph network can automatically identify high-risk relationship patterns (e.g., when two of a company's three direct guarantors have loan defaults, the model judges the company's associated risk by aggregating the features of adjacent nodes). Second, by dynamically updating the triples in the knowledge graph (e.g., real-time capture of business registration changes and judicial litigation data), the graph network parameters are fine-tuned, enabling real-time monitoring and prediction of relationship risks. This solves the lag problem of traditional methods relying on periodic manual due diligence. Experiments show that this technology can improve the accuracy of risk identification in guarantee chains by 40% and effectively reduce the default rate caused by related relationships.
[0133] The following are the advantages of this application compared to related technologies.
[0134]
[0135]
[0136] Figure 4 This is a structural diagram of a risk assessment device according to an embodiment of this application, such as... Figure 4 As shown, the device includes:
[0137] The acquisition module 41 is used to acquire a first dataset, which includes data related to loan companies.
[0138] The first processing module 42 is used to perform word segmentation on the text data in the first dataset to obtain a word sequence, and to process the word sequence using a pre-trained language model to obtain an output sequence, wherein the output sequence includes the embedding representation of each word in the word sequence.
[0139] The second processing module 43 is used to perform feature extraction processing on the output sequence using a pre-trained text convolutional neural network and a pre-trained bidirectional long short-term memory network to obtain text feature vectors.
[0140] The determination module 44 is used to determine the knowledge graph corresponding to the loan enterprise based on the first dataset, wherein the knowledge graph includes at least: triple data, node attribute data and edge attribute data.
[0141] The third processing module 45 is used to process triplet data, node attribute data and edge attribute data using a pre-trained graph network to obtain node embedding feature vectors that include risk information.
[0142] The generation module 46 is used to process the text feature vector and node embedding feature vector using a pre-trained risk assessment model to obtain the risk score of the loan enterprise.
[0143] Optionally, the first processing module 42 is further configured to perform the following steps: add classification markers and delimiter markers to the beginning and end positions of the word sequence respectively to obtain a tokenized sequence, and map the tokenized sequence to the corresponding symbolized sequence; process the symbolized sequence using a pre-trained language model to obtain an output sequence, wherein the output sequence includes at least: a word vector matrix, and the word vector matrix includes: the hidden layer dimension corresponding to each symbol in the symbolized sequence.
[0144] Optionally, the second processing module 43 is further configured to perform the following steps: extract the first vector corresponding to the label position of the classification label in the word vector matrix, and determine the first vector as the text global feature vector; extract local semantic features in the word vector matrix using multiple convolution kernels of different sizes in a pre-trained text convolutional neural network to obtain the first feature vector; capture long-distance dependencies in the word vector matrix using a pre-trained bidirectional long short-term memory network to obtain multiple temporal features, and perform mean pooling on the multiple temporal features to obtain the second feature vector; and concatenate the text global feature vector, the first feature vector, and the second feature vector to obtain the text feature vector.
[0145] Optionally, the generation module 46 is further configured to perform the following steps: imputing missing values and normalizing the tabular data in the first dataset to obtain target tabular data; determining the month-on-month revenue growth rate within m months based on the revenue of the current month and the revenue of the previous m months, and determining the time-series features based on the month-on-month revenue growth rate within m months, where m is a positive integer; determining the interest coverage ratio based on earnings before interest and taxes and interest expenses, and determining the ratio features based on the interest coverage ratio; determining derived features based on the time-series features and the ratio features; performing feature extraction on the target tabular data, and adding derived features to the feature extraction results to obtain tabular features; and processing the text feature vector, node embedding feature vector, and tabular features using a pre-trained risk assessment model to obtain the risk score of the loan enterprise.
[0146] Optionally, the generation module 46 is further configured to perform the following steps: process the table features using a fully connected layer to obtain a hot-start node embedding vector; process the hot-start node embedding vector using a pre-trained graph network to obtain a graph structure feature vector; fuse the graph structure feature vector with the text feature vector to obtain a first target feature vector; fuse the first target feature vector with the node embedding feature vector to obtain a second target feature vector; and process the second target feature vector using a pre-trained risk assessment model to obtain a risk score for the loan company.
[0147] Optionally, the language model and text convolutional neural network are trained using the following method: Evaluation data is acquired and divided into training and testing sets; the text data in the training set is segmented to obtain a first word sequence, and the language model is used to process the first word sequence to obtain a first output sequence; the language model and text convolutional neural network are trained using the first output sequence. The text convolutional neural network and bidirectional long short-term memory network are trained using the following method: the text convolutional neural network and bidirectional long short-term memory network are trained using the first word sequence to obtain a first text vector output by the text convolutional neural network and bidirectional long short-term memory network; the first output sequence is merged into the first text vector as an intermediate feature to obtain first feature data; the corresponding label of the first feature data is used as the first target variable, and the first target variable is combined with historical target variables to fine-tune the parameters of the text convolutional neural network to obtain the parameters of the first target variable, wherein the historical target variable is the risk category label already labeled in the text data of the training set, used to supervise model training.
[0148] Optionally, the graph network is trained using the following method: A first knowledge graph is determined based on the training set; the graph network is used to process the first triplet data, first node attribute data, and first edge attribute data in the first knowledge graph to obtain a first node embedding feature vector including risk information; the first feature data and the first node embedding feature vector are fused to obtain second feature data; the corresponding label of the second feature data is used as a second target variable, and the second target variable is combined with historical target variables to fine-tune the parameters of the graph network, obtaining the second target variable parameters; the risk assessment model is trained using the following method: the risk assessment model is used to process the first text vector and the first node embedding feature vector to obtain a first output result; the second feature data and the first output result are fused to obtain third feature data; the corresponding label of the third feature data is used as a third target variable, and the third target variable is combined with historical target variables to fine-tune the parameters of the risk assessment model, obtaining the third target variable parameters.
[0149] It should be noted that the above Figure 4 The modules in the above can be program modules (e.g., a set of program instructions that implement a specific function) or hardware modules. For the latter, they can be represented in the following forms, but are not limited to these: each of the above modules is represented by a processor, or the functions of each of the above modules are implemented by a processor.
[0150] It should be noted that, Figure 4 Preferred embodiments of the shown examples can be found in [reference needed]. Figure 1 The relevant descriptions of the embodiments shown will not be repeated here.
[0151] Figure 5 A hardware block diagram of a computer terminal for implementing a risk assessment method is shown. Figure 5 As shown, the computer terminal 50 may include one or more processors 502 (shown as 502a, 502b, ..., 502n in the figure) 502 (processor 502 may include, but is not limited to, a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 504 for storing data, and a transmission module 506 for communication functions. In addition, it may also include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of a BUS bus), a network interface, a power supply, and / or a camera. Those skilled in the art will understand that... Figure 5 The structure shown is for illustrative purposes only and does not limit the structure of the aforementioned electronic device. For example, computer terminal 50 may also include... Figure 5 The more or fewer components shown, or having the same Figure 5 The different configurations shown.
[0152] It should be noted that the aforementioned one or more processors 502 and / or other data processing circuits are generally referred to herein as "data processing circuits". These data processing circuits may be embodied, in whole or in part, in software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuits may be a single, independent processing module, or may be integrated, in whole or in part, into any other element within the computer terminal 50. As involved in the embodiments of this application, the data processing circuits serve as a processor control mechanism (e.g., selection of a variable resistor termination path connected to an interface).
[0153] The memory 504 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the risk assessment method in this embodiment. The processor 502 executes various functional applications and data processing by running the software programs and modules stored in the memory 504, thereby realizing the aforementioned risk assessment method. The memory 504 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 504 may further include memory remotely located relative to the processor 502, and these remote memories can be connected to the computer terminal 50 via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0154] The transmission module 506 is used to receive or send data via a network. Specific examples of the network described above may include a wireless network provided by the communication provider of the computer terminal 50. In one example, the transmission module 506 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission module 506 may be a Radio Frequency (RF) module, used for wireless communication with the Internet.
[0155] The display may be, for example, a touchscreen liquid crystal display (LCD) that allows the user to interact with the user interface of the computer terminal 50.
[0156] It should be noted here that, in some optional embodiments, the above... Figure 5 The computer terminal shown may include hardware elements (including circuitry), software elements (including computer code stored on a computer-readable medium), or a combination of both hardware and software elements. It should be noted that... Figure 5This is only one instance of a specific particular instance, and is intended to illustrate the types of components that may exist in the aforementioned computer terminal.
[0157] It should be noted that, Figure 5 The computer terminal shown is used to execute Figure 1 The risk assessment method shown above applies to this electronic device as well, and will not be repeated here.
[0158] This application also provides a non-volatile storage medium, which includes a stored program, wherein the program, when running, controls the device where the storage medium is located to execute the above-mentioned risk assessment method.
[0159] A non-volatile storage medium performs the following functions: It acquires a first dataset, which includes data related to loan companies; it segments the text data in the first dataset to obtain word sequences, and processes these word sequences using a pre-trained language model to obtain an output sequence, where the output sequence includes the embedding representation of each word in the word sequence; it extracts features from the output sequence using a pre-trained text convolutional neural network and a pre-trained bidirectional long short-term memory network to obtain text feature vectors; based on the first dataset, it determines the knowledge graph corresponding to the loan companies, where the knowledge graph includes at least: triple data, node attribute data, and edge attribute data; it processes the triple data, node attribute data, and edge attribute data using a pre-trained graph network to obtain node embedding feature vectors including risk information; and it processes the text feature vectors and node embedding feature vectors using a pre-trained risk assessment model to obtain the risk score of the loan companies.
[0160] This application also provides an electronic device, including: a memory and a processor, wherein the processor is used to run a program stored in the memory, wherein the program executes the risk assessment method described above when it runs.
[0161] The processor is used to run a program that performs the following functions: acquires a first dataset, which includes data related to loan companies; segments the text data in the first dataset into word sequences, and processes the word sequences using a pre-trained language model to obtain an output sequence, wherein the output sequence includes the embedding representation of each word in the word sequence; performs feature extraction processing on the output sequence using a pre-trained text convolutional neural network and a pre-trained bidirectional long short-term memory network to obtain a text feature vector; determines the knowledge graph corresponding to the loan companies based on the first dataset, wherein the knowledge graph includes at least: triple data, node attribute data, and edge attribute data; processes the triple data, node attribute data, and edge attribute data using a pre-trained graph network to obtain a node embedding feature vector including risk information; and processes the text feature vector and node embedding feature vector using a pre-trained risk assessment model to obtain a risk score for the loan companies.
[0162] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0163] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0164] In the above embodiments of this application, the information collected is information and data authorized by the user or fully authorized by all parties, and the collection, storage, use, processing, transmission, provision, disclosure and application of the relevant data all comply with relevant laws, regulations and standards, take necessary protective measures, do not violate public order and good morals, and provide corresponding operation entry points for users to choose to authorize or refuse.
[0165] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0166] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0167] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0168] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0169] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.
Claims
1. A risk assessment method, characterized in that, include: Obtain a first dataset, wherein the first dataset includes data related to loan companies; The text data in the first dataset is segmented to obtain a word sequence, and the word sequence is processed using a pre-trained language model to obtain an output sequence, wherein the output sequence includes the embedding representation of each word in the word sequence; The output sequence is processed by pre-trained text convolutional neural network and pre-trained bidirectional long short-term memory network to obtain text feature vectors. Based on the first dataset, a knowledge graph corresponding to the loan enterprise is determined, wherein the knowledge graph includes at least: triple data, node attribute data, and edge attribute data; The triplet data, node attribute data, and edge attribute data are processed using a pre-trained graph network to obtain node embedding feature vectors that include risk information. The risk score of the lending company is obtained by processing the text feature vector and the node embedding feature vector using a pre-trained risk assessment model.
2. The method according to claim 1, characterized in that, The word sequence is processed using a pre-trained language model to obtain an output sequence, including: A classification marker and a delimiter marker are added to the beginning and end of the word sequence to obtain a tokenized sequence, and the tokenized sequence is mapped to the corresponding symbolic sequence. The pre-trained language model is used to process the symbolized sequence to obtain an output sequence, wherein the output sequence includes at least a character vector matrix, and the character vector matrix includes the hidden layer dimension corresponding to each symbol in the symbolized sequence.
3. The method according to claim 2, characterized in that, The output sequence is processed by a pre-trained text convolutional neural network and a pre-trained bidirectional long short-term memory network to extract features, resulting in a text feature vector, including: Extract the first vector corresponding to the label position of the classification label in the word vector matrix, and determine the first vector as the text global feature vector; The first feature vector is obtained by extracting local semantic features from the character vector matrix using multiple convolutional kernels of different sizes in the pre-trained text convolutional neural network. The pre-trained bidirectional long short-term memory network is used to capture long-distance dependencies in the word vector matrix to obtain multiple temporal features. The average pooling of the multiple temporal features is then performed to obtain a second feature vector. The text global feature vector, the first feature vector, and the second feature vector are concatenated to obtain the text feature vector.
4. The method according to claim 1, characterized in that, The risk score of the lending company is obtained by processing the text feature vector and the node embedding feature vector using a pre-trained risk assessment model, including: Missing values are filled and normalized in the tabular data of the first dataset to obtain the target tabular data; Based on the revenue of the current month and the revenue of the previous m months, determine the month-on-month growth rate of revenue over the m months, and determine the time series characteristics based on the month-on-month growth rate of revenue over the m months, where m is a positive integer; The interest coverage ratio is determined based on earnings before interest and taxes (EBIT) and interest expense, and the ratio characteristics are determined based on the interest coverage ratio. Based on the time-series characteristics and the ratio characteristics, derived characteristics are determined; The target table data is subjected to feature extraction, and the derived features are added to the feature extraction results to obtain table features; The pre-trained risk assessment model is used to process the text feature vector, the node embedding feature vector, and the table features to obtain the risk score of the loan enterprise.
5. The method according to claim 4, characterized in that, The pre-trained risk assessment model is used to process the text feature vector, the node embedding feature vector, and the table features to obtain the risk score of the lending company, including: The table features are processed using a fully connected layer to obtain the hot-start node embedding vector; The hot-start node embedding vector is processed using the pre-trained graph network to obtain the graph structure feature vector; The graph structure feature vector is fused with the text feature vector to obtain the first target feature vector; The first target feature vector is fused with the node embedding feature vector to obtain the second target feature vector; The second target feature vector is processed using the pre-trained risk assessment model to obtain the risk score of the lending enterprise.
6. The method according to claim 1, characterized in that, The language model and the text convolutional neural network were trained using the following method: Obtain evaluation data and divide the evaluation data into training set and test set; The text data in the training set is segmented to obtain a first word sequence, and the first word sequence is processed using a language model to obtain a first output sequence. The language model and the text convolutional neural network are trained using the first output sequence, respectively. The text convolutional neural network and the bidirectional long short-term memory network were trained using the following method: Using the first word sequence, the text convolutional neural network and the bidirectional long short-term memory network are trained respectively to obtain the first text vector output by the text convolutional neural network and the bidirectional long short-term memory network; The first output sequence is used as an intermediate feature and merged into the first text vector to obtain the first feature data; The corresponding label of the first feature data is used as the first target variable, and the first target variable is combined with the historical target variable to fine-tune the parameters of the text convolutional neural network to obtain the parameters of the first target variable. The historical target variable is the risk category label that has been labeled in the text data of the training set, which is used to supervise the training of the model.
7. The method according to claim 6, characterized in that, The graph network was trained using the following method: Based on the training set, a first knowledge graph is determined; The first triplet data, first node attribute data, and first edge attribute data in the first knowledge graph are processed using a graph network to obtain the first node embedding feature vector including risk information. The first feature data and the first node embedded feature vector are fused to obtain the second feature data; The corresponding label of the second feature data is used as the second target variable, and the second target variable is combined with the historical target variable to fine-tune the parameters of the graph network, thereby obtaining the parameters of the second target variable; The risk assessment model was trained using the following method: The first text vector and the first node embedding feature vector are processed using a risk assessment model to obtain the first output result; The second feature data and the first output result are fused to obtain the third feature data; The corresponding label of the third feature data is used as the third target variable, and the third target variable is combined with the historical target variable to fine-tune the parameters of the risk assessment model, thereby obtaining the parameters of the third target variable.
8. A risk assessment device, characterized in that, include: The acquisition module is used to acquire a first dataset, wherein the first dataset includes data related to loan companies; The first processing module is used to perform word segmentation on the text data in the first dataset to obtain a word sequence, and to process the word sequence using a pre-trained language model to obtain an output sequence, wherein the output sequence includes the embedding representation of each word in the word sequence; The second processing module is used to perform feature extraction processing on the output sequence using a pre-trained text convolutional neural network and a pre-trained bidirectional long short-term memory network to obtain a text feature vector. The determination module is used to determine the knowledge graph corresponding to the loan enterprise based on the first dataset, wherein the knowledge graph includes at least: triple data, node attribute data and edge attribute data; The third processing module is used to process the triplet data, the node attribute data and the edge attribute data using a pre-trained graph network to obtain a node embedding feature vector including risk information. The generation module is used to process the text feature vector and the node embedding feature vector using a pre-trained risk assessment model to obtain the risk score of the lending enterprise.
9. A non-volatile storage medium, characterized in that, The non-volatile storage medium includes a stored program, wherein, when the program is executed, it controls the device containing the non-volatile storage medium to perform the risk assessment method according to any one of claims 1 to 7.
10. An electronic device, characterized in that, include: A memory and a processor, the processor being configured to run a program stored in the memory, wherein the program, when running, performs the risk assessment method according to any one of claims 1 to 7.
11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the risk assessment method according to any one of claims 1 to 7.
Citation Information
Cited By
Risk detection method, server, medium and equipment
CN121304325A