Market subject evaluation method and system based on risk entropy model and big data processing
By integrating multi-source data and constructing an enterprise risk knowledge graph through risk entropy model and big data processing, explicit and implicit risk paths are identified, which solves the shortcomings of existing market entity assessment methods, realizes the comprehensiveness and dynamism of market entity credit assessment, and generates scientific and interpretable risk assessment results.
Patent Information
- Application Number
- CN202510798824.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2045-06-16
AI Technical Summary
Existing market entity assessment methods rely on single structured data, making it difficult to fully integrate the dynamic behavior of enterprises and changes in the external environment. Traditional models are insufficient in the integration of multi-dimensional risk characteristics, lack adaptive assessment and update mechanisms, and cannot reflect the credit evolution process of market entities in real time.
By employing a risk entropy model and big data processing approach, a corporate risk knowledge graph is constructed through multi-source data fusion and natural language processing. This graph identifies explicit and implicit risk paths, and combines fuzzy membership degree and time decay factor to conduct multi-dimensional risk assessment, ultimately generating the subject's risk level score.
It has achieved efficient integration and standardization of market entity credit data, improved the comprehensiveness and dynamic adaptability of risk assessment, and generated scientific and interpretable risk assessment results.
Smart Images

Figure CN120975902A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of big data analysis and risk assessment, and particularly relates to a market subject evaluation method and system based on a risk entropy value model and big data processing. BACKGROUND
[0002] In recent years, with the rapid development of domestic and foreign financial markets, the financial industry has gradually evolved into a typical industry highly dependent on information technology support, and the circulation efficiency, coverage and data quality of information in financial activities have become increasingly prominent. Financial risk management and decision-making processes increasingly rely on comprehensive acquisition and accurate identification of enterprise subject information, and the timeliness, authenticity and integrity of information have become the core prerequisite for ensuring the safety and efficiency of the financial system. However, in actual market operation, information distortion, data missing, and update lag frequently occur, directly affecting the risk judgment and decision-making effect of financial assets.
[0003] The existing market subject evaluation method mainly relies on a single, structured historical data source, which is difficult to comprehensively integrate the dynamic behavior of enterprises and changes in the external environment. More importantly, traditional models are insufficient in multi-dimensional risk feature fusion, lack adaptive evaluation and updating mechanisms based on big data processing, and are difficult to reflect the credit evolution process of market subjects in real time. Therefore, there is an urgent need for a method that can accurately identify risks by deeply integrating multi-source heterogeneous data and intelligently identifying unstructured risk factors, to achieve the comprehensiveness, objectivity and dynamics of market subject evaluation. SUMMARY
[0004] To solve the above problems in the prior art, the application provides a market subject evaluation method and system based on a risk entropy value model and big data processing, The object of the application can be achieved by the following technical solutions: A market subject evaluation method based on a risk entropy value model and big data processing, comprising: S1: acquiring market credit data, outputting market subject benchmark credit data by multi-source data fusion and feature standardization processing of the market credit data, and storing the market subject benchmark credit data in batches using a distributed framework; S2: constructing an enterprise risk knowledge graph based on the market subject benchmark credit data, and identifying a risk factor set in the enterprise risk knowledge graph through a relation extraction model; S3: presetting a multi-dimensional risk entropy value evaluation model, and outputting a multi-dimensional risk entropy value sequence through the multi-dimensional risk entropy value evaluation model for the risk factor set; S4: performing subject scoring based on the multi-dimensional risk entropy value sequence to obtain a subject risk level score result, and visualizing.
[0005] Preferably, the processing procedure of the market subject baseline credit data in step S1 comprises: S101: the market credit data comprises structured data and unstructured data; S102: obtaining market credit structured data by normalizing the structured data; S103: obtaining market credit unstructured data by processing the unstructured data through natural language processing; S104: forming a set of trusted labels by labeling the market credit structured data and the market credit unstructured data through a confidence correction mechanism; generating a multi-source market credit set by storing the set of trusted labels; S105: outputting market subject baseline credit data by unifying the format of the multi-source market credit set.
[0006] Preferably, the natural language processing in step S103 comprises word segmentation, named entity recognition, and syntax analysis.
[0007] Preferably, the processing procedure of the confidence correction mechanism in step S104 is as follows: S104-1: obtaining a confidence score by calculating the confidence of the market credit structured data and the market credit unstructured data; S104-2: pre-setting a confidence threshold, if the confidence score is less than the confidence threshold, labeling the data; if the confidence score is greater than the confidence threshold, storing it in the set of trusted labels.
[0008] Preferably, the identification procedure of the risk factor set in step S2 is as follows: S201: identifying entity relationships in the market subject baseline credit data; S202: identifying implicit risk paths to generate the risk factor set based on the entity relationships and the enterprise risk knowledge graph through a path reasoning algorithm.
[0009] Preferably, the operation procedure of the path reasoning algorithm in step S202 comprises: S202-1: mapping the entity relationships to a node set and semantic edges in the enterprise risk knowledge graph to generate a basic risk path set; S202-2: obtaining nodes and edges of the basic risk path set and performing weighted summation to obtain path weights; S202-3: pre-setting a risk threshold, retaining paths with path weights greater than the risk threshold to form a candidate risk path set; S202-4: extracting associated risk factors based on the candidate risk path set to output the risk factor set.
[0010] Preferably, the modeling process of the multi-dimensional risk entropy value evaluation model in step S3 is as follows: S301: Obtain a risk factor in the set of risk factors, and calculate a fuzzy membership degree based on the risk factor; S302: Calculate a fuzzy entropy value based on the fuzzy membership degree; S303: Obtain a historical risk factor triggering time, and calculate a time-decay risk entropy based on the historical risk factor triggering time and the fuzzy entropy value; S304: Output the multi-dimensional risk entropy value sequence by storing the time-decay risk entropy.
[0011] Preferably, the scoring process of the subject score in step S4 is as follows: S401: Obtain a time variation rate and a triggering intensity, and obtain a multi-dimensional risk entropy value in the multi-dimensional risk entropy value sequence; S402: Calculate a dynamic scoring weight based on the time variation rate, the triggering intensity, and the multi-dimensional risk entropy value; S403: Obtain a final score based on the dynamic scoring weight; S404: Map the final score to a grading level to obtain a subject risk level score result.
[0012] Preferably, the mapping process in step S404 is as follows: S404-1: Predefine a mapping rule, a first threshold value, and a second threshold value, and map the final score to a scoring interval based on the mapping rule to obtain a risk level; The judgment process of the mapping is as follows: If the final score is greater than the first threshold value, the risk level is “high risk”; Otherwise, if the final score is greater than the second threshold value, the risk level is “medium risk”; if the final score is less than the second threshold value, the risk level is “low risk”; S404-2: Match a color code according to the risk level, and dynamically display the subject score result and the associated risk path in an interactive dashboard.
[0013] A market subject evaluation system based on a risk entropy value model and big data processing, comprising a data preprocessing module, a relationship extraction module, a risk evaluation module, and a subject evaluation module, and comprising: The data preprocessing module is configured to obtain market credit data, fuse and standardize features of the market credit data to output market subject benchmark credit data, and store the market subject benchmark credit data in batches by using a distributed framework; The relationship extraction module is configured to construct an enterprise risk knowledge graph based on the market subject benchmark credit data, and identify a risk factor set in the enterprise risk knowledge graph through a relationship extraction model; The risk assessment module is configured to preset a multi-dimensional risk entropy value evaluation model, and output a multi-dimensional risk entropy value sequence of the risk factor set through the multi-dimensional risk entropy value evaluation model. The subject assessment module is configured to perform subject scoring based on the multi-dimensional risk entropy value sequence to obtain a subject risk level score result, and perform visualization.
[0014] The present application has the following advantages: (1) Through structured data normalization and unstructured data natural language processing, combined with a confidence correction mechanism and a distributed storage framework, efficient integration and standardization of multi-source market credit data are achieved. This significantly improves the efficiency of data processing and ensures the credibility and consistency of the data, providing high-quality input for subsequent risk assessment.
[0015] (2) Based on the enterprise risk knowledge graph and path reasoning algorithm, explicit and implicit risk paths can be identified, combined with fuzzy membership and time decay factors, to dynamically correct the timeliness and severity weight of risk factors. This mechanism enhances the depth of risk factor mining and improves the comprehensiveness and dynamic adaptability of risk assessment.
[0016] (3) Through the multi-dimensional risk entropy value evaluation model, risk factors are converted into fuzzy entropy values and time decay risk entropy, quantifying the risk impact of different dimensions. Combined with time variation rate, trigger strength and dynamic scoring weight, the subject risk level score result is finally generated. This modeling method makes the risk assessment results more scientific and interpretable. BRIEF DESCRIPTION OF DRAWINGS
[0017] In order to facilitate understanding by those skilled in the art, the present application will be further described below with reference to the accompanying drawings.
[0018] Figure 1 The flowchart of the market subject assessment method based on the risk entropy value model and big data processing of the present application. DETAILED DESCRIPTION
[0019] In order to further illustrate the technical means and effects adopted by the present application to achieve the predetermined invention purpose, the specific embodiments, structures, features and effects according to the present application are described in detail below with reference to the accompanying drawings and preferred embodiments.
[0020] Please refer to Figure 1 A market subject assessment method based on a risk entropy value model and big data processing, comprising: S1: Obtain market credit data, output market subject baseline credit data by multi-source data fusion and feature standardization processing of the market credit data, and store the market subject baseline credit data in batches by using a distributed framework; S2: Construct an enterprise risk knowledge graph based on the market subject baseline credit data, and identify a risk factor set in the enterprise risk knowledge graph by using a relation extraction model; S3: Pre-set a multi-dimensional risk entropy value evaluation model, and output a multi-dimensional risk entropy value sequence by using the multi-dimensional risk entropy value evaluation model on the risk factor set; S4: Perform subject scoring based on the multi-dimensional risk entropy value sequence to obtain a subject risk level score result, and perform visualization.
[0021] Specifically, the market credit data in the step S1 includes structured data and unstructured data, the structured data includes financial statements, credit records, and business registration information, and the unstructured data includes public opinion texts, announcements, news, and social media content.
[0022] Specifically, the processing process of the market subject baseline credit data in the step S1 includes: S101: The market credit data includes structured data and unstructured data; S102: Obtain market credit structure data by normalizing the structured data; S103: Obtain market credit unstructured data by processing the unstructured data by using natural language processing; S104: Form a trusted label set by labeling the market credit structure data and the market credit unstructured data by using a confidence correction mechanism, and generate a multi-source market credit set by storing the trusted label set; S105: Output market subject baseline credit data by unifying the format of the multi-source market credit set.
[0023] Specifically, the natural language processing in the step S103 includes word segmentation, named entity recognition, and syntax analysis.
[0024] Specifically, the processing process of the confidence correction mechanism in the step S104 is as follows: S104-1: Obtain a confidence score by calculating the confidence of the market credit structure data and the market credit unstructured data; S104-2: Pre-set a confidence threshold, if the confidence score is less than the confidence threshold, label the data, and if the confidence score is greater than the confidence threshold, store it in the trusted label set.
[0025] In the embodiment, the indicators considered in the calculation of the confidence score mainly include source credibility, content consistency, information integrity, and context signals (negative words and speculative words in the text), and then the confidence score is obtained by weighted summation of the indicators.
[0026] Specifically, the market subject baseline credit data in the step S105 is a triple data in the form of <subject ID, feature type, feature value>.
[0027] Specifically, the enterprise risk knowledge graph in the step S2 is composed of a node set and semantic edges, the node set includes enterprises, events, and personnel, and the semantic edges include litigation relationships, penalty relationships, and investment relationships.
[0028] Specifically, the identification process of the risk factor set in the step S2 is as follows: S201: identifying the entity relationship in the market subject baseline credit data; S202: identifying the implicit risk path to generate the risk factor set based on the entity relationship and the enterprise risk knowledge graph through a path reasoning algorithm.
[0029] Specifically, the entity relationship in the step S201 is represented as: <subject, risk behavior / event, time / associated object>; The attributes contained in each factor in the risk factor set in the step S202 are: r i =<type, trigger time, severity, frequency, associated subject>, wherein r i represents a factor in the risk factor set.
[0030] Specifically, the operation process of the path reasoning algorithm in the step S202 includes: S202-1: mapping the entity relationship to the node set and semantic edges in the enterprise risk knowledge graph to generate a basic risk path set; S202-2: obtaining the nodes and edges of the basic risk path set and performing weighted summation to obtain a path weight; S202-3: presetting a risk threshold, and retaining the paths with a path weight greater than the risk threshold to form a candidate risk path set; S202-4: extracting associated risk factors based on the candidate risk path set to output the risk factor set.
[0031] Specifically, the modeling process of the multi-dimensional risk entropy evaluation model in the step S3 is as follows: S301: Obtain a risk factor in the set of risk factors, and calculate a fuzzy membership degree based on the risk factor; The calculation formula is: , wherein, μ i is the fuzzy membership degree, e represents a logarithmic function calculation, γ is a decay rate, t now is a current time, t is an event start time, severity(r i ) represents a severity of the risk factor r i , η is a nonlinear adjustment parameter, and k is a smoothing constant; S302: Calculate a fuzzy entropy value through the fuzzy membership degree; The calculation expression of the fuzzy entropy value is: , wherein, H f is the fuzzy entropy value, i represents an i-th risk factor, n represents a total number of risk factors, μ i is the fuzzy membership degree, and ε is a minimum numerical value; S303: Obtain a historical risk factor trigger time, and calculate a time-decay risk entropy according to the historical risk factor trigger time and the fuzzy entropy value; The calculation expression of the time-decay risk entropy is: , wherein, H t represents the time-decay risk entropy, H f is the fuzzy entropy value, λ is a time-decay coefficient, and Δt represents an interval between a current time and an event occurrence time; S304: Output the multi-dimensional risk entropy value sequence by storing the time-decay risk entropy.
[0032] Specifically, the scoring process of the main body score in the step S4 is: S401: Obtain a time change rate and a trigger intensity, and obtain a multi-dimensional risk entropy value in the multi-dimensional risk entropy value sequence; S402: Calculate a dynamic scoring weight according to the time change rate, the trigger intensity, and the multi-dimensional risk entropy value; The calculation expression of the dynamic scoring weight is: , wherein, w t is the dynamic scoring weight, is a decay coefficient, H t represents the time-decay risk entropy, and I tFor the trigger strength, At represents the time change rate, and ε is a smoothing constant; S403: obtaining a final score according to the dynamic score weight; The mathematical expression of the final score is: , Wherein, S represents the final score, t represents the tth time decay risk entropy, m represents the number of time decay risk entropies, w t is the dynamic score weight, H t represents the time decay risk entropy; S404: mapping the final score to a grading level to obtain the subject risk level score result.
[0033] Specifically, the mapping process in the step S404 is: S404-1: presetting a mapping rule, a first threshold value, and a second threshold value, and mapping the final score to a score interval based on the mapping rule to obtain a risk level; The judgment process of the mapping is: If the final score is greater than the first threshold value, the risk level is "high risk"; If not, if the final score is greater than the second threshold value, the risk level is "medium risk"; if the final score is less than the second threshold value, the risk level is "low risk"; S404-2: matching color coding according to the risk level, and dynamically displaying the subject score result and the associated risk path in the interactive dashboard.
[0034] Specifically, the color coding includes: if the risk level is low risk, the color coding is green; if the risk level is medium risk, the color coding is orange; if the risk level is high risk, the color coding is red.
[0035] In this embodiment, a market subject evaluation system based on a risk entropy value model and big data processing includes a data preprocessing module, a relationship extraction module, a risk assessment module, and a subject evaluation module, comprising: The data preprocessing module is configured to obtain market credit data, fuse and standardize the market credit data to output market subject benchmark credit data, and store the market subject benchmark credit data in batches using a distributed framework; The relationship extraction module is configured to construct an enterprise risk knowledge graph based on the market subject benchmark credit data, and identify a risk factor set in the enterprise risk knowledge graph through a relationship extraction model; The risk assessment module is configured to preset a multi-dimensional risk entropy value evaluation model, and output a multi-dimensional risk entropy value sequence by inputting the risk factor set into the multi-dimensional risk entropy value evaluation model. The subject evaluation module is configured to perform subject scoring based on the multi-dimensional risk entropy value sequence to obtain a subject risk grade scoring result and perform visualization.
[0036] The above is only a preferred embodiment of the present application, and does not limit the present application in any form. Although the present application has been disclosed as above with a preferred embodiment, it is not intended to limit the present application. Any person skilled in the art can make slight changes or modifications to the above disclosed technical content to obtain equivalent embodiments with equivalent changes, without departing from the technical solution of the present application. Any simple modification, equivalent change and modification of the above embodiments based on the technical essence of the present application are still within the scope of the technical solution of the present application.
Claims
1. A market entity assessment method based on a risk entropy model and big data processing, characterized in that, Includes the following steps: S1: Acquire market credit data, process the market credit data through multi-source data fusion and feature standardization to output benchmark credit data of market entities, and use a distributed framework to store the benchmark credit data of market entities in batches; S2: Construct an enterprise risk knowledge graph based on the benchmark credit data of the market entities, and identify the risk factor set in the enterprise risk knowledge graph through a relation extraction model; S3: Preset a multidimensional risk entropy value assessment model, and output a multidimensional risk entropy value sequence from the risk factor set through the multidimensional risk entropy value assessment model; S4: Based on the multidimensional risk entropy value sequence, the subject is scored to obtain the subject risk level score result, and then visualized.
2. The market entity assessment method based on risk entropy model and big data processing according to claim 1, characterized in that, The processing procedure for market entity benchmark credit data in step S1 includes: S101: The market credit data includes structured data and unstructured data; S102: Obtain market credit structure data by normalizing the structured data; S103: Obtain market credit unstructured data by processing the unstructured data using natural language; S104: Label the market credit structure data and the market credit unstructure data using a confidence correction mechanism to form a trusted label set; generate a multi-source market credit set by storing the trusted label set; S105: Output benchmark credit data of market entities by unifying the format of the multi-source market credit set.
3. The market entity assessment method based on risk entropy model and big data processing according to claim 2, characterized in that, The natural language processing in step S103 includes word segmentation, named entity recognition, and syntactic analysis.
4. The market entity assessment method based on risk entropy model and big data processing according to claim 2, characterized in that, The confidence correction mechanism in step S104 is processed as follows: S104-1: Obtain a confidence score by calculating the confidence levels of the market credit structure data and the market credit unstructure data; S104-2: A pre-set confidence threshold is used. If the confidence score is less than the confidence threshold, the data is labeled. If the confidence score is greater than the confidence threshold, it is stored in the trusted tag set.
5. The market entity assessment method based on risk entropy model and big data processing according to claim 1, characterized in that, The process of identifying the risk factor set in step S2 is as follows: S201: By identifying entity relationships in the benchmark credit data of the market entities; S202: Based on the entity relationships and the enterprise risk knowledge graph, the risk factor set is generated by identifying implicit risk paths through a path reasoning algorithm.
6. The market entity assessment method based on risk entropy model and big data processing according to claim 5, characterized in that, The path reasoning algorithm in step S202 includes the following steps: S202-1: Map the entity relationships to the node set and semantic edges in the enterprise risk knowledge graph to generate a basic risk path set; S202-2: Obtain the nodes and edges of the basic risk path set, and perform a weighted summation to obtain the path weights; S202-3: Preset a risk threshold, and retain paths with path weights greater than the risk threshold to form a candidate risk path set; S202-4: Extract associated risk factors based on the candidate risk path set and output the risk factor set.
7. The market entity assessment method based on risk entropy model and big data processing according to claim 1, characterized in that, The modeling process of the multidimensional risk entropy value assessment model in step S3 is as follows: S301: Obtain the risk factors in the risk factor set, and calculate the fuzzy membership degree based on the risk factors; S302: Calculate the fuzzy entropy value using the fuzzy membership degree; S303: Obtain the historical risk factor trigger time, and calculate the time decay risk entropy based on the historical risk factor trigger time and the fuzzy entropy value; S304: Output the multidimensional risk entropy value sequence by storing the time decay risk entropy.
8. The market entity assessment method based on risk entropy model and big data processing according to claim 1, characterized in that, The scoring process for the subject score in step S4 is as follows: S401: Obtain the time change rate and trigger intensity, and obtain the multidimensional risk entropy value in the multidimensional risk entropy value sequence; S402: Calculate the dynamic scoring weight based on the time change rate, the trigger intensity, and the multidimensional risk entropy value; S403: Obtain the final score based on the dynamic scoring weights; S404: Map the final score to the grade level to obtain the subject risk level score result.
9. The market entity assessment method based on risk entropy model and big data processing according to claim 1, characterized in that, The mapping process in step S404 is as follows: S404-1: Preset mapping rules, first threshold, and second threshold; based on the mapping rules, map the final score to a score range to obtain the risk level. The process for determining the mapping is as follows: Determine whether the final score is greater than the first threshold; if yes, the risk level is "high risk". No, if the final score is greater than the second threshold, the risk level is "medium risk"; if the final score is less than the second threshold, the risk level is "low risk". S404-2: Match color codes according to the risk level and dynamically display the subject's scoring results and associated risk paths in the interactive dashboard.
10. A market entity assessment system based on a risk entropy model and big data processing, comprising a data preprocessing module, a relationship extraction module, a risk assessment module, and an entity assessment module, characterized in that, include: The data preprocessing module is used to acquire market credit data, process the market credit data through multi-source data fusion and feature standardization to output benchmark credit data of market entities, and use a distributed framework to store the benchmark credit data of market entities in batches. The relationship extraction module is used to construct an enterprise risk knowledge graph based on the benchmark credit data of the market entities, and to identify the risk factor set in the enterprise risk knowledge graph through the relationship extraction model; The risk assessment module is used to preset a multidimensional risk entropy value assessment model and output a multidimensional risk entropy value sequence from the risk factor set through the multidimensional risk entropy value assessment model. The subject assessment module is used to score the subject based on the multidimensional risk entropy value sequence to obtain the subject risk level score result, and then visualize it.
Citation Information
Patent Citations
Supply chain financial risk assessment method and system based on big data
CN114841790A
Contract open risk management method and system in electric power spot transaction
CN117391705A
Business processing method and device, equipment, medium and program product
CN119313439A
Digital economic risk assessment method and device based on computer statistics
CN119359016A
Enterprise risk assessment method and device, computer equipment and storage medium
CN119398504A