Risk quantification assessment method and apparatus, electronic device, and storage medium
Patent Information
- Application Number
- CN202611206757.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-10
- Publication Date
- 2026-09-29
AI Technical Summary
然而,上述方法侧重于宏观实体的风险传染建模,输出结果多为定性的概率评分,导致风险识别结果的客观性、可解释性与可执行性存在明显不足
[0019]本发明还提供一种非暂态计算机可读存储介质,其上存储有计算机程序,该计算机程序被处理器执行时实现如上述任一种所述风险量化评估方法。
Smart Images

Figure CN122838575A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of natural language processing technology, and in particular to a risk quantification assessment method, apparatus, electronic device, and storage medium. Background Technology
[0002] Automatically identifying risk transmission logic from unstructured disclosure texts to aid assessment and decision-making has become an important direction in the field of intelligent text analysis.
[0003] Current risk assessment methods typically collect data to construct knowledge graphs and use models such as Graph Neural Networks (GNNs) to perform reasoning simulations on the graphs, thereby outputting a comprehensive risk score and evolution path. However, these methods focus on modeling the risk contagion of macro-entities, and the output results are mostly qualitative probability scores, resulting in significant deficiencies in the objectivity, interpretability, and feasibility of risk identification results. Summary of the Invention
[0004] This invention provides a risk quantification assessment method, apparatus, electronic device, and storage medium to address the deficiencies in the prior art.
[0005] This invention provides a method for quantifying and assessing risk, comprising the following steps: A directed acyclic graph is obtained, which is constructed based on standardized nodes obtained by semantic clustering of risk causal links extracted from risk disclosure texts; Based on the node information of the non-terminating node in the directed acyclic graph, the quantization execution logic corresponding to the non-terminating node is generated. The quantization execution logic is used to characterize the logic for data extraction and quantization judgment of the text. Based on the quantization execution logic, the text to be evaluated is parsed to obtain the node state of the non-termination node; Based on the node status of each non-terminating node, the risk transmission path is traced along the directed acyclic graph to obtain the risk quantification assessment result.
[0006] According to a risk quantification assessment method provided by the present invention, the step of generating quantification execution logic corresponding to the non-terminating node based on the node information of the non-terminating node in the directed acyclic graph includes: Input the normalized name of the non-terminating node and the list of merging factors into the language model, and obtain the quantization execution logic output by the language model, which includes the data lookup path, quantization calculation formula and segmentation decision function.
[0007] According to a risk quantification assessment method provided by the present invention, the language model generates the quantification execution logic based on the following method: Using the language model, under the constraints of a preset data source whitelist, the data search path for extracting the target field is determined, and it is marked whether the target field requires manual inference. The language model is used to generate the quantization calculation formula for performing logical operations on the target field; Based on the different calculation result ranges of the quantization calculation formula, a mapping relationship is constructed to map the calculation results to true, false, and uncertain, respectively, and the mapping relationship is used as the segmentation judgment function; The data lookup path, the quantization calculation formula, and the segmentation judgment function are combined to output the quantization execution logic.
[0008] According to a risk quantification assessment method provided by the present invention, the step of parsing the text to be assessed based on the quantification execution logic to obtain the node state of the non-terminating node includes: Extract the target field from the text to be evaluated according to the data search path; Perform a consistency check on the target field; The target field that has passed the verification is substituted into the quantization calculation formula for calculation, and the calculation result is input into the segmentation judgment function to obtain the node state.
[0009] According to a risk quantification assessment method provided by the present invention, the step of extracting target fields from the text to be assessed according to the data search path includes: According to a preset priority order, multiple table parsing strategies with different formatting rules are used sequentially to parse the text to be evaluated; If the current table parsing strategy successfully matches the data corresponding to the target field, then the target field is output and parsing stops; If no match is found, the table parsing strategy of the next lower priority will be switched to continue parsing.
[0010] According to a risk quantification assessment method provided by the present invention, the consistency verification of the target field includes: If the target field is a composite field containing a whole field and multiple subfields, then calculate the sum of the values of each subfield; If the difference between the sum of the numerical values and the value of the overall field is less than the preset tolerance, then the target field is determined to have passed the consistency check.
[0011] According to a risk quantification assessment method provided by the present invention, the step of parsing the text to be assessed based on the quantification execution logic to obtain the node state of the non-terminating node includes: Based on the quantization execution logic, the text to be evaluated is parsed to determine the extraction result of the target field from the text to be evaluated; If the extraction result is that the target field is not extracted, then the node status is marked as a data missing status; If the extraction result is the extraction of the target field, and the data search path carries a marker that requires manual inference, then the node status is marked as pending inference.
[0012] According to a risk quantification assessment method provided by the present invention, the step of tracing the risk transmission path along the directed acyclic graph based on the node states of each non-terminating node to obtain the risk quantification assessment result includes: An XOR decision gateway is set between each non-terminating node and its downstream node in the directed acyclic graph. If the node state of the non-terminating node is true, it is transmitted to the corresponding downstream node via the XOR decision gateway. When the node states of all nodes on the propagation path from the trigger node to the termination node are valid, the corresponding risk propagation path is activated. The overall risk level of the text to be evaluated is determined by combining the number of activated risk transmission paths and the supporting weights of each path, and is used as the result of the risk quantification assessment.
[0013] According to a risk quantification assessment method provided by the present invention, the step of extracting the risk causal link includes: Based on a given target risk seed, recall target paragraphs from the risk disclosure text that are semantically related to the target risk seed; The target paragraph and the target risk seed are input into the language model, and the risk factors contained in the target paragraph and the causal transmission relationship between each risk factor are extracted to obtain multiple step stages with sequential orientation as the risk causal link.
[0014] According to a risk quantification assessment method provided by the present invention, the step of recalling semantically relevant target paragraphs from the risk disclosure text based on a given target risk seed includes: The risk disclosure text is divided into multiple candidate paragraphs; Each candidate paragraph and the target risk seed are mapped to a semantic vector; Calculate the spatial similarity between the semantic vector of each candidate paragraph and the semantic vector of the target risk seed; Candidate paragraphs whose spatial similarity meets the preset sorting conditions are retained as the target paragraphs.
[0015] According to a risk quantification assessment method provided by the present invention, the multiple steps and stages correspond to different role types, and the role types include triggering factors, intermediate transmission, and risk outcomes. The process of extracting risk factors contained in the target paragraph and the causal transmission relationships between these risk factors, resulting in multiple sequentially pointing step stages as the risk causal chain, includes: Extract the risk factors contained in the target paragraph and the causal transmission relationship between each risk factor to generate a candidate link containing multiple step stages with sequential orientation; In the candidate links, the candidate links whose role type in the first step is the triggering factor, whose role type in the last step is the risk result, and which are consistent with the target risk seed, are designated as the risk causal links.
[0016] According to a risk quantification assessment method provided by the present invention, the step of obtaining the directed acyclic graph includes: The original risk factors in the risk causal chain are semantically equivalently aggregated to form the standardized nodes, and the original risk factors that cannot be quantified are discarded. Count the number of times adjacent standardized nodes appear as related node pairs in different texts to remove duplicates; Using the number of deduplications as the support strength of the corresponding directed edges, a directed acyclic graph is constructed with the target risk seed as the endpoint.
[0017] The present invention also provides a risk quantification assessment device, comprising the following modules: The acquisition module is used to acquire a directed acyclic graph, which is constructed based on standardized nodes obtained by semantic clustering of risk causal links extracted from risk disclosure texts; The generation module is used to generate the quantization execution logic corresponding to the non-terminal node based on the node information of the non-terminal node in the directed acyclic graph. The quantization execution logic is used to characterize the logic for data extraction and quantization judgment of the text. The parsing module is used to parse the text to be evaluated based on the quantization execution logic to obtain the node state of the non-termination node; The evaluation module is used to trace the risk transmission path along the directed acyclic graph based on the node status of each non-terminating node, and obtain the risk quantification evaluation result.
[0018] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the risk quantification assessment method as described above.
[0019] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the risk quantification assessment method as described above.
[0020] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the risk quantification assessment method as described above.
[0021] The risk quantification assessment method, apparatus, electronic device, and storage medium provided by this invention construct a directed acyclic graph (DAG) by semantic clustering of risk causal links extracted from risk disclosure texts, thereby converting the risk transmission logic in unstructured natural language into a reusable knowledge graph. Quantification execution logic is generated using node information from non-terminating nodes in the DAG, avoiding the lack of interpretability caused by reliance on black-box model scoring. Furthermore, the quantification execution logic parses the text to be assessed to determine node states, and traces the risk transmission path along the DAG to obtain the risk quantification assessment result. This achieves end-to-end traceable risk rating, improving the objectivity and cross-sample comparability of risk identification in complex business scenarios, and avoiding the business pain points of traditional risk control's excessive reliance on manual review and the difficulty in verifying judgment criteria. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating the risk quantification assessment method provided by the present invention.
[0024] Figure 2 This is a flowchart illustrating the method for constructing a directed acyclic graph provided by the present invention.
[0025] Figure 3 This is a flowchart illustrating another risk quantification assessment method provided by the present invention.
[0026] Figure 4 This is a schematic diagram of the risk quantification assessment device provided by the present invention.
[0027] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0029] With the increasing maturity of Large Language Models (LLM) and related fintech technologies, more and more tasks that heavily rely on human experience, such as credit due diligence and risk factor scanning, are gradually being assisted or even replaced by machines. In intelligent analysis of financial texts, standardized structured data is relatively easy to automate parsing and judgment. However, risk disclosures presented in unstructured natural language are more difficult to automate due to the complex causal transmission logic and cross-validation of multi-source signals involved. In particular, identifying the complete causal link from "triggering factors to intermediate transmission to risk outcomes" disclosed in risk sections is quite challenging. Automatically summarizing the causal transmission paths between risk factors and transforming them into quantitative judgment rules that can be directly executed in new reports is the core link in achieving an end-to-end closed loop from risk knowledge mining to automated execution.
[0030] Current methods for enterprise risk identification and risk transmission analysis often employ risk semantic parsing and evolution path identification techniques based on unstructured text. These techniques construct evolutionary dependencies by parsing text structure and risk semantic units, thereby outputting a qualitative risk evolution schedule. Alternatively, intelligent early warning methods based on dynamic knowledge graphs and GNN inference are used. These methods construct dynamic graphs using enterprise entities, related parties, and risk events as nodes, and use inference models to simulate the propagation process to output a comprehensive risk score.
[0031] However, the aforementioned text-oriented semantic parsing methods typically process single texts, failing to perform semantic alignment and cross-document confidence weighting for risk factor descriptions with different names across multiple similar disclosures. This results in a lack of industry-level normalization and cross-sample comparability in the generated risk paths. Furthermore, methods based on knowledge graphs and model reasoning essentially depict macro-level risk contagion relationships between enterprises, lacking fine-grained modeling capabilities for the micro-level causal transmission relationships between financial and operational indicators within a single enterprise. In addition, the results produced by these methods often remain at the level of qualitative evolutionary schedules or probability scores as a black box of the model. The graph nodes lack clear quantitative rules corresponding to each other, failing to provide specific, objective, and verifiable evidence. They also fail to establish an end-to-end execution path that reverses to new reports, making it difficult to support an automated quantitative assessment closed loop for specific business texts.
[0032] To address this issue, this invention provides a risk quantification assessment method. It aims to construct a standardized directed acyclic graph (DAG) based on micro-causal links extracted from risk disclosure texts and semantic clustering. This DAG automatically generates executable logic for nodes in the graph, including data extraction and quantification judgments. This logic is then used to analyze the node states and trace the paths in the text to be assessed, achieving an end-to-end closed loop from unstructured risk knowledge mining to rule-based asset execution. This improves the objectivity, cross-sample comparability, and interpretability of complex risk identification results in business implementation. This method can be applied to automated enterprise operational risk assessment, intelligent financial report scanning, and investment and financing decision support scenarios; however, this invention does not specifically limit its application. For ease of understanding, the following embodiments are all illustrated using an automated enterprise operational risk assessment scenario as an example.
[0033] in, Figure 1 This is a flowchart illustrating the risk quantification assessment method provided by the present invention, as shown below. Figure 1 As shown, the method includes steps 110, 120, 130 and 140.
[0034] Step 110: Obtain the directed acyclic graph. The directed acyclic graph is constructed based on standardized nodes obtained by semantic clustering of risk causal links extracted from risk disclosure texts.
[0035] Here, risk disclosure text can be understood as a long document publicly disclosed by a company that includes risk statements and financial data. This can be a prospectus, a private placement prospectus, an annual report, etc. Risk disclosure text is used to characterize a company's operating conditions and potential risks at a specific time.
[0036] As an alternative embodiment, vector retrieval can be used to extract target paragraphs related to the target risk seed from the risk disclosure text, and then input into a language model to identify the micro-causal logic chain from the triggering factor to the intermediate transmission and then to the risk result as the risk causal link.
[0037] The risk causal chain contains the specific logic of risk evolution and originates from different risk disclosure texts, often exhibiting the same thing under different names. Therefore, knowledge integration can be performed based on the original risk causal chains from multiple risk disclosure texts. After identifying multiple original chains, cross-document semantic clustering is conducted to aggregate semantically equivalent but differently described original risk factors into standardized nodes. These standardized nodes can be understood as risk element units with unified industry quantitative standards and deduplicated. For example, if different reports contain multiple descriptions of increased borrowing costs, clustering and alignment yields standardized nodes for interest expense nodes.
[0038] Furthermore, a directed acyclic graph (DAG) is a graph network consisting of multiple standardized nodes connected by unidirectional edges with causal relationships, without forming closed loops. Specifically, a DAG can be constructed by statistically analyzing the frequency of the occurrence of standardized nodes in different texts, such as using the number of duplicate occurrences of adjacent nodes as associated nodes as the support strength of the directed edges.
[0039] Similarly, directed acyclic graphs (DAGs) present fine-grained risk structure relationships, and thus DAGs can also represent an overall industry-level element map for a specific risk issue.
[0040] Step 120: Based on the node information of non-terminal nodes in the directed acyclic graph, generate the quantization execution logic corresponding to the non-terminal nodes. The quantization execution logic is used to represent the logic for data extraction and quantization judgment of the text.
[0041] Specifically, non-terminating nodes can be understood as various upstream inducing nodes and intermediate transmission nodes in a directed acyclic graph, excluding the final risk conclusion node. These can be various specific financial indicator nodes or business behavior nodes. Quantitative execution logic is used to characterize the logic for data extraction and quantitative judgment of text. For example, for financial nodes that rely on fixed formula calculations, quantitative execution logic can typically be an execution rule containing a clear data source and addition, subtraction, multiplication, and division operators; for indicators that need to be compared with industry averages, quantitative execution logic can typically be a segmented mapping rule containing a clear judgment threshold.
[0042] As an alternative implementation, quantization execution logic can be obtained by inputting specific prompt word templates into a large language model. For example, the normalized name of a non-terminating node and its merged original factor list can be input into the language model to generate a decision criterion configuration suitable for that node.
[0043] Furthermore, considering that relying solely on human experience to quantify execution logic may be limited by subjective human judgment and inefficient, leading to a lack of consistency and scalability in the rules—for example, manually formulating rules is time-consuming and difficult to adapt to complex and ever-changing corporate reports on a large scale. Evaluating based solely on external fixed knowledge bases or black-box deep learning models may detach from the micro-characteristics of a company's own financial statements, resulting in warning results failing to pinpoint the change logic of specific accounts, such as failing to provide specific financial formulas to support the establishment of specific nodes.
[0044] Based on this, in order to achieve the objective and procedural conversion of rule assets, this embodiment combines the node information of non-terminal nodes to determine the quantitative execution logic, avoiding the lack of interpretability caused by relying on a simple black-box model.
[0045] As an optional embodiment, based on the node information of non-terminating nodes in the directed acyclic graph, the extraction object can be selected from a preset data source whitelist, and a ternary quantization rule containing a specific data search path, an executable quantization calculation formula, and a segmented judgment function can be determined as the quantization execution logic.
[0046] Step 130: Based on the quantization execution logic, parse the text to be evaluated to obtain the node status of non-terminating nodes.
[0047] Specifically, the text to be evaluated can be understood as the text that requires risk scanning and objective analysis, which can be the company's financial summary or annual report for the new year. Node status is used to characterize whether a specific risk factor indicator under the current text to be evaluated has been truly triggered or whether there is objective evidence. For example, for nodes where all data can be found and meet the trigger threshold, the node status is usually "established"; for nodes where financial table data has not been fully disclosed, the node status is usually "data missing" or "artificially inferred".
[0048] As an optional embodiment, the target field can be located and extracted from the text to be evaluated based on the data lookup path in the quantization execution logic, the initial calculation input of the node can be determined, and then numerical logic operations can be performed in combination with the quantization calculation formula. The result is then input into the segmentation decision function to determine the final node state.
[0049] Considering the heterogeneous layout and format of structured tables in the text to be evaluated in real-world scenarios, a single report may have multiple format forms. To avoid data extraction failure due to a single parsing tool, this embodiment preferably adopts a method of sequentially calling multiple table parsing strategies with different format rules according to a preset priority order, serially attempting and extracting the text to be evaluated, and performing multi-column cross-consistency checks on composite fields after successful extraction. Finally, the reliable data that has passed the checks is substituted into the formula to calculate the value, thereby accurately determining the node status.
[0050] Step 140: Based on the node status of each non-terminating node, trace the risk transmission path along the directed acyclic graph to obtain the risk quantification assessment result.
[0051] Specifically, risk transmission path tracing can be understood as using a directed acyclic graph as a routing network to verify whether a risk has spread downstream by sequentially tracing the causal connections between nodes. The risk quantification assessment results are used to characterize the overall risk severity and structured attribution diagnostic report for the entity corresponding to the text being assessed. For example, for a company with stable overall operations, the risk quantification assessment results typically show low-level risks and path information without any red-point triggers; for a company with multiple risk accumulation points, the risk quantification assessment results typically show high-level risks and a list of continuously triggering links colored by hot paths.
[0052] As an optional implementation, several activated end-to-end risk transmission paths can be determined by traversing from the starting point to the ending point of the directed acyclic graph based on the objective node states of each non-terminating node. Then, based on the states of the activated paths, a comprehensive risk quantification assessment result is determined.
[0053] Considering that directly applying the topology graph for traversal may be difficult to drive smoothly in actual engineering systems, in order to avoid confusion in the transmission evaluation caused by unclear logic control gateways, this embodiment preferably adopts the method of explicitly setting an XOR decision gateway between each non-terminal node and downstream node in the directed acyclic graph. If the node state of a non-terminal node is determined to be true, a trigger signal is transmitted to the corresponding downstream node through the gateway. When all nodes on the entire transmission path from start to finish are true, the path is activated. Finally, based on the number of activated risk transmission paths and the graph support weight corresponding to each path, a comprehensive risk level is determined as the final risk quantification assessment result.
[0054] The risk quantification assessment method provided in this embodiment constructs a directed acyclic graph (DAG) by semantic clustering of the risk causal links extracted from risk disclosure texts, thereby converting the risk transmission logic in unstructured natural language into a reusable knowledge graph. Quantification execution logic is generated using node information from non-terminating nodes in the DAG, avoiding the lack of interpretability caused by reliance on black-box model scoring. Furthermore, the quantitative execution logic parses the text to be assessed to determine node states, and traces the risk transmission path along the DAG to obtain the risk quantification assessment result. This achieves end-to-end traceable risk rating, improving the objectivity and cross-sample comparability of risk identification in complex business scenarios, and avoiding the business pain points of traditional risk control's excessive reliance on manual review and the difficulty in verifying judgment criteria.
[0055] Based on the above embodiments, step 120 generates quantized execution logic corresponding to the non-terminating nodes according to the node information of the non-terminating nodes in the directed acyclic graph, including: Input the normalized names of non-terminating nodes and the list of merging factors into the language model, and obtain the quantization execution logic output by the language model, which includes the data lookup path, quantization calculation formula, and segmentation decision function.
[0056] Specifically, a normalized name refers to a standardized name for risk factors with the same semantic meaning; for example, a normalized name could be "days of accounts receivable." A clustering factor list refers to the set of all original risk factor descriptions that are included in a non-terminating node during the clustering process. A language model refers to a deep learning model trained on a large amount of text that possesses the ability to understand and generate natural language.
[0057] After inputting the normalized names of non-terminating nodes and the list of merging factors into the language model, the language model outputs a quantitative execution logic that includes data search paths, quantitative calculation formulas, and segmentation judgment functions. This can transform qualitative text nodes into a closed-loop rule that integrates searching, calculation, and judgment, which can be directly called by the automated evaluator. This ensures that the data sources in the evaluation process are verifiable, the calculation process is objective and transparent, and the judgment results are mutually exclusive and clear.
[0058] Here, the data lookup path is used to indicate the specific location information of the target data in the text to be evaluated and extracted; the quantization calculation formula refers to the operator expression used to perform mathematical logic operations on the extracted data; and the piecewise decision function refers to the logic function used to map the calculation result of the quantization calculation formula to different decision states.
[0059] As an optional implementation, the quantization execution logic, which includes the data lookup path, quantization calculation formula, and segmentation decision function, can be output through a language model. In this process, the quantization execution logic corresponding to non-terminating nodes can be formally defined as R... v =(Q v F v J v ), R v Q represents the quantization execution logic corresponding to the non-termination node. v Indicates the data lookup path, F v J represents the quantitative calculation formula. v This represents the piecewise decision function.
[0060] Based on any of the above embodiments, the language model generates quantization execution logic in the following manner: including: Step 121: Using a language model, under the constraints of a preset data source whitelist, determine the data search path for the target field and mark whether the target field requires manual inference.
[0061] Specifically, the pre-defined whitelist of data sources refers to a predefined list of compliant and credible sources of supporting documents; the target field refers to the specific financial or operational data items that need to be extracted from the text to be evaluated; and human-assisted inference refers to the process of making inferences based on human experience for data that cannot be directly obtained from the text by machines or logically judged.
[0062] As an optional implementation, the language model's generation scope can be limited to a preset data source whitelist, guiding the language model to output precise drafts, classifications, and specific target fields, thus forming a data lookup path. Simultaneously, target fields lacking clear quantification conditions are marked as requiring manual inference. Specifically, the data lookup path can be represented as a set Q of data lookup paths.v Q v ={Draft q, Category q, Field q, ai q}, where q represents the draft information selected from the preset data source whitelist, q represents the data category information selected from the preset data source whitelist, q represents the target field to be extracted, and ai q This is an identifier parameter that indicates whether the target field requires manual inference; its value can be 0 or 1.
[0063] Step 122: Generate a quantitative calculation formula for performing logical operations on the target field using a language model.
[0064] Specifically, this embodiment can mathematically combine multiple interrelated target fields according to professional financial analysis logic, thereby transforming basic data into quantitative indicator values that can be directly used for evaluation and comparison.
[0065] As an optional embodiment, execution logic for processing basic data can be generated through a language model, which serves as a quantization calculation formula, and this quantization calculation formula can be expressed as F. v (x1, x2, ..., x m ), x1 to x m This represents the various target field variables extracted from the search path based on the data mentioned above.
[0066] Step 123: Based on the different calculation result intervals of the quantization calculation formula, construct a mapping relationship that maps the calculation results to true, false, and uncertain, and use the mapping relationship as a piecewise judgment function.
[0067] Here, the calculation result range refers to the pre-defined non-overlapping range of values output by the quantitative calculation formula; "valid" means that the risk of the corresponding node objectively exists and is triggered; "invalid" means that the risk status of the corresponding node is at a normal level; and "uncertain" means that a clear conclusion cannot be drawn due to the value being in a specific fuzzy range or the lack of basic target fields.
[0068] As an optional embodiment, the potential output results of the quantization calculation formula can be divided into mutually exclusive intervals, and then a mapping relationship can be constructed as a piecewise decision function. Specifically, this piecewise decision function can be formally defined as a piecewise mapping structure: If F v ∈Θyes,J v (F v )=triggered; if F v ∈Θno, J v (F v ) = normal; if F v∈Θunc or field missing, J v (F v =uncertain.
[0069] Where, triggered represents the state where the mapping is true, normal represents the state where the mapping is false, uncertain represents the state where the mapping is uncertain, Θyes represents the threshold interval corresponding to the state where the mapping is true, Θno represents the threshold interval corresponding to the state where the mapping is false, and Θunc represents the threshold interval corresponding to the state where the mapping is uncertain.
[0070] Step 124: Combine the data search path, quantization calculation formula, and segmentation judgment function to output the quantization execution logic.
[0071] Specifically, combination refers to the operation of organizing rule dimensions of different dimensions together according to a unified data structure. This involves encapsulating the determined data lookup path, the generated quantization calculation formula, and the constructed segmented decision function into a triplet rule object, which serves as the final quantization execution logic. Combined with the language model's calling logic, this process can be represented as R... v =LLM(canonical v members v D white T rule LLM stands for Language Model Call Function, canonical v Members represents the normalized name of a non-terminating node. v D represents the list of all factors for merging non-terminating nodes. white This indicates a pre-defined whitelist of data sources, T rule This indicates an additional rule generation instruction template.
[0072] Based on any of the above embodiments, step 130 parses the text to be evaluated based on the quantization execution logic to obtain the node states of non-terminating nodes, including: Step 131: Extract the target field from the text to be evaluated according to the data search path.
[0073] Specifically, the target field refers to the specific item field actually extracted from the text. This extraction process can be performed according to a pre-generated data search path, performing multi-format field extraction operations to obtain the target field from the text to be evaluated. This extraction process can be formally defined as the extraction function parsing process, specifically expressed as: x i =Extract(d label i , [parser HTML , parser MD , parserLaTeX ]); Where, x i This represents the i-th target field extracted, where Extract represents the field extraction function, and d The label represents the text to be evaluated. i The parser represents the specific tag corresponding to the target field. HTML , parser MD and parser LaTeX These represent the Hypertext Markup Language table parser, the Lightweight Markup Language pipeline table parser, and the typesetting system aligned table parser, respectively. By matching different format parsers sequentially according to priority, a match is returned, thus completing the extraction of the target field.
[0074] Step 132: Perform a consistency check on the target field.
[0075] Considering that when parsing long documents or complex tables, incomplete or erroneous data may be extracted due to formatting errors, character recognition errors, etc., directly inputting dirty data into subsequent calculations will cause systematic deviations in risk assessment. Therefore, consistency checks are performed on the target fields to filter out and intercept erroneous data that is parsed abnormally in advance.
[0076] Consistency verification refers to checking whether the extracted data conforms to established financial accounting or business logic identities. Optionally, to support the verification of structured disclosure data such as aging tables and bad debt provision tables, multi-column positioning and cross-validation rules can be configured separately for the extracted target fields. For example, the cross-validation rule for composite aging fields can be expressed as: aging over1y =Σ{i∈{1-2y, 2-3y, 3y+}}aging i ≡aging total -aging within1y ; Among them, aging over1y The target field value represents the account aging period of more than one year. i∈{1-2y, 2-3y, 3y+} represents the sub-ranges of 1 to 2 years, 2 to 3 years, and more than 3 years. i This represents the aging value for the corresponding range. total This represents the total amount of accounts receivable aging. within1y This represents the value for accounts receivable within one year. If the difference calculated from both ends of the above formula is less than the preset tolerance, denoted as ε, then the target field is deemed to have passed the consistency check, and the extraction result is considered reliable.
[0077] Step 133: Substitute the verified target fields into the quantization calculation formula for calculation, and input the calculation result into the segmentation judgment function to obtain the node state.
[0078] As an optional implementation, quantization execution logic can be invoked to perform evaluation operations on the field set. This node-level state determination process can be represented as follows: S(v)=J v (F v (Q v (X(d )))); Where S(v) represents the node state of a non-terminating node v, J v Let F represent the piecewise decision function. v This represents the quantitative calculation formula, Q. v X(d) represents the set of data lookup paths. ) indicates that from the text to be evaluated d Extract and validate the set of fields that pass the above evaluation. After the above evaluation calculation, the final output value of the node state S(v) belongs to the preset state set, S(v)∈{triggered, normal, ai} required no data}, triggered indicates a triggered state, normal indicates a normal false state, ai required This indicates a state where manual inference is required, no. data This indicates a missing data state.
[0079] Based on any of the above embodiments, step 131 extracts the target field from the text to be evaluated according to the data search path, including: Step 1311: According to the preset priority order, use a variety of table parsing strategies with different format rules to parse the text to be evaluated.
[0080] Here, the preset priority order refers to the order in which the pre-arranged parsing algorithms are called, different format rules refer to the structural recognition logic set for different typesetting features, and table parsing strategy refers to the specific algorithmic means of extracting table structure data from the document.
[0081] The priority order of parsers can be set based on the common format distribution of the converted document. For example, after the source documents such as annual reports are converted to other formats, the same report may contain three forms: Hypertext Markup Language tables, Lightweight Markup Language pipeline tables, and typesetting system aligned tables. Therefore, a multi-parser chaining strategy can be constructed to try parsing these three table parsing strategies with different formatting rules in sequence according to a pre-set priority order.
[0082] Step 1312: If the current table parsing strategy successfully matches the data corresponding to the target field, output the target field and stop parsing.
[0083] Specifically, if the current table parsing strategy successfully matches the data corresponding to the target field, it indicates that the required value has been obtained. To confirm the parsing conclusion, in this case, this embodiment outputs the target field and stops parsing, ensuring the efficiency of the extraction process and avoiding redundant interference caused by multiple parsings. Here, the current table parsing strategy refers to the specific parsing algorithm being executed in the priority sequence, and matching refers to the process of successfully locating and identifying the required data in the text. Optionally, when the highest-priority Hypertext Markup Language table parser is called, once the data corresponding to the target field is successfully identified and matched at the corresponding position in the text to be evaluated, the extraction is confirmed to have hit, and the target field is directly output for use by downstream processes, while the subsequent call process of other parsers is stopped.
[0084] Step 1313: If no match is found, switch to the next priority table parsing strategy to continue parsing.
[0085] Specifically, if a match is not found, it indicates that the current strategy has failed. To implement a fallback mechanism, this embodiment switches to the next priority table parsing strategy to continue parsing. Here, the next priority table parsing strategy refers to the backup parsing algorithm that follows the current strategy in a preset sequence. As an optional embodiment, if a match is not found in the current paragraph of the text to be evaluated using the preferred Hypertext Markup Language table parser, a fallback process is performed, switching to the next priority Lightweight Markup Language Pipeline table parser to continue parsing; if a match is still not found, it further switches to the typesetting system aligned table parser to continue trying.
[0086] Based on any of the above embodiments, consistency verification of the target field includes: If the target field is a composite field containing the overall field and multiple subfields, then calculate the sum of the values of each subfield; If the difference between the sum of the values and the overall field value is less than the preset tolerance, then the target field is determined to have passed the consistency check.
[0087] Here, a composite field refers to a set of fields consisting of a single overall field and multiple subfields. The overall field represents the total or sum of values within the composite field, while the subfields are the individual detailed breakdown items that make up the overall field. For example, when the target field to be extracted is structured disclosure data such as an aging table or a bad debt provision multi-list, it is identified as a composite field containing an overall field and multiple subfields. For this composite field, all its subfields are extracted individually, and the sum of the values of these subfields is calculated.
[0088] Furthermore, if the difference between the sum of the extracted values and the overall field value is less than the preset tolerance, it indicates that the extracted subfields conform to the inherent preset logic with the overall field, and the data parsing results are accurate and reliable. At this point, it can be determined that the target field has passed the consistency check. The preset tolerance refers to a pre-set threshold for the allowable error range.
[0089] Based on any of the above embodiments, the text to be evaluated is parsed based on the quantization execution logic to obtain the node states of non-terminating nodes, including: Based on the quantitative execution logic, the text to be evaluated is parsed to determine the extraction results of the target fields from the text to be evaluated; If the extraction result is that the target field was not extracted, the node status will be marked as missing data. If the extraction result is the target field, and the data search path carries a marker that requires manual inference, then the node status is marked as pending inference.
[0090] Specifically, the extraction result refers to the status information obtained after performing data addressing and extraction operations on the text to be evaluated based on the quantitative execution logic. It usually includes two cases: successfully obtaining the target field or not finding the corresponding target field.
[0091] If the extraction result indicates that the target field was not extracted, it means that the relevant structured fields were not disclosed in the text to be evaluated, resulting in a lack of underlying objective data for subsequent quantitative calculations at that node. Consequently, the node status is marked as missing data. The missing data status corresponds to the "no source" status in the evaluation system, representing a data gap that cannot support subsequent formula calculations. Optionally, when the extraction result during the data extraction stage indicates that the target field was not extracted, the node status of the corresponding non-terminating node is marked as missing data. In the subsequently generated visual diagnostic report, nodes marked as missing data can be clearly identified with contrasting color blocks, such as gray blocks, to facilitate reviewers' intuitive identification of the missing risk verification steps in the current report.
[0092] If the extraction result shows that the target field has been extracted, and the data search path carries a marker indicating that manual inference is required, it indicates that although relevant information has been successfully obtained, the target field is a non-standardized expression or lacks clear direct quantification conditions. Relying solely on machine-defined arithmetic formulas and thresholds is insufficient for accurate judgment, thus marking the node state as a state awaiting inference. Here, the marker indicating that manual inference is required refers to a special identifier set for certain fields during the quantification execution logic generation stage, signifying that the field requires qualitative judgment using natural language reasoning or human intervention. A state awaiting inference refers to an intermediate node state in the evaluation system that explicitly indicates the need for external experience intervention for state confirmation.
[0093] As an optional implementation, when the extraction result is a target field, the parameter attributes in its corresponding data lookup path are further examined. If the data lookup path carries a marker marked 1, indicating that manual inference is required, the target field is not forcibly substituted into a regular arithmetic formula. Instead, the node state of the non-terminal node is marked as a state to be inferred. Similarly, in the visual evaluation display, nodes in the state to be inferred can be clearly marked with a different color block, such as an amber block, thereby actively exposing the known boundaries in the evaluation process and forming explicit human-machine collaborative nodes.
[0094] Based on any of the above embodiments, step 140 involves tracing the risk transmission path along the directed acyclic graph according to the node states of each non-terminating node to obtain a risk quantification assessment result, including: Step 141: Set up an XOR decision gateway between each non-terminating node and its downstream node in the directed acyclic graph.
[0095] Here, a downstream node refers to the next level node that receives the signal in the causal direction, and an XOR decision gateway is a single control unit used to determine whether to allow the signal to be passed to the downstream branch based on the current node's state. Specifically, an XOR decision gateway can be inserted to the right of each non-terminating node in the merged directed acyclic graph (DAG), forming a routing-based DAG.
[0096] As an optional embodiment, for each non-terminating node in the weighted directed acyclic graph, an XOR decision gateway is inserted between it and the downstream node to form a routing directed acyclic graph. The initial weighted directed acyclic graph can be represented as G=(V,E,w), where G represents the weighted directed acyclic graph, V represents the normalized node set, E represents the weighted directed edge set, and w represents the support strength corresponding to the edge.
[0097] For each non-terminating node v in the weighted directed acyclic graph G, insert an XOR decision gateway between it and its downstream nodes. v The formal definition of this XOR decision gateway can be expressed as: XOR v J v (F v (X))∈{triggered, normal, uncertain}; [YES] The branch propagates to the downstream node Out(v) of v; [NO] Branch points to safe termination box Stop v .
[0098] Among them, XOR vThis represents the XOR decision gateway corresponding to the non-terminating node v, X represents the set of extracted data fields, Out(v) represents the set of all downstream adjacent nodes of the non-terminating node v, and Stop... v The safety termination box indicates that conduction has stopped.
[0099] Step 142: If the node status of a non-terminating node is true, then the result is transmitted to the corresponding downstream node via the XOR decision gateway.
[0100] After setting up the XOR decision gateway, if the node status of a non-terminating node is true, it indicates that the risk factor indicator corresponding to the non-terminating node has a real objective trigger basis. In order to realistically simulate the process of risk spreading backward along the causal chain in real business, under this case, it is transmitted to the corresponding downstream node through the XOR decision gateway.
[0101] As an optional embodiment, the XOR decision gateway configured above is used to perform routing scheduling on the node status of non-terminating nodes. If the node status of a non-terminating node is valid, that is, the status is equivalent to triggered, the corresponding YES branch is triggered, and the risk signal is smoothly transmitted to the corresponding downstream node through the XOR decision gateway. Conversely, if the node status is invalid, the NO branch is selected, the transmission is cut off, and it no longer spreads downstream.
[0102] Step 143: When the node states of all nodes on the propagation path from the trigger node to the termination node are valid, activate the corresponding risk propagation path.
[0103] After simulating signal transmission between nodes, considering that a local triggering of a single node does not equate to the occurrence of the final risk outcome, only a complete causal chain can confirm the actual implementation of the risk. Therefore, when the node states of all nodes along the transmission path from the triggering node to the terminating node are valid, the corresponding risk transmission path is activated. This ensures that the end-to-end risk conclusion no longer depends on a single node mutation, but is based on the continuity of global structural evidence, possessing rigorous logical traceability. Here, the triggering node refers to the initial node at the forefront of the graph that triggers subsequent reactions; the terminating node refers to the final node representing the final target risk issue in the graph; the transmission path refers to an end-to-end sequence of lines connecting nodes at all levels; and activating the corresponding risk transmission path means recognizing that the causal chain has actually occurred in the current text to be evaluated and presents a warning response state.
[0104] Specifically, the activation condition can be expressed as: like v i When ∈π\{r}, S(v i )=triggered, π=(v1, v2,..., v q ,r),A(π)=1; Otherwise, A(π) = 0.
[0105] A(π) represents the path activation function, where π represents the end-to-end propagation path ending at node r; 1 indicates activation of the corresponding risk propagation path, and 0 indicates deactivation of the corresponding risk propagation path. i Let S(v) represent the i-th non-terminating node on the propagation path π, and r represent the terminating node. i ) represents node v i The node states are defined by the term "triggered", which indicates a valid state. Further traversing all propagation paths yields the set of active paths Πactive={π|A(π)=1}, where Πactive represents the set of active paths comprised of all activated risk propagation paths.
[0106] Thus, each non-terminating node v obtains a rule object R that can be directly consumed by the automated evaluator. v The terminating node r does not invoke the language model, and its decision function is fixed as "it is true if triggered by any upstream path", formally defined as follows: J r =triggered, if ∃π∈Π(r): π has been activated; Among them, J r Let r be the decision function for the termination node r, triggered indicates the state of triggering, ∃ indicates the existence quantifier (i.e., at least one exists), π indicates the end-to-end risk transmission path, and Π(r) is the set of all end-to-end transmission paths ending at r.
[0107] As an optional implementation, the final state of the termination node is determined based on the number of paths contained in the active path set and the verification status of each node. The specific judgment logic can be expressed as follows: If |Πactive|≥1, S(r)=triggered; If |Πactive|=0 and v:S(v)≠uncertain, S(r)=normal; Otherwise, S(r) = uncertain.
[0108] Where S(r) represents the final state of the terminating node r, |Πactive| represents the number of paths contained in the active path set, v represents a non-terminating node, and S(v) represents the node state of node v.
[0109] Step 144: Based on the number of activated risk transmission paths and the supporting weights of each path, determine the overall risk level of the text to be evaluated, which will serve as the result of the risk quantification assessment.
[0110] After identifying the activated risk transmission paths, considering the varying contributions and credibility of different paths to the final risk conclusion, and to output comparable results with a unified standard, the overall risk level of the text to be evaluated is determined by comprehensively considering the number of activated risk transmission paths and the corresponding support weights of each path. This serves as the result of the risk quantification assessment, mapping the complex multi-dimensional graph features into an intuitive business classification, thus meeting the macro-level quantitative control needs of different application scenarios. Here, support weights refer to the edge weight values representing the strength of cross-document node associations, statistically obtained during the graph construction phase, while the overall risk level refers to the final comprehensive risk severity level of the corresponding entity.
[0111] The risk level mapping mechanism can be represented as follows: W t =Σ{v:S(v)=triggered}w in (v), w in (v)=Σ{u:(u,v)∈E}; n t =|{v:S(v)=triggered}|; L(d If S(r) = HIGH, and S(r) = triggered and (n) t ≥θ h or W t ≥ω h ); L(d If S(r) = MEDIUM and the HIGH condition is not satisfied; L(d )=LOW, if S(r)=normal.
[0112] Among them, W t Let S(v) represent the weighted confidence level, and w represent the node state. in (v) represents the sum of the weights of the incoming edges supporting a non-terminating node v, which is equal to the sum of the weights of all incoming edges pointing to v; n t L(d) represents the number of triggering nodes, used to measure the scale of nodes triggered by the number of activated risk transmission paths; ) indicates the text to be evaluated, d. The overall risk level; HIGH represents high risk, MEDIUM represents medium risk, and LOW represents low risk; S(r) represents the final state of the terminating node r, which is triggered when there is at least one active path; θ h ω represents the preset trigger threshold. hThis indicates a preset support weight threshold; "normal" indicates a false condition. Therefore, the overall risk level is output as the risk quantification assessment result. Simultaneously, the trigger nodes and edges can be colored with hot paths in the corresponding display interface to generate interpretable structured diagnostic information.
[0113] Based on any of the above embodiments, the steps for extracting risk causal links include: Step 210: Based on the given target risk seed, recall the target paragraphs that are semantically related to the target risk seed from the risk disclosure text.
[0114] Here, the target risk seed refers to a pre-specified target risk issue, such as the risk of bad debt recovery of accounts receivable, and the target paragraph refers to a text block in the text that has a high degree of relevance to the target risk seed in the semantic vector space. The entire risk disclosure text can be divided into a paragraph set P(d) by paragraphs or semantic blocks. k )={p1, p2, ..., p M Then, calculate the semantic relevance of each paragraph to the target risk seed. This semantic relevance calculation method can be expressed as: sim(p j ,r)=cos(Emb(p j ),Emb(r)); In the above formula, sim(p j r) represents the j-th paragraph p j The semantic relevance to the target risk seed r is defined by Emb, which represents the pre-trained vector encoding function. By retaining the top K paragraphs with the highest semantic relevance, the function can accurately recall target paragraphs that are semantically related to the target risk seed.
[0115] Optionally, a concurrent or serial processing flow is initiated for the input document library, sequentially performing the aforementioned paragraph semantic recall and structured extraction operations based on a large language model on all K texts to be processed, ultimately obtaining the complete set of original causal links. This convergence process can be expressed as the formula: C = ⋃{k=1}^{K}C k ; In the above formula, C represents the total set of all original causal links obtained by summarizing, k represents the sequence index of the text to be processed, and K represents the total number of texts to be processed, i.e., the total number of texts in the document library. k This represents the set of all original causal links extracted from the k-th text to be processed.
[0116] Step 220: Input the target paragraph and target risk seed into the language model, extract the risk factors contained in the target paragraph and the causal transmission relationship between each risk factor, and obtain multiple step stages with sequential orientation as risk causal links.
[0117] Specifically, risk factors refer to specific business variables or operational events that lead to or are affected by risks; causal transmission relationships refer to the logical chains of triggering, intermediate evolution, or resulting outcomes among different risk factors; and steps and stages refer to the specific independent nodes in the causal chain and their corresponding role attributes.
[0118] As an optional embodiment, the recalled target paragraphs and target risk seeds are input together as a joint context into the language model, along with a structured extraction instruction template. This requires the language model to identify the complete link from the triggering factor to the risk conclusion. This structured extraction process can be represented as follows: C k =LLM(R(d k ,r),r,T chain ); In the above formula, C k ={c k1 c k2 c kn}, C k This indicates that from the k-th risk disclosure text d k The complete set of original causal links extracted from R(d), where LLM represents the language model call function, and R(d k ,r) indicates from the risk disclosure text d k The set of target segments to be recalled, r represents the target risk seed, T chain This represents a structured extraction instruction template.
[0119] Furthermore, each extracted risk causal link consists of multiple sequentially pointing steps, and its formal definition can be expressed as: c ki =⟨s1, s2, ..., s L >; In the above formula, c ki Represents set C k The i-th risk causal link in the equation, ⟨s1, s2, ..., s L > represents an ordered sequence consisting of steps 1 to L, s L This represents the Lth step stage. Each step stage includes a description of the corresponding risk factor, a role value, and a supporting text excerpt for back-end tracing. The role value is selected from three options: trigger, intermediate, and conclusion.
[0120] To ensure the integrity of causal transmission, the constraints can be set as follows: the role of the first step in each risk causal link is anchored as the trigger, and the role of the last step is anchored as the conclusion, with its risk factor description strictly aligned to the target risk seed r. Thus, the language model escapes unconstrained generation and outputs machine-readable causal links according to the aforementioned three-stage paradigm, obtaining multiple step stages with sequential pointing as risk causal links.
[0121] Based on any of the above embodiments, step 210, based on a given target risk seed, retrieves semantically relevant target paragraphs from the risk disclosure text, including: Step 211: Divide the risk disclosure text into multiple candidate paragraphs.
[0122] Specifically, candidate paragraphs refer to text blocks with relatively independent semantics obtained after segmenting a long document. The entire risk disclosure text can be segmented into natural paragraphs or semantic blocks to obtain a set of paragraphs containing multiple candidate paragraphs. This set of paragraphs can be represented as P(d k )={p1, p2, ..., p M}, P(d k ) indicates the risk disclosure text d k The corresponding set of paragraphs, p1-p M These represent the first to the Mth candidate segments obtained from the segmentation.
[0123] Step 212: Map each candidate paragraph and target risk seed into a semantic vector.
[0124] Here, semantic vectors refer to high-dimensional floating-point arrays output by a pre-trained encoding model that represent the deep semantic information of the text. Each candidate paragraph and the target risk seed can be encoded separately using a pre-trained vector encoding function; this process can be represented as Emb(p j And Emb(r), where Emb represents the pre-trained vector encoding function, p j Let r represent the j-th candidate paragraph, and let Emb(p) represent the target risk seed. j ) indicates candidate paragraph p j The corresponding semantic vector, Emb(r), represents the semantic vector corresponding to the target risk seed r.
[0125] Step 213: Calculate the spatial similarity between the semantic vector of each candidate paragraph and the semantic vector of the target risk seed.
[0126] Here, spatial similarity refers to a numerical metric used in a high-dimensional vector space to measure how close two vectors are in direction or distance. Specifically, cosine similarity can be used to assess the correlation between candidate segments and target risk seeds.
[0127] Step 214: Retain candidate paragraphs whose spatial similarity meets the preset sorting conditions as target paragraphs.
[0128] After calculating the spatial similarity of each candidate paragraph, considering that extracting all candidate paragraphs would lead to a waste of subsequent computing resources and contain a large amount of irrelevant information, it is necessary to filter redundant data and retain candidate paragraphs whose spatial similarity meets the preset sorting conditions as target paragraphs.
[0129] Among them, the preset sorting conditions refer to the sorting thresholds or ranking rules set in advance for truncating and filtering qualified results in the similarity inverted list, and the target paragraph refers to the related text that is finally selected after filtering for downstream causal chain extraction.
[0130] As an optional embodiment, the spatial similarity of all candidate paragraphs can be sorted in descending order, and the top K candidate paragraphs in terms of spatial similarity can be retained as the recall paragraph set, i.e., the target paragraphs. Here, K is a preset recall threshold. This recall paragraph set can be represented as R(d k ,r), R represents the recall function, d k This represents the risk disclosure text, and 'r' represents the target risk seed.
[0131] Based on any of the above embodiments, the multiple steps and stages correspond to different role types, including triggering factors, intermediate transmission, and risk outcomes. Extract the risk factors contained in the target paragraph and the causal transmission relationships between these risk factors to obtain multiple sequentially pointing steps as the risk causal chain, including: Extract the risk factors contained in the target paragraph and the causal transmission relationship between the risk factors to generate candidate links containing multiple step stages with sequential orientation; In the candidate links, the candidate links in which the role type of the first step is the triggering factor, the role type of the last step is the risk outcome, and the candidate links are consistent with the target risk seed are regarded as risk causal links.
[0132] Here, candidate links refer to initially generated directed sequences with sequential orientations that have not yet undergone rigorous finality filtering. Multiple steps correspond to different role types, including triggering factors, intermediate transmissions, and risk outcomes. As an optional implementation, the target paragraph and a structured extraction instruction template can be input into a language model to generate candidate links containing multiple sequential steps. These candidate links can be formally defined as c... ki =⟨s1, s2, ..., s L >,c kiLet ⟨s1, s2, ..., s be the candidate link extracted from the k-th text. L > represents an ordered sequence consisting of steps 1 to L, s1-s L These represent specific steps or stages in an ordered sequence. Furthermore, each step or stage can be represented as s. l =(factor l role l evidence l ), s l Indicates the l-th step stage, factor l This indicates the description of the risk factors corresponding to this step / stage. l This indicates the role type corresponding to this step, and its value is selected from one of three items: triggering factor, intermediate transmission, risk outcome, or evidence. l This indicates a passage from the original text that supports the factor, used for reverse tracing.
[0133] After generating candidate links containing multiple steps with sequential directions, considering that the language model may produce invalid links lacking a source or deviating from the target during the free generation process, if they are not truncated and filtered, a large amount of irrelevant noise will be introduced into the downstream graph construction stage. Therefore, in the candidate links, the candidate links whose role type in the first step is the triggering factor and whose role type in the last step is the risk result, and which are consistent with the target risk seed, are regarded as risk causal links. This ensures that the links that are ultimately retained are all risk inducements at the logical starting point and are all anchored to the preset risk issues at the logical ending point, thereby improving the quality of the extracted knowledge.
[0134] In this context, the first step stage refers to the first node in the candidate link sequence, and the last step stage refers to the last node in the candidate link sequence. Role type refers to the logical role that a step stage plays in the causal chain, triggering factor refers to the source cause that leads to a series of chain reactions, and risk outcome refers to the negative state that the evolutionary chain ultimately results in.
[0135] As an optional implementation, a traversal check can be performed on all generated candidate links, imposing structured constraints. Specifically, the constraints are: verifying the first step of the candidate link, requiring its role type to be assigned as a triggering factor; verifying the last step of the candidate link, requiring its role type to be assigned as a risk result, and further verifying the risk factor description in the last step to ensure it is semantically anchored as a target risk seed. Candidate links that satisfy all the above constraints are retained and confirmed as risk causal links.
[0136] Based on any of the above embodiments, the steps for obtaining a directed acyclic graph include: Step 310: Perform semantic equivalence aggregation on the original risk factors in the risk causal chain to form standardized nodes, and discard the original risk factors that cannot be quantified.
[0137] Specifically, semantic equivalence aggregation refers to the operation of grouping descriptions that differ in expression but inherently refer to the same fact or indicator into the same category. Unquantifiable original risk factors refer to descriptions that lack clear numerical indicators or financial formulas, making automated judgment difficult.
[0138] As an optional implementation, all original risk factors of non-conclusion roles are extracted from the entire risk causal chain and deduplicated to obtain the original factor set, which can be represented as: F={factor l | c∈C, s l ∈c, role l ≠ result}; Where F represents the original set of factors, factor l This represents the risk factor description corresponding to the l-th step stage, c represents the risk causal link, C represents the set of all original causal links, and s l The role represents a step or stage in the link. l This indicates the role type corresponding to the step stage, and result indicates the risk outcome role.
[0139] Furthermore, the original set of factors is input into the language model, and the constraint output is several cluster groups. This process can be represented as: G=LLM(F,T cluster )={g1, g2, ..., g N}; g n =⟨canonical n formula n members n >; Where G represents the cluster set, LLM represents the language model call function, and T cluster Represents a clustering instruction template, g n Canonical represents the nth cluster group, where N represents the total number of cluster groups. n The formula represents the normalization metric name corresponding to the nth cluster group. n This indicates the standard calculation formula for this indicator, members n This represents the original list of factors that were clustered into this group.
[0140] During this process, the following constraints are applied: unquantifiable raw factors are discarded and do not enter any cluster group; different financial indicators cannot be merged; different directional descriptions of the same indicator are all grouped into the same group. Subsequently, the factors in each link are replaced with normalized nodes through a mapping function. For nodes mapped to unquantifiable identifiers, the rule of truncating or discarding the entire link is executed, and adjacency deduplication is performed to obtain normalized links. This mapping process can be represented as: v l =(r, if role) l ==result, otherwise it is (factor l )); Among them, v l Let r represent the l-th normalized node after mapping, and r represent the target risk seed. l The original role is represented by "result," and the conclusion role is represented by "result." The factor represents the mapping function from the original factors to the standardized nodes. l This indicates a description of the original factors.
[0141] After performing the above clustering and constraints, the mapping function from the original factors to the standardized nodes is obtained, which can be expressed as: :F→V∪{⊥}; Where V = {canonical1, ..., canonical} N}∪{r}, This represents the mapping function from the original factors to the standardized nodes, where F represents the set of original factors, V represents the set of standardized nodes, and ⊥ represents the unquantifiable and discarded label of the original risk factor. 1和 canonical N represents the normalized index name corresponding to the 1st and Nth cluster groups respectively, and r represents the target risk seed.
[0142] Step 320: Count the number of times adjacent standardized nodes appear as related node pairs in different texts to remove duplicates.
[0143] Here, adjacent normalized nodes refer to two nodes that are closely connected in the same normalized link, associated node pairs refer to a combination of two nodes that have a sequential pointing relationship, and deduplication count refers to the frequency calculated after removing duplicate occurrences in different document dimensions.
[0144] As an optional implementation, traversing all normalized links and counting the number of times each adjacent node pair appears as a deduplicated pair of different document identifiers and link identifiers across the document dimension can be represented as: w(u,v)=|{(k,i)|(u,v)∈Adj(cmber ki )}|; Where w(u, v) represents the number of times nodes u and v are deduplicated, u and v represent two adjacent normalized nodes, (k, i) represents the i-th link identifier combination in the k-th text, and Adj(c̃) ki ) represents a normalized link c̃ ki The set of all adjacent node pairs in the array.
[0145] Step 330: Using the number of deduplications as the support strength of the corresponding directed edges, construct a directed acyclic graph with the target risk seed as the endpoint.
[0146] This embodiment uses the number of deduplication steps as the support strength of the corresponding directed edges to construct a directed acyclic graph (DAG) with the target risk seed as the endpoint. This enables the disparate microscopic transmission chains to be woven into a global knowledge network. Here, support strength refers to a numerical indicator reflecting the reliability of causal transmission between related node pairs, and a directed acyclic graph refers to a network structure composed of nodes and directed edges that does not contain loops.
[0147] As an optional embodiment, a weighted directed network is constructed based on the standardized nodes and the support strength between them. The directed acyclic graph can be represented as G=(V,E,w), where G represents the directed acyclic graph, V represents the set of standardized nodes, r∈V and is the unique terminal node, E={(u,v)|w(u,v)>0} is the set of weighted directed edges, and w represents the support strength corresponding to the edge.
[0148] Furthermore, the normalized propagation probability of an edge is defined, reflecting the proportion of all links originating from node u that propagate to downstream node v. The calculation formula can be expressed as: P(v|u)=w(u,v) / Σ{v'∈Out(u)}w(u,v'); Where P(v|u) represents the normalized propagation probability, w(u,v) represents the support strength of node u pointing to node v, v' represents each node in the set Out(u) of all downstream adjacent nodes belonging to node u, and w(u,v') represents the support strength of node u pointing to each downstream node v'.
[0149] Figure 2 This is a flowchart illustrating the directed acyclic graph construction method provided by the present invention, as shown below. Figure 2As shown, the system receives the original risk causal links from different texts (such as document d1, document d2, and document d3). For example, the original risk causal link of document d1 includes "the proportion of 1-2 year aging accounts increases", the original risk causal link of document d2 includes "the aging structure deteriorates", and the original risk causal link of document d3 includes "the aging of accounts extends". Although these three statements are different, they all further point to "the increase in bad debt provisions" and ultimately transmit to the target risk seed, i.e., the endpoint r.
[0150] Subsequently, a semantic clustering mapping function is executed using a language model to perform semantic equivalence aggregation on the three original risk factors with different expressions, mapping them uniformly to the same node v2. Next, a link normalization operation is performed, mapping the three original risk causal links to normalized links with the same structure, i.e., v2→v3 (normalized node)→r. Based on this, a cross-document deduplication count operation is performed to count the frequency of adjacent normalized node pairs as associated node pairs in different texts.
[0151] Finally, the calculated number of deduplications is used as the support strength of the corresponding directed edges to construct the weighted directed acyclic graph at the bottom. Specifically, v2 (long-term accounts receivable ratio) is transmitted to the standardized node v3 (bad debt provision balance) through the directed edge with support strength w=3, and v3 is then transmitted to the final target risk seed r through the directed edge with support strength w=3.
[0152] Figure 3 This is a flowchart illustrating another risk quantification assessment method provided by the present invention, as shown below. Figure 3 As shown, firstly, a set of risk disclosure texts to be processed, D={d1, d2, ..., d...}, is received. K}, and the target risk seed r.
[0153] For each pending risk disclosure text d K Based on a given target risk seed, each target paragraph R(d) is retrieved from the risk disclosure text that is semantically related to the target risk seed. K The target paragraph and target risk seed are input into the LLM, and the risk factors contained in the target paragraph and the causal transmission relationship between each risk factor are extracted to obtain multiple step stages with sequential orientation as the risk causal link C. k .
[0154] Perform cross-document semantic merging on all risk causal links from all texts to be processed: C=∪ k C kThe original risk factors in the risk causal chain are semantically equivalently aggregated to form standardized nodes V. The number of times adjacent standardized nodes appear as related node pairs in different texts is counted and deduplicated. The number of deduplications is used as the support strength of the corresponding directed edges to construct a directed acyclic graph (DAG) G with the target risk seed as the endpoint.
[0155] For each non-terminal node in the directed acyclic graph, generate: v∈V{r}, using LLM to automatically generate quantized execution logic (Q) v F v J v And set up an XOR decision gateway between each non-terminating node and the downstream node in the directed acyclic graph to form a routing directed acyclic graph (routing DAG).
[0156] Receive new text to be evaluated (d) ), from which the target field (X(d)) required for the quantization execution logic is extracted. The algorithm parses the text to be evaluated based on the quantitative execution logic to obtain the node state (S(v)) of the non-terminating node. When the node states of all nodes on the transmission path from the trigger node to the termination node are valid, the corresponding risk transmission path is activated along the routing directed acyclic graph. The risk level L of the text to be evaluated is determined by combining the number of activated risk transmission paths and the support weights of each path, which is used as the risk quantification evaluation result. Finally, the end-to-end interpretable risk level determination result is output.
[0157] The risk quantification assessment device provided by the present invention is described below. The risk quantification assessment device described below and the risk quantification assessment method described above can be referred to in correspondence.
[0158] Based on any of the above embodiments Figure 4 This is a schematic diagram of the risk quantification assessment device provided by the present invention, as shown below. Figure 4 As shown, the device includes: The acquisition module 410 is used to acquire a directed acyclic graph, which is constructed based on standardized nodes obtained by semantic clustering of risk causal links extracted from risk disclosure texts. The generation module 420 is used to generate the quantization execution logic corresponding to the non-terminal node based on the node information of the non-terminal node in the directed acyclic graph. The quantization execution logic is used to represent the logic of data extraction and quantization judgment of the text. The parsing module 430 is used to parse the text to be evaluated based on the quantization execution logic to obtain the node status of non-terminal nodes. The evaluation module 440 is used to trace the risk transmission path along the directed acyclic graph based on the node status of each non-terminating node, and obtain the risk quantification evaluation result.
[0159] Figure 5 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 5 As shown, the electronic device may include a processor 510, a communications interface 520, a memory 530, and a communication bus 540, wherein the processor 510, the communications interface 520, and the memory 530 communicate with each other through the communication bus 540. The processor 510 can call logical instructions in the memory 530 to execute a risk quantification assessment method.
[0160] Furthermore, the logical instructions in the aforementioned memory 530 can be implemented as software functional units and, when sold or used as independent products, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, essentially, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0161] On the other hand, the present invention also provides a computer program product, which includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer is able to perform the risk quantification assessment methods provided by the above methods.
[0162] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the risk quantification assessment methods provided by the methods described above.
[0163] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0164] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0165] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for quantifying and assessing risk, characterized in that, include: A directed acyclic graph is obtained, which is constructed based on standardized nodes obtained by semantic clustering of risk causal links extracted from risk disclosure texts; Based on the node information of the non-terminating node in the directed acyclic graph, the quantization execution logic corresponding to the non-terminating node is generated. The quantization execution logic is used to characterize the logic for data extraction and quantization judgment of the text. Based on the quantization execution logic, the text to be evaluated is parsed to obtain the node state of the non-termination node; Based on the node status of each non-terminating node, the risk transmission path is traced along the directed acyclic graph to obtain the risk quantification assessment result.
2. The risk quantification assessment method according to claim 1, characterized in that, The step of generating quantized execution logic corresponding to the non-terminal node based on the node information of the non-terminal node in the directed acyclic graph includes: Input the normalized name of the non-terminating node and the list of merging factors into the language model, and obtain the quantization execution logic output by the language model, which includes the data lookup path, quantization calculation formula and segmentation decision function.
3. The method according to claim 2, characterized in that, The language model generates the quantization execution logic based on the following method: Using the language model, under the constraints of a preset data source whitelist, the data search path for extracting the target field is determined, and it is marked whether the target field requires manual inference. The language model is used to generate the quantization calculation formula for performing logical operations on the target field; Based on the different calculation result ranges of the quantization calculation formula, a mapping relationship is constructed to map the calculation results to true, false, and uncertain, respectively, and the mapping relationship is used as the segmentation judgment function; The data lookup path, the quantization calculation formula, and the segmentation judgment function are combined to output the quantization execution logic.
4. The method according to claim 3, characterized in that, The step of parsing the text to be evaluated based on the quantization execution logic to obtain the node state of the non-terminating node includes: Extract the target field from the text to be evaluated according to the data search path; Perform a consistency check on the target field; The target field that has passed the verification is substituted into the quantization calculation formula for calculation, and the calculation result is input into the segmentation judgment function to obtain the node state.
5. The method according to claim 4, characterized in that, The step of extracting the target field from the text to be evaluated according to the data search path includes: According to a preset priority order, multiple table parsing strategies with different formatting rules are used sequentially to parse the text to be evaluated; If the current table parsing strategy successfully matches the data corresponding to the target field, then the target field is output and parsing stops; If no match is found, the table parsing strategy of the next lower priority will be switched to continue parsing.
6. The method according to claim 4, characterized in that, The consistency check of the target field includes: If the target field is a composite field containing a whole field and multiple subfields, then calculate the sum of the values of each subfield; If the difference between the sum of the numerical values and the value of the overall field is less than the preset tolerance, then the target field is determined to have passed the consistency check.
7. The method according to claim 6, characterized in that, The step of parsing the text to be evaluated based on the quantization execution logic to obtain the node state of the non-terminating node includes: Based on the quantization execution logic, the text to be evaluated is parsed to determine the extraction result of the target field from the text to be evaluated; If the extraction result is that the target field is not extracted, then the node status is marked as a data missing status; If the extraction result is the extraction of the target field, and the data search path carries a marker that requires manual inference, then the node status is marked as pending inference.
8. The method according to any one of claims 1 to 7, characterized in that, The step of tracing the risk transmission path along the directed acyclic graph based on the node states of each non-terminating node to obtain the risk quantification assessment result includes: An XOR decision gateway is set between each non-terminating node and its downstream node in the directed acyclic graph. If the node state of the non-terminating node is true, it is transmitted to the corresponding downstream node via the XOR decision gateway. When the node states of all nodes on the propagation path from the trigger node to the termination node are valid, the corresponding risk propagation path is activated. The overall risk level of the text to be evaluated is determined by combining the number of activated risk transmission paths and the supporting weights of each path, and is used as the result of the risk quantification assessment.
9. The method according to any one of claims 1 to 7, characterized in that, The steps for extracting the risk causal link include: Based on a given target risk seed, recall target paragraphs from the risk disclosure text that are semantically related to the target risk seed; The target paragraph and the target risk seed are input into the language model, and the risk factors contained in the target paragraph and the causal transmission relationship between each risk factor are extracted to obtain multiple step stages with sequential orientation as the risk causal link.
10. The method according to claim 9, characterized in that, The step of recalling semantically relevant target paragraphs from the risk disclosure text based on a given target risk seed includes: The risk disclosure text is divided into multiple candidate paragraphs; Each candidate paragraph and the target risk seed are mapped to a semantic vector; Calculate the spatial similarity between the semantic vector of each candidate paragraph and the semantic vector of the target risk seed; Candidate paragraphs whose spatial similarity meets the preset sorting conditions are retained as the target paragraphs.
11. The method according to claim 9, characterized in that, The multiple steps and stages each correspond to different role types, including triggering factors, intermediate transmission, and risk outcomes. The process of extracting risk factors contained in the target paragraph and the causal transmission relationships between these risk factors, resulting in multiple sequentially pointing step stages as the risk causal chain, includes: Extract the risk factors contained in the target paragraph and the causal transmission relationship between each risk factor to generate a candidate link containing multiple step stages with sequential orientation; In the candidate links, the candidate links whose role type in the first step is the triggering factor, whose role type in the last step is the risk result, and which are consistent with the target risk seed, are designated as the risk causal links.
12. The method according to claim 11, characterized in that, The steps for obtaining the directed acyclic graph include: The original risk factors in the risk causal chain are semantically equivalently aggregated to form the standardized nodes, and the original risk factors that cannot be quantified are discarded. Count the number of times adjacent standardized nodes appear as related node pairs in different texts to remove duplicates; Using the number of deduplications as the support strength of the corresponding directed edges, a directed acyclic graph is constructed with the target risk seed as the endpoint.
13. A risk quantification assessment device, characterized in that, include: The acquisition module is used to acquire a directed acyclic graph, which is constructed based on standardized nodes obtained by semantic clustering of risk causal links extracted from risk disclosure texts; The generation module is used to generate the quantization execution logic corresponding to the non-terminal node based on the node information of the non-terminal node in the directed acyclic graph. The quantization execution logic is used to characterize the logic for data extraction and quantization judgment of the text. The parsing module is used to parse the text to be evaluated based on the quantization execution logic to obtain the node state of the non-termination node; The evaluation module is used to trace the risk transmission path along the directed acyclic graph based on the node status of each non-terminating node, and obtain the risk quantification evaluation result.
14. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the risk quantification assessment method as described in any one of claims 1 to 12.
15. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the risk quantification assessment method as described in any one of claims 1 to 12.