Decision risk assessment method and device, equipment and storage medium
By employing dual clustering and risk assessment methods, the problem of inaccurate risk assessment for multiple entities was solved, resulting in more precise risk assessment and decision support.
Patent Information
- Application Number
- CN202511685515.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-18
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2045-11-18
AI Technical Summary
Existing technologies do not provide sufficient detail in risk assessments involving multiple stakeholders, leading to inaccurate decision-making outcomes.
The dual clustering method first performs a first clustering based on basic qualification characteristics to determine the risk-driving characteristics and types of the first cluster centers. Then, a second clustering is performed based on the feature weights determined by the intra-cluster dispersion. Combining the risk calculation radius and difference assessment, entities that meet the investment promotion decision are selected.
It improves the accuracy of risk assessment, enabling a more precise determination of the risk level of each entity, and overcomes the shortcomings of traditional methods in terms of insufficient classification and analysis.
Smart Images

Figure CN121146531A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of risk assessment technology, and more specifically, relates to a decision risk assessment method, apparatus, equipment and storage medium. Background Technology
[0002] In various decision-making and evaluation scenarios, such as investment promotion, risk assessments are typically required to select suitable companies for invitation and recommendation from a large pool of information. Current technologies usually involve initial screening using assessment models followed by manual judgment. However, this approach lacks sufficient detail in classifying and analyzing multiple entities, leading to inaccurate risk assessments and impacting decision-making outcomes. Summary of the Invention
[0003] The purpose of this application is to provide a decision risk assessment method, apparatus, device and storage medium that can improve the accuracy of risk assessment for multiple subjects and make the decision results more accurate.
[0004] A first aspect of this application provides a decision risk assessment method, including: In response to receiving an investment promotion decision request, the system extracts features from the investment promotion decision request to obtain a text feature vector, and selects target keywords that match the text feature vector from a keyword library. There are multiple target keywords. The query is performed based on the target keywords to obtain initial query information, which contains multiple entities. The first clustering is performed on multiple entities to obtain the first clustering result, which includes N first cluster centers and N first clusters; N is a positive integer. Determine the risk-driving characteristics of each first cluster center, and determine the risk type of each first cluster center based on the risk-driving characteristics; The N first cluster centers are used as the initial centers for the second clustering. Based on the intra-cluster dispersion of each basic qualification feature in each first cluster in the first clustering, the feature weights of each basic qualification feature in the second clustering are determined. The second clustering is performed on each of the N first clusters to obtain M second cluster centers and M second clusters corresponding to each first cluster; M is a positive integer. The risk calculation radius for each second cluster is determined based on the risk type. For each second cluster, perform the following risk assessment: The average risk score of the second cluster and the risk score of each subject in the second cluster are calculated based on the risk calculation radius. The first difference is obtained by calculating the difference between the risk score of each subject and the average risk score of the second cluster. Risk assessment is performed on each entity based on the comparison between the absolute value of the first difference and the preset threshold.
[0005] A second aspect of this application provides a decision risk assessment device, comprising: The feature extraction unit is used to respond to the received investment promotion decision request information, extract features from the investment promotion decision request information to obtain a text feature vector, and select target keywords that match the text feature vector from the keyword library. There are multiple target keywords. The query unit is used to search based on target keywords and obtain initial query information, which contains multiple entities. The first clustering unit is used to perform the first clustering of multiple subjects to obtain the first clustering result, which includes N first cluster centers and N first clusters; N is a positive integer. The risk type determination unit is used to determine the risk-driving characteristics of each first cluster center and to determine the risk type of each first cluster center based on the risk-driving characteristics. The second clustering unit is used to take the N first cluster centers as the initial centers for the second clustering. Based on the intra-cluster dispersion of each basic qualification feature in each first cluster in the first clustering, the feature weights of each basic qualification feature in the second clustering are determined. The second clustering is performed on the N first clusters respectively to obtain M second cluster centers and M second clusters corresponding to each first cluster; M is a positive integer. A calculation unit is used to determine the risk calculation radius of each second cluster based on the risk type; For each second cluster, the risk assessment unit is used to perform the following risk assessment operations: The average risk score of the second cluster and the risk score of each subject in the second cluster are calculated based on the risk calculation radius. The first difference is obtained by calculating the difference between the risk score of each subject and the average risk score of the second cluster. Risk assessment is performed on each entity based on the comparison between the absolute value of the first difference and the preset threshold.
[0006] A third aspect of this application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and running on the processor, wherein the processor executes the computer program to implement the steps of the decision risk assessment method described above.
[0007] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the decision risk assessment method described above.
[0008] The beneficial effects of the decision risk assessment method, apparatus, device, and storage medium provided in this application are as follows: This application embodiment extracts text feature vectors based on investment promotion decision request information, selects target keywords matching the text feature vectors from a keyword library, and performs queries based on these target keywords to obtain initial query information. This allows for the rapid and accurate identification of multiple entities closely related to investment promotion decisions. This application embodiment achieves classification of multiple entities through a dual clustering approach. In the second clustering, the feature weights of the second clustering are determined based on the intra-cluster dispersion of each first cluster in the first clustering. This method fully considers the distribution of different features in different clusters, making the determination of feature weights more scientific and reasonable. It can more accurately reflect the role of each basic qualification feature in distinguishing the risks of different entities, thereby improving the accuracy of the clustering results.
[0009] In addition, this embodiment uses two clustering processes to filter out multiple entities that meet the investment promotion decision-making information. The risk calculation radius of each second cluster is determined based on the risk type. The entities obtained from the double clustering are then further filtered based on the risk calculation radius to obtain the final filtering results. Furthermore, the risk status of each entity can be assessed through its risk score. The method provided in this embodiment allows for a direct observation of the difference between each entity and the average level of its cluster, thereby more accurately determining the risk level of each entity. This overcomes the shortcomings of traditional methods in classifying and analyzing multiple entities without sufficient detail, and improves the accuracy of risk assessment. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1 A flowchart illustrating a decision risk assessment method provided in an embodiment of this application; Figure 2 A flowchart illustrating a decision risk assessment method provided in another embodiment of this application; Figure 3 A structural block diagram of a decision risk assessment device provided in an embodiment of this application; Figure 4 This is a schematic block diagram of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0014] Please refer to Figure 1 and Figure 2 , Figure 1 This is a flowchart illustrating a decision risk assessment method provided in an embodiment of this application. The method can be executed by a risk assessment platform and may include: S101: In response to receiving the investment promotion decision request information, extract features from the investment promotion decision request information to obtain a text feature vector, and select target keywords that match the text feature vector from the keyword library.
[0015] In this embodiment, there are multiple target keywords. Investment promotion decision request information refers to the key technical support needed by enterprises or other institutions for the research project, including various descriptions, requirements, and conditions related to the research project, specifically including but not limited to the industry type, investment scale, and qualification requirements for participating enterprises. After receiving the investment promotion decision request information, the risk assessment platform needs to extract features from the information to obtain a text feature vector.
[0016] This embodiment utilizes a bag-of-words model and a convolutional neural network model to extract features from investment promotion decision request information, resulting in a text feature vector. This vector contains multiple text features. For example, if the investment promotion decision request information is "Planning to carry out a smart city cooperation project, requiring the introduction of environmentally friendly and energy-saving manufacturing enterprises," then the text features would be "energy saving," "environmental protection," "manufacturing," and "smart city," and the text feature vector would be a set of text features. This embodiment considers that directly extracting features from the investment promotion decision request information is not conducive to subsequent data processing and could lead to inaccurate risk assessment results. Therefore, after obtaining the text feature vector, the text features can be expanded. The expansion method can be keyword matching. That is, target keywords matching the text feature vector are selected from a keyword library. Target keywords can be synonyms of the text features, or they can be hypernyms, hyponyms, or strongly related scene words.
[0017] S102: Perform a query based on the target keywords to obtain initial query information, which contains multiple entities.
[0018] In this embodiment, a query is performed on a preset data source based on the target keywords. The query logic can be exact matching, fuzzy matching, or a combination of both. The purpose of the query is to initially filter out entities that meet the investment promotion requirements from massive amounts of data.
[0019] S103: Perform the first clustering on multiple subjects to obtain the first clustering result, which includes N first cluster centers and N first clusters.
[0020] In one embodiment, a first clustering is performed on multiple subjects to obtain a first clustering result, including: The N value and initial cluster centers for the first clustering are determined based on the basic qualification characteristics of multiple subjects. The first clustering is then performed on multiple subjects based on the N value and initial cluster centers to obtain the first clustering result, where N is a positive integer.
[0021] The basic qualification characteristics of multiple entities represent the fundamental qualification features that each entity possesses. These basic qualification characteristics include, but are not limited to, industry attributes, company size, R&D capabilities, industry experience, and credit history. The first clustering result contains N first cluster centers and N first clusters, with each first cluster corresponding to one first cluster center. The first clustering essentially performs a coarse classification of the entities' basic qualification characteristics, yielding a "macro-level difference" among multiple entities. If risk assessment is directly performed on each entity based on the first clustering result, the results will be inaccurate. Therefore, a second clustering of multiple entities can be performed to achieve a more refined breakdown of risk.
[0022] S104: Determine the risk-driving characteristics of each first cluster center, and determine the risk type of each first cluster center based on the risk-driving characteristics.
[0023] In this embodiment, risk-driven features represent key characteristics affecting the risk level of the group to which the first cluster center belongs. For example, if the first cluster center is "environmental protection + credit," then its risk-driven feature could be "solid waste treatment compliance rate." Based on the above approach, the risk-driven features of each first cluster center are determined sequentially according to the common characteristics of each first cluster, and the risk type of each first cluster center is determined based on the risk-driven features. The risk type of the first cluster center can include three types: high risk, low risk, and medium risk. High risk indicates that all entities in its corresponding first cluster center pose a high risk to the investment promotion agency; low risk indicates that all entities in its corresponding first cluster center pose a relatively low risk to the investment promotion agency; and medium risk indicates that all entities in its corresponding first cluster center pose a moderate risk to the investment promotion agency.
[0024] S105: Using the N first cluster centers as the initial centers for the second clustering, based on the intra-cluster dispersion of each basic qualification feature in the first clustering, determine the feature weights of each basic qualification feature for the second clustering, and perform the second clustering on the N first clusters respectively to obtain M second cluster centers and M second clusters corresponding to each first cluster.
[0025] In this embodiment, the first cluster center is the center point of each first cluster after the first clustering. It can reflect the core aggregation direction of multiple subjects in the first cluster. Using N first cluster centers as the initial centers of the second clustering can avoid the randomness of selecting the initial centers from zero in the second clustering and reduce the instability of the clustering results. In this embodiment, M is a positive integer.
[0026] In this embodiment, intra-cluster dispersion is a quantitative indicator that measures the degree of dispersion of all entities within a first cluster on a certain basic qualification feature. Its core function is to reflect the consistency of entities within the cluster on that feature. The lower the intra-cluster dispersion, the more similar all entities within the cluster are on that feature; the higher the intra-cluster dispersion, the greater the differences among entities within the cluster on that feature. When, in the first clustering, the intra-cluster dispersion of a certain basic qualification feature is below the dispersion threshold, it indicates that the feature has high discriminative power for the cluster. Therefore, the feature weight of this feature can be increased in the second clustering, making the clustering results more focused on the core difference dimension.
[0027] In one embodiment, the feature weights of each basic qualification feature in the second cluster are determined based on the intra-cluster dispersion of each basic qualification feature in the first cluster in the first cluster, including: For each basic qualification feature: the average dispersion of the basic qualification feature within each first cluster is taken to obtain the average dispersion of the basic qualification feature in all first clusters. The intra-cluster dispersion is the standard deviation of the basic qualification feature within the cluster. Based on the average dispersion of the basic qualification feature in all first clusters and the dispersion of the basic qualification feature in all subjects, the feature discrimination is calculated. The feature discrimination of all basic qualification features is normalized to obtain the feature weight of each basic qualification feature in the second cluster.
[0028] In this embodiment, the feature discriminant can be obtained by comparing the average dispersion of the basic qualification features across all first clusters with the dispersion of the basic qualification features across all subjects. A higher feature discriminant indicates a worse discrimination effect of the first clustering on that feature. When the feature discriminant is high, the feature weights of each basic qualification feature in the second clustering can be increased, thereby making the second clustering more focused on features that were not sufficiently distinguished in the first clustering and improving the overall refinement of the clustering.
[0029] Based on the determined feature weights and initial centers, a second clustering is performed on multiple subjects. For each first cluster, the first cluster center of that first cluster is used as the initial center. The feature weights of each basic qualification feature in the second clustering are determined by comparing the intra-cluster dispersion of the basic qualification feature within that cluster with the dispersion threshold. For example, if the intra-cluster dispersion is greater than the dispersion threshold, the feature weight of that feature in the first clustering is increased by a factor of A to obtain the feature weight of that feature in the second clustering. After the second clustering is completed, the number of subjects in each first cluster is the same as the number of subjects in the M second clusters.
[0030] S106: Determine the risk calculation radius for each second cluster based on the risk type.
[0031] In this embodiment, risk types include high-risk, low-risk, and medium-risk types. The risk calculation radius for each second cluster is determined based on the risk type, including: Set a radius baseline value 'a', and adjust the radius baseline value 'a' based on the risk type to obtain the risk calculation radius of each second cluster.
[0032] Alternatively, a mapping table between risk types and risk calculation radii can be constructed, and the risk calculation radius of each second cluster can be determined based on the risk type.
[0033] Alternatively, determine the radius benchmark value b corresponding to each risk type; Calculate the intra-cluster dispersion of the basic qualification characteristics within each second cluster, and then weight the radius benchmark value b corresponding to the risk type of the first cluster center to which each second cluster belongs with the intra-cluster dispersion to obtain the risk calculation radius of each second cluster.
[0034] In this embodiment, the intra-cluster dispersion of the basic qualification characteristics within the second cluster refers to the degree of numerical dispersion of a certain basic qualification characteristic (such as environmental qualification level) among all subjects within a single second cluster, and is commonly calculated using standard deviation. The risk calculation radius of each second cluster is obtained by weighting the baseline value b corresponding to the risk type of the first cluster center to which each second cluster belongs with the intra-cluster dispersion. At this point, the risk calculation radius of each second cluster is greater than b, less than b, or equal to b.
[0035] In one embodiment, the decision risk assessment method further includes: If the risk calculation radius of the second cluster is greater than the preset maximum threshold, then the risk calculation radius of the second cluster is set to the preset maximum threshold. In the above method, when the risk calculation radius of the second cluster is greater than the preset maximum threshold, the risk calculation radius is promptly truncated to the preset maximum value to avoid the risk calculation radius becoming too large due to outliers.
[0036] S107: Perform a risk assessment operation for each second cluster.
[0037] refer to Figure 2 The risk assessment process includes: S1071: Calculate the average risk score of the second cluster and the risk score of each subject in the second cluster based on the risk calculation radius, and calculate the first difference by comparing the risk score of each subject with the average risk score of the second cluster. S1072: Risk assessment is performed on each subject based on the comparison result of the absolute value of the first difference and the preset threshold.
[0038] In this embodiment, after determining the risk calculation radius of each second cluster, it is necessary to filter multiple target entities with low risk scores that meet the investment promotion decision request information. The specific operation is as follows: This embodiment uses the risk calculation radius as the boundary to calculate the risk score and the first difference. For each second cluster, based on the assessment dimension corresponding to the risk type of the second cluster, the subjects within the cluster are quantitatively scored in multiple dimensions, and then the risk score of each subject is calculated by combining feature weights. At the same time, the effective subjects are determined based on the risk type and risk calculation radius corresponding to the second cluster. For example, in a high-risk second cluster, the risk score of subjects within the risk calculation radius is greater than the risk score of subjects outside the risk radius; in a low-risk second cluster, the risk score of subjects within the risk calculation radius is less than the risk score of subjects outside the risk radius. Therefore, all subjects outside the risk calculation radius can be considered as effective subjects. The average risk score of the effective subjects is calculated to obtain the average risk score of the second cluster. The average risk score is subtracted from the risk score of each effective subject to obtain the first difference reflecting the difference between individual and group risk.
[0039] In this embodiment, after the above subject screening, the average risk score of the effective subjects is lower than that of other subjects. However, in order to eliminate the calculation error of individual effective subjects, this embodiment uses the comparison result of the absolute value of the first difference with the preset threshold as the judgment basis. If the absolute value of the first difference is greater than the preset threshold, it means that the subject's risk level is significantly different from the average risk level of the cluster, and it is judged as an abnormal subject. Thus, the abnormal value is removed, and the screening and risk assessment of effective subjects are finally completed.
[0040] As can be seen from the above, this embodiment extracts text feature vectors based on investment promotion decision request information, selects target keywords matching the text feature vectors from the keyword library, and performs queries based on the target keywords to obtain initial query information. This allows for the rapid and accurate identification of multiple entities closely related to investment promotion decisions. This embodiment also achieves the classification of multiple entities through dual clustering. In the second clustering, the feature weights of the second clustering are determined based on the intra-cluster dispersion of each first cluster in the first clustering. This approach fully considers the distribution of different features in different clusters, making the determination of feature weights more scientific and reasonable. It can more accurately reflect the role of each basic qualification feature in distinguishing the risks of different entities, thereby improving the accuracy of the clustering results.
[0041] In addition, this embodiment uses two clustering processes to filter out multiple entities that meet the investment promotion decision-making information. The risk calculation radius of each second cluster is determined based on the risk type. The entities obtained from the double clustering are then further filtered based on the risk calculation radius to obtain the final filtering results. Furthermore, the risk status of each entity can be assessed through its risk score. The method provided in this embodiment allows for a direct observation of the difference between each entity and the average level of its cluster, thereby more accurately determining the risk level of each entity. This overcomes the shortcomings of traditional methods in classifying and analyzing multiple entities without sufficient detail, and improves the accuracy of risk assessment.
[0042] In one embodiment of this application, determining the risk-driven characteristics of each first cluster center includes: For each first cluster center: Extract multiple risk association features of the main entities within the first cluster corresponding to the first cluster center; Calculate the correlation coefficient between each risk-related feature and the risk events of the main body in the first cluster. The risk-related features whose absolute value of the correlation coefficient is greater than the preset coefficient threshold are determined as the risk-driving features of the first cluster center.
[0043] In this embodiment, risk-related features indicate that the feature may affect the risk assessment of the subject, but may not be a key factor causing the risk, requiring further analysis. Risk-driving features, on the other hand, indicate that the feature is strongly correlated with the risk and can directly affect the probability of the risk occurring. For each first cluster center, multiple risk-related features of the subjects within the first cluster corresponding to the first cluster center can be extracted. These risk-related features include all the subject's basic qualification features and corporate relationship features. Corporate relationship features include cooperative relationships, social relationships, and industry relationships. Cooperative relationships include the qualifications of upstream suppliers, the types of downstream customers, and the credit rating of partners. Social relationships include the social identity of the corporate entity, the number of affiliated companies, and interaction records with regulatory agencies. Industry relationships include the position in the industry chain and membership in industry associations.
[0044] The correlation coefficients between each risk-related feature and the risk events of the main body within the first cluster are calculated separately. Risk-related features whose absolute values of correlation coefficients are greater than a preset threshold are identified as risk-driving features of the first cluster center. This can be understood as feature screening, that is, by using the correlation relationship between risk-related features and the risk events of the main body, features with high correlation and easy to cause risk are screened out from multiple risk-related features to obtain risk-driving features.
[0045] This embodiment obtains risk-driven features through feature filtering, which can reduce the amount of data computation and the time of the clustering process.
[0046] In one embodiment of this application, determining the risk type of each first cluster center based on risk-driven characteristics includes: For each first cluster center: The risk-driving features of the first cluster center are quantified to obtain the quantified values of each risk-driving feature. Each risk-driving feature and its relative quantified value are combined to form a feature combination. Calculate the matching degree between each feature combination and each preset feature combination in the preset risk type system. The preset risk type system includes multiple preset feature combinations and the preset risk types corresponding to the preset feature combinations. If the matching degree of multiple feature combinations is greater than the matching degree threshold, the relative preset risk types of the multiple feature combinations are statistically analyzed to obtain the statistical results, and the risk type of the first cluster center is determined based on the statistical results. If there are multiple feature combinations whose matching degree is less than or equal to the matching degree threshold, then the risk type of the first cluster center is determined to be low risk.
[0047] In this embodiment, for numerical risk-driven features, quantification can be performed directly based on numerical values; for non-numerical risk-driven features, the risk-driven features can be encoded, and then the encoded data can be quantized. The matching degree between each feature combination and each preset feature combination in the preset risk type system is calculated. If multiple feature combinations have matching degrees greater than a matching degree threshold (e.g., feature combination 1, feature combination 2, and feature combination 3 have matching degrees greater than the matching degree threshold, and feature combination 1 corresponds to a high-risk type, feature combination 2 corresponds to a high-risk type, and feature combination 3 corresponds to a medium-risk type), then the risk type of the first cluster center is high-risk. If multiple feature combinations have matching degrees less than or equal to the matching degree threshold, it indicates that none of them match, and the risk type of the first cluster center can be determined as low-risk.
[0048] This embodiment, based on quantitative features and explicit matching rules, makes the source of risk types traceable, thereby improving the credibility of decision-making.
[0049] In one embodiment of this application, the risk types include high-risk, medium-risk, and low-risk types; The average risk score of the second cluster is calculated based on the risk calculation radius, including: If the risk type of the second cluster is high risk, a circular region is determined with the second cluster center corresponding to the second cluster as the center and the risk calculation radius as the radius. The average risk score of the second cluster is calculated based on the subjects outside the circular region. If the risk type of the second cluster is medium risk, a circular region is determined with the second cluster center corresponding to the second cluster as the center and the risk calculation radius as the radius. The average risk score of the second cluster is calculated based on the subjects within the circular region or based on the subjects outside the circular region. If the risk type of the second cluster is low risk, a circular region is determined with the second cluster center corresponding to the second cluster as the center and the risk calculation radius as the radius. The average risk score of the second cluster is calculated based on the subjects within the circular region.
[0050] In one embodiment, calculating the average risk score of the second cluster based on the main bodies within the circular region includes: At least one risk assessment dimension is determined based on the risk-driven characteristics of the first cluster center to which the second cluster belongs; For each risk assessment dimension, the feature values of all subjects within the circular area are converted into standardized risk scores, where the feature values are standardized values used to describe any attribute of the subject. Based on the feature weights of each risk assessment dimension in the second clustering, the weighted total risk score of each subject in each region is calculated; The average of the weighted total risk scores for all subjects is calculated to obtain the average risk score for the second cluster.
[0051] In this embodiment, the feature weights of each risk assessment dimension are the feature weights corresponding to the risk-driving features. A specific example is as follows: Assume that four risk assessment dimensions are determined based on the risk-driving features of the first cluster center to which the second cluster belongs: technology R&D dimension, intellectual property reserve dimension, team personnel quantity dimension, and funding dimension. For the technology R&D dimension, the feature values of all entities within the circular area can include the proportion of enterprise R&D investment, the proportion of R&D personnel, etc. Converting the feature values of all entities within the circular area into standardized risk scores can be understood as converting the R&D investment proportion into a standardized risk score. For example, the R&D investment proportion ranges from (0,1). The closer the R&D investment proportion is to 1, the higher the standardized risk score; the closer the R&D investment proportion is to 0, the lower the standardized risk score. Based on the second clustering results, the feature weights of each risk assessment dimension are determined as follows: technology R&D dimension 0.35, intellectual property reserve dimension 0.25, team personnel quantity dimension 0.2, and funding dimension 0.2. The weighted total risk score for each enterprise is calculated using the following formula: Weighted Total Risk Score = Technology R&D Dimension Score × 0.35 + Intellectual Property Reserve Dimension Score × 0.25 + Team Member Quantity Dimension Score × 0.2 + Capital Dimension Score × 0.2. The weighted total risk scores of all enterprises within the circular area are statistically analyzed, and their average is calculated to obtain the average risk score for the second cluster, providing a crucial reference for investment attraction decisions in the science and technology industrial park.
[0052] Corresponding to the decision risk assessment method in the above embodiments, Figure 3 This is a structural block diagram of a decision risk assessment device provided in one embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 3 The decision risk assessment device 20 includes: a feature extraction unit 21, a query unit 22, a first clustering unit 23, a risk type determination unit 24, a second clustering unit 25, a calculation unit 26, and a risk assessment unit 27.
[0053] Among them, the feature extraction unit 21 is used to extract features from the investment promotion decision request information in response to receiving the investment promotion decision request information to obtain a text feature vector, and select target keywords that match the text feature vector from the keyword library. There are multiple target keywords. Query unit 22 is used to query based on target keywords to obtain initial query information, which contains multiple subjects; The first clustering unit 23 is used to perform the first clustering of multiple subjects to obtain the first clustering result, which includes N first cluster centers and N first clusters; N is a positive integer; Risk type determination unit 24 is used to determine the risk driving characteristics of each first cluster center and determine the risk type of each first cluster center based on the risk driving characteristics; The second clustering unit 25 is used to take N first cluster centers as the initial centers of the second clustering, determine the feature weights of each basic qualification feature in the second clustering based on the intra-cluster dispersion of each basic qualification feature in each first cluster in the first clustering, and perform the second clustering on each of the N first clusters to obtain M second cluster centers and M second clusters corresponding to each first cluster; M is a positive integer. Calculation unit 26 is used to determine the risk calculation radius of each second cluster based on the risk type; For each second cluster, risk assessment unit 27 is used to perform the following risk assessment operations: The average risk score of the second cluster and the risk score of each subject in the second cluster are calculated based on the risk calculation radius. The first difference is obtained by calculating the difference between the risk score of each subject and the average risk score of the second cluster. Risk assessment is performed on each entity based on the comparison between the absolute value of the first difference and the preset threshold.
[0054] In one embodiment of this application, the risk type determination unit 24 is specifically used for: For each first cluster center: Extract multiple risk association features of the main entities within the first cluster corresponding to the first cluster center; Calculate the correlation coefficient between each risk-related feature and the risk events of the main body in the first cluster. The risk-related features whose absolute value of the correlation coefficient is greater than the preset coefficient threshold are determined as the risk-driving features of the first cluster center.
[0055] In one embodiment of this application, the risk type determination unit 24 is specifically used for: For each first cluster center: The risk-driving features of the first cluster center are quantified to obtain the quantified values of each risk-driving feature. Each risk-driving feature and its relative quantified value are combined to form a feature combination. Calculate the matching degree between each feature combination and each preset feature combination in the preset risk type system. The preset risk type system includes multiple preset feature combinations and the preset risk types corresponding to the preset feature combinations. If the matching degree of multiple feature combinations is greater than the matching degree threshold, the relative preset risk types of the multiple feature combinations are statistically analyzed to obtain the statistical results, and the risk type of the first cluster center is determined based on the statistical results. If there are multiple feature combinations whose matching degree is less than or equal to the matching degree threshold, then the risk type of the first cluster center is determined to be low risk.
[0056] In one embodiment of this application, the second clustering unit 25 is specifically used for: For each basic qualification feature: The average dispersion of the basic qualification features within each first cluster is obtained by averaging the dispersion of the basic qualification features within each first cluster. The dispersion within a cluster is the standard deviation of the basic qualification features within the cluster. The feature discrimination is calculated based on the average dispersion of basic qualification features in all first clusters and the dispersion of basic qualification features in all subjects; The feature discrimination of all basic qualification features is normalized to obtain the feature weights of each basic qualification feature in the second cluster.
[0057] In one embodiment of this application, the computing unit 26 is specifically used for: Determine the corresponding radius benchmark value for each risk type; Calculate the intra-cluster dispersion of the basic qualification characteristics within each second cluster, and then weight the baseline value of the risk type corresponding to the first cluster center to which each second cluster belongs with the intra-cluster dispersion to obtain the risk calculation radius of each second cluster.
[0058] In one embodiment of this application, the risk types include high-risk, medium-risk, and low-risk types; the calculation unit 26 is specifically used for: If the risk type of the second cluster is high risk, a circular region is determined with the second cluster center corresponding to the second cluster as the center and the risk calculation radius as the radius. The average risk score of the second cluster is calculated based on the subjects outside the circular region. If the risk type of the second cluster is medium risk, a circular region is determined with the second cluster center corresponding to the second cluster as the center and the risk calculation radius as the radius. The average risk score of the second cluster is calculated based on the subjects within the circular region or based on the subjects outside the circular region. If the risk type of the second cluster is low risk, a circular region is determined with the second cluster center corresponding to the second cluster as the center and the risk calculation radius as the radius. The average risk score of the second cluster is calculated based on the subjects within the circular region.
[0059] In one embodiment of this application, the computing unit 26 is specifically used for: At least one risk assessment dimension is determined based on the risk-driven characteristics of the first cluster center to which the second cluster belongs; For each risk assessment dimension, the feature values of all subjects within the circular area are converted into standardized risk scores, where the feature values are standardized values used to describe any attribute of the subject. Based on the feature weights of each risk assessment dimension in the second clustering, the weighted total risk score of each subject in each region is calculated; The average of the weighted total risk scores for all subjects is calculated to obtain the average risk score for the second cluster.
[0060] See Figure 4 , Figure 4 This is a schematic block diagram of an electronic device provided according to an embodiment of this application. Figure 4 The electronic device 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the units in the aforementioned device embodiments, for example... Figure 3 The functions of the feature extraction unit 21, query unit 22, first clustering unit 23, risk type determination unit 24, second clustering unit 25, calculation unit 26, and risk assessment unit 27 are shown.
[0061] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0062] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.
[0063] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory.
[0064] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in the decision risk assessment method provided in the embodiments of this application, or they can execute the implementation methods of the electronic devices described in the embodiments of this application, which will not be repeated here.
[0065] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0066] The computer-readable storage medium can be an internal storage unit of the electronic device in any of the foregoing embodiments, such as a hard disk or memory of the electronic device. The computer-readable storage medium can also be an external storage device of the electronic device, such as a plug-in hard disk, smart media card (SMC), secure digital card (SD), flash card, etc., equipped on the electronic device. Furthermore, the computer-readable storage medium can include both internal and external storage units of the electronic device. The computer-readable storage medium is used to store computer programs and other programs and data required by the electronic device. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0067] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0068] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working process of the electronic devices and units described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0069] In the several embodiments provided in this application, it should be understood that the disclosed electronic devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.
[0070] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0071] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0072] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A decision-making risk assessment method, characterized in that, include: In response to receiving an investment promotion decision request, the system performs feature extraction on the investment promotion decision request to obtain a text feature vector, and selects target keywords that match the text feature vector from a keyword library. There are multiple target keywords. A query is performed based on the target keywords to obtain initial query information, which contains multiple entities. The multiple entities are first clustered to obtain a first clustering result, which includes N first cluster centers and N first clusters; N is a positive integer. Determine the risk-driving characteristics of each first cluster center, and determine the risk type of each first cluster center based on the risk-driving characteristics; The N first cluster centers are used as the initial centers for the second clustering. Based on the intra-cluster dispersion of each basic qualification feature in each first cluster in the first clustering, the feature weights of each basic qualification feature in the second clustering are determined. The N first clusters are then subjected to the second clustering to obtain M second cluster centers and M second clusters corresponding to each first cluster; M is a positive integer. The risk calculation radius for each second cluster is determined based on the aforementioned risk type; For each second cluster, perform the following risk assessment: The average risk score of the second cluster and the risk score of each subject in the second cluster are calculated based on the risk calculation radius. The first difference is obtained by calculating the difference between the risk scores of each subject and the average risk score of the second cluster. Risk assessment is performed on each entity based on the comparison between the absolute value of the first difference and the preset threshold.
2. The decision risk assessment method as described in claim 1, characterized in that, The determination of the risk-driven characteristics of each first cluster center includes: For each first cluster center: Extract multiple risk association features of the subjects within the first cluster corresponding to the first cluster center; The correlation coefficients between each risk-related feature and the risk events of the main body within the first cluster are calculated respectively. Risk-related features whose absolute values of correlation coefficients are greater than preset coefficient thresholds are determined as risk-driving features of the first cluster center.
3. The decision risk assessment method as described in claim 1, characterized in that, The determination of the risk type of each first cluster center based on the risk-driven characteristics includes: For each first cluster center: The risk-driving features of the first cluster center are quantified to obtain the quantified values of each risk-driving feature. Each risk-driving feature and its relative quantified value are combined to form a feature combination. Calculate the matching degree between each feature combination and each preset feature combination in the preset risk type system, wherein the preset risk type system includes multiple preset feature combinations and the preset risk types corresponding to the preset feature combinations; If the matching degree of multiple feature combinations is greater than the matching degree threshold, then the preset risk types of the multiple feature combinations are statistically analyzed to obtain statistical results, and the risk type of the first cluster center is determined based on the statistical results. If there are multiple feature combinations whose matching degree is less than or equal to the matching degree threshold, then the risk type of the first cluster center is determined to be low risk.
4. The decision risk assessment method as described in claim 1, characterized in that, The determination of feature weights for each basic qualification feature in the second clustering, based on the intra-cluster dispersion of each basic qualification feature in the first clustering, includes: For each basic qualification feature: The average dispersion of the basic qualification feature within each first cluster is obtained by averaging the dispersion of the basic qualification feature within each first cluster. The dispersion within each cluster is the standard deviation of the basic qualification feature within the cluster. The feature discrimination is calculated based on the average dispersion of the basic qualification features in all first clusters and the dispersion of the basic qualification features in all subjects; The feature discrimination of all basic qualification features is normalized to obtain the feature weights of each basic qualification feature in the second cluster.
5. The decision risk assessment method as described in claim 1, characterized in that, The determination of the risk calculation radius for each second cluster based on the risk type includes: Determine the radius benchmark value corresponding to each of the aforementioned risk types; Calculate the intra-cluster dispersion of the basic qualification characteristics within each second cluster, and then weight the radius benchmark value corresponding to the risk type of the first cluster center to which each second cluster belongs with the intra-cluster dispersion to obtain the risk calculation radius of each second cluster.
6. The decision risk assessment method as described in claim 5, characterized in that, The risk types include high-risk, medium-risk, and low-risk types; The calculation of the average risk score of the second cluster based on the risk calculation radius includes: If the risk type of the second cluster is high risk, a circular region is determined with the second cluster center corresponding to the second cluster as the center and the risk calculation radius as the radius. The average risk score of the second cluster is calculated based on the main body outside the circular region. If the risk type of the second cluster is medium risk, a circular region is determined with the second cluster center corresponding to the second cluster as the center and the risk calculation radius as the radius. The average risk score of the second cluster is calculated based on the subjects within the circular region or based on the subjects outside the circular region. If the risk type of the second cluster is low risk, a circular region is determined with the second cluster center corresponding to the second cluster as the center and the risk calculation radius as the radius. The average risk score of the second cluster is calculated based on the subjects within the circular region.
7. The decision risk assessment method as described in claim 6, characterized in that, The calculation of the average risk score of the second cluster based on the main body within the circular region includes: At least one risk assessment dimension is determined based on the risk-driven characteristics of the first cluster center to which the second cluster belongs; For each risk assessment dimension, the feature values of all subjects within the circular area are converted into standardized risk scores, where the feature values are standardized values used to describe any attribute of the subject. Based on the feature weights of each risk assessment dimension in the second clustering, the weighted total risk score of each subject in each region is calculated; The average of the weighted total risk scores for all subjects is calculated to obtain the average risk score for the second cluster.
8. A decision risk assessment device, characterized in that, include: The feature extraction unit is used to, in response to receiving investment promotion decision request information, extract features from the investment promotion decision request information to obtain a text feature vector, and select target keywords that match the text feature vector from a keyword library, wherein there are multiple target keywords; The query unit is used to query based on the target keywords to obtain initial query information, which includes multiple entities; The first clustering unit is used to perform the first clustering of the multiple entities to obtain the first clustering result, which includes N first cluster centers and N first clusters; A risk type determination unit is used to determine the risk-driving characteristics of each first cluster center, and to determine the risk type of each first cluster center based on the risk-driving characteristics; The second clustering unit is used to take the N first cluster centers as the initial centers for the second clustering, determine the feature weights of each basic qualification feature in the second clustering based on the intra-cluster dispersion of each basic qualification feature in each first cluster in the first clustering, and perform the second clustering on each of the N first clusters to obtain M second cluster centers and M second clusters corresponding to each first cluster; M is a positive integer. A calculation unit is used to determine the risk calculation radius of each second cluster based on the risk type; For each second cluster, the risk assessment unit is used to perform the following risk assessment operations: The average risk score of the second cluster and the risk score of each subject in the second cluster are calculated based on the risk calculation radius. The first difference is obtained by calculating the difference between the risk scores of each subject and the average risk score of the second cluster. Risk assessment is performed on each entity based on the comparison between the absolute value of the first difference and the preset threshold.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Enterprise risk assessment method and device, computer equipment and storage medium
CN115759742A
Object acquisition method and device, electronic equipment and storage medium
CN116541474A
Risk area generation method and device, computer equipment and storage medium
CN119250987A
Risk assessment method and device, computer readable storage medium and electronic equipment
CN119624623A
Company risk analyzing method, apparatus, computer device, and storage medium
WO2020107872A1