Risk assessment methods, devices, equipment, and media based on big data analytics
Patent Information
- Application Number
- CN202611077219.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-20
- Publication Date
- 2026-09-11
AI Technical Summary
[0003]然而,现有的风险评估方法一般依赖人工排查,这会降低数据评估的效率,同时人工排查很难捕捉多源数据内部包含的数据关联关系,从而降低了数据风险评估的准确性
[0035]The aforementioned risk assessment method, apparatus, equipment, and medium based on big data analytics acquire business activity data, business resource data, business approval-related data, and external supervision data of the target business from various business-related platforms. Feature extraction is then performed on this multi-source business-related data to obtain multi-dimensional business characteristics of the target business. Subsequently, a graph attention network is used to process these multi-dimensional business characteristics to obtain the current risk assessment parameters of the target business. Based on these current risk assessment parameters and preset risk benchmark parameters, the risk assessment result of the target business is determined. Compared to related technologies that rely on manual screening to analyze business risks contained in multi-source business-related data, this method, combined with a graph attention network, performs a full-link analysis of the multi-dimensional business characteristics of multi-source business-related data. This effectively captures data changes along the link, obtains more accurate current risk assessment parameters, and thus improves the accuracy of subsequent data risk assessments.
Smart Images

Figure CN122736343A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a risk assessment method, apparatus, device, and medium based on big data analysis. Background Technology
[0002] With the deepening of digital transformation, in order to improve the prevention and control of data risks, methods for big data risk assessment have emerged. For example, the risk assessment of current data can be carried out by combining business process logic and historical data flow.
[0003] However, existing risk assessment methods generally rely on manual screening, which reduces the efficiency of data assessment. At the same time, manual screening makes it difficult to capture the data relationships contained within multi-source data, thereby reducing the accuracy of data risk assessment. Summary of the Invention
[0004] Therefore, it is necessary to provide a risk assessment method, apparatus, equipment, and medium based on big data analysis to address the aforementioned technical problems and improve the accuracy of data risk assessment.
[0005] Firstly, this application provides a risk assessment method based on big data analysis, including:
[0006] In response to risk assessment requests for the target business, multi-source business-related data for the target business is obtained from various business-related platforms; among which, multi-source business-related data includes business activity data, business resource data, business approval-related data, and external supervision data;
[0007] Feature extraction is performed on multi-source business-related data to obtain multi-dimensional business features of the target business.
[0008] By using graph attention networks, multi-dimensional business features are processed to obtain the current risk assessment parameters for the target business.
[0009] Based on the current risk assessment parameters and the preset risk benchmark parameters, determine the risk assessment results for the target business.
[0010] In one embodiment, the multidimensional business features include abnormal business frequency features, business object association features, business process deviation features, resource flow matching features, abnormal resource fluctuation features, resource approval link features, business node risk features, and supervision feedback association features. Feature extraction is performed on multi-source business-related data to obtain the multidimensional business features of the target business, including: feature extraction from business activity data to obtain abnormal business frequency features, business object association features, and business process deviation features; determining resource flow matching features, abnormal resource fluctuation features, and resource approval link features based on business resource data and the target business's business information; and determining business node risk features and supervision feedback association features based on business approval-related data and external supervision data.
[0011] In one embodiment, a graph attention network is used to process multi-dimensional business features to obtain current risk assessment parameters for the target business. This includes: determining a feature correlation matrix based on the correlation between multi-dimensional business features using the graph attention network; processing the feature correlation matrix based on historical business risk data to obtain a target correlation matrix; and determining the current risk assessment parameters for the target business based on the target correlation matrix.
[0012] In one embodiment, the feature association matrix is processed based on historical business risk data to obtain a target association matrix, including: extracting each feature association link in the feature association matrix and the feature association type of each feature association link; determining the historical association frequency and historical association strength of each feature association link based on historical business risk data; determining the feature association confidence level of each feature association link based on the feature association type, historical association frequency, and historical association strength; and fusing the feature association matrix and the feature association confidence levels of each feature association link to obtain the target association matrix.
[0013] In one embodiment, determining the current risk assessment parameters of the target business based on the target association matrix includes: processing the target association matrix using a gradient boosting tree model to obtain risk association features; determining the current risk assessment parameters of the target business based on the matrix parameters of the links where the risk association features are located in the target association matrix; wherein the matrix parameters include feature association confidence and information gain value; the information gain value is determined based on the feature association confidence of the link and the risk probability of the link in historical business risk data.
[0014] In one embodiment, determining the risk assessment result of the target business based on current risk assessment parameters and preset risk benchmark parameters includes: determining the current risk level and risk events in the target business based on the current risk assessment parameters and preset risk benchmark parameters; wherein, the preset risk benchmark parameters are determined based on the business domain of the target business and the object hierarchy of the related processing objects of the target business; constructing a risk transmission map based on the business characteristics associated with the risk events in the multi-dimensional business characteristics; determining risk propagation information based on preset risk diffusion parameters and the risk transmission map; and determining the risk assessment result of the target business based on the current risk level, the risk transmission map, and the risk propagation information.
[0015] Secondly, this application also provides a risk assessment device based on big data analysis, comprising:
[0016] The data acquisition module is used to respond to risk assessment requests for the target business by acquiring multi-source business-related data of the target business from various business-related platforms. Among them, the multi-source business-related data includes business activity data, business resource data, business approval-related data, and external supervision data.
[0017] The feature extraction module is used to extract features from multi-source business-related data to obtain multi-dimensional business features of the target business.
[0018] The parameter acquisition module is used to process multi-dimensional business features through a graph attention network to obtain the current risk assessment parameters of the target business.
[0019] The risk assessment module is used to determine the risk assessment results of the target business based on the current risk assessment parameters and preset risk benchmark parameters.
[0020] Thirdly, this application also provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to perform the following steps:
[0021] In response to risk assessment requests for the target business, multi-source business-related data for the target business is obtained from various business-related platforms; among which, multi-source business-related data includes business activity data, business resource data, business approval-related data, and external supervision data;
[0022] Feature extraction is performed on multi-source business-related data to obtain multi-dimensional business features of the target business.
[0023] By using graph attention networks, multi-dimensional business features are processed to obtain the current risk assessment parameters for the target business.
[0024] Based on the current risk assessment parameters and the preset risk benchmark parameters, determine the risk assessment results for the target business.
[0025] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the following steps:
[0026] In response to risk assessment requests for the target business, multi-source business-related data for the target business is obtained from various business-related platforms; among which, multi-source business-related data includes business activity data, business resource data, business approval-related data, and external supervision data;
[0027] Feature extraction is performed on multi-source business-related data to obtain multi-dimensional business features of the target business.
[0028] By using graph attention networks, multi-dimensional business features are processed to obtain the current risk assessment parameters for the target business.
[0029] Based on the current risk assessment parameters and the preset risk benchmark parameters, determine the risk assessment results for the target business.
[0030] Fifthly, this application also provides a computer program product, including a computer program that, when executed by a processor, performs the following steps:
[0031] In response to risk assessment requests for the target business, multi-source business-related data for the target business is obtained from various business-related platforms; among which, multi-source business-related data includes business activity data, business resource data, business approval-related data, and external supervision data;
[0032] Feature extraction is performed on multi-source business-related data to obtain multi-dimensional business features of the target business.
[0033] By using graph attention networks, multi-dimensional business features are processed to obtain the current risk assessment parameters for the target business.
[0034] Based on the current risk assessment parameters and the preset risk benchmark parameters, determine the risk assessment results for the target business.
[0035] The aforementioned risk assessment method, apparatus, equipment, and medium based on big data analytics acquire business activity data, business resource data, business approval-related data, and external supervision data of the target business from various business-related platforms. Feature extraction is then performed on this multi-source business-related data to obtain multi-dimensional business characteristics of the target business. Subsequently, a graph attention network is used to process these multi-dimensional business characteristics to obtain the current risk assessment parameters of the target business. Based on these current risk assessment parameters and preset risk benchmark parameters, the risk assessment result of the target business is determined. Compared to related technologies that rely on manual screening to analyze business risks contained in multi-source business-related data, this method, combined with a graph attention network, performs a full-link analysis of the multi-dimensional business characteristics of multi-source business-related data. This effectively captures data changes along the link, obtains more accurate current risk assessment parameters, and thus improves the accuracy of subsequent data risk assessments. Attached Figure Description
[0036] To more clearly illustrate the technical solutions in the embodiments of this application or related technologies, the drawings used in the description of the embodiments of this application or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.
[0037] Figure 1 This is a flowchart illustrating a risk assessment method based on big data analysis in one embodiment;
[0038] Figure 2 This is a flowchart illustrating the process of determining multi-dimensional business characteristics in one embodiment;
[0039] Figure 3 This is a flowchart illustrating the process of determining the current risk assessment parameters in one embodiment;
[0040] Figure 4 This is a flowchart illustrating the process of determining the target association matrix in one embodiment;
[0041] Figure 5 This is a flowchart illustrating the process of determining the current risk assessment parameters in one embodiment;
[0042] Figure 6 This is a flowchart illustrating the process of determining risk assessment results in one embodiment;
[0043] Figure 7 This is a flowchart illustrating a risk assessment method based on big data analysis in another embodiment;
[0044] Figure 8 This is a structural block diagram of a risk assessment device based on big data analysis in one embodiment;
[0045] Figure 9 This is an internal structural diagram of a computer device in one embodiment. Detailed Implementation
[0046] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0047] With the deepening of digital transformation, in order to improve the prevention and control of data risks, methods for big data risk assessment have emerged. For example, the risk assessment of current data can be carried out by combining business process logic and historical data flow.
[0048] However, existing risk assessment methods generally rely on manual screening, which reduces the efficiency of data assessment. At the same time, manual screening makes it difficult to capture the data relationships contained within multi-source data, thereby reducing the accuracy of data risk assessment.
[0049] Based on this, in an exemplary embodiment, a risk assessment method based on big data analysis is provided. The method is illustrated using an application to a server as an example. Figure 1 As shown, the specific steps include:
[0050] S101, in response to a risk assessment request for the target business, obtains multi-source business-related data of the target business from various business-related platforms.
[0051] The target business can be understood as the business that requires risk assessment. The risk assessment request instructs the server to perform a risk assessment on the target business. The business-related platform can be understood as a data platform related to the target business. Multi-source business-related data can be understood as data related to the target business across multiple dimensions. Multi-source business-related data includes business activity data, business resource data, business approval-related data, and external supervision data.
[0052] Business activity data can be understood as data generated by activities related to the target business. Business resource data can be understood as data related to resources within the target business; resources here can refer to information such as resources themselves. Business approval-related data can be understood as data generated during the approval process of each execution step of the target business. External supervision data can be understood as relevant data generated when external organizations monitor the target.
[0053] Optionally, upon detecting a risk assessment request for a target business, multi-source business association data of the target business can be obtained from various business-related platforms based on the identification information of the target business.
[0054] For example, business activity data can be collected by connecting to the business management system through interface call protocols, including core fields such as business travel, reception, and procurement; business resource data can be collected by connecting to the resource payment system and the bank's docking platform through standardized data interaction interfaces, covering information such as resource allocation, reimbursement, and account transactions; business approval-related data can be collected by relying on the data sharing channel of the operation supervision platform, covering the entire process of approval and project management; and external supervision data can be collected by connecting to the supervision platform, monitoring system, and third-party supervision feedback channels through API interfaces.
[0055] S102, extract features from multi-source business-related data to obtain multi-dimensional business features of the target business.
[0056] Among them, multidimensional business characteristics can be understood as the business characteristics of multi-source business-related data under multiple feature dimensions. For example, multidimensional business characteristics include business frequency anomaly characteristics, business object association characteristics, business process deviation characteristics, resource flow matching characteristics, abnormal resource fluctuation characteristics, resource approval chain characteristics, business node risk characteristics, and supervision feedback association characteristics.
[0057] Optionally, feature extraction can be performed on multi-source business-related data under preset dimensions to obtain multi-dimensional business features of the target business.
[0058] For example, time-series analysis can be used to extract abnormal business frequency features from multi-source business association data; association rule mining algorithms can be used to extract business object association features from multi-source business association data; process specification comparison methods can be used to extract business process deviation features from multi-source business association data; resource flow matching features can be extracted from multi-source business association data using business activity number, participants, and timestamp as association keys; time-series trend analysis can be used to mine abnormal fluctuations in resource quotas to obtain abnormal resource fluctuation features; resource approval link features can be extracted by decomposing the process link of the target business; and node extraction algorithms can be used to locate key nodes in the target business and match corresponding supervision information to extract business node risk features and supervision feedback association features.
[0059] Furthermore, the various business features can be normalized, and redundant features with a correlation coefficient ≥ 0.8 can be removed by Pearson correlation coefficient analysis. Then, effective features can be integrated with feature fusion algorithms to construct a multidimensional business feature that includes business features under each dimension.
[0060] S103 uses a graph attention network to process multi-dimensional business features and obtain the current risk assessment parameters for the target business.
[0061] Graph attention networks are used to uncover the relationships between features. Current risk assessment parameters are used to characterize the risk profile of multi-source business-related data.
[0062] Optionally, a graph attention network can be used to process multi-dimensional business features to obtain the correlation and transmission relationships between multi-dimensional business features. Then, by combining the correlation and transmission relationships with the business processing logic of the target business standard, the current risk assessment parameters of the target business can be determined.
[0063] S104. Determine the risk assessment result of the target business based on the current risk assessment parameters and the preset risk benchmark parameters.
[0064] The preset risk benchmark parameters can be understood as pre-set values that measure the magnitude of the current risk assessment parameters. For example, a domain-hierarchical classification modeling method can be used to set preset risk benchmark parameters for different businesses. The risk assessment result can be understood as the result obtained after conducting a risk assessment on the target business.
[0065] Optionally, one can first determine preset risk benchmark parameters by combining the business attribute information of the target business; then, the current risk assessment parameters are compared with the preset risk benchmark parameters to obtain the current risk level of the target business.
[0066] Furthermore, a risk management plan can be determined by combining the current risk level of the target business; then, a risk assessment result for the target business can be generated based on the risk management plan and the current risk level.
[0067] The aforementioned risk assessment method based on big data analytics acquires business activity data, business resource data, business approval-related data, and external supervision data of the target business from various business-related platforms. Feature extraction is then performed on this multi-source business-related data to obtain multi-dimensional business characteristics of the target business. Subsequently, a graph attention network is used to process these multi-dimensional business characteristics to obtain the current risk assessment parameters for the target business. Based on these current risk assessment parameters and preset risk benchmark parameters, the risk assessment result for the target business is determined. Compared to related technologies that rely on manual screening to analyze business risks contained in multi-source business-related data, this method, combined with a graph attention network, performs a full-link analysis of the multi-dimensional business characteristics of multi-source business-related data. This effectively captures data changes along the link, obtaining more accurate current risk assessment parameters, thereby improving the accuracy of subsequent data risk assessments.
[0068] Based on the above embodiments, this application provides an optional method for obtaining multi-source business-related data, specifically including the following steps:
[0069] The first step is to connect to the business management system via an interface call protocol to collect business activity data.
[0070] The business activity data includes business travel records, business reception registration information, and business procurement filing information.
[0071] Specifically, by using an interface call protocol to connect with the business management system, and through a pre-defined list of data collection fields, relevant information about business activities is collected in batches. These fields include: for business travel records, the traveler, destination, travel time, mode of transportation, accompanying persons, reason for travel, and approval authority; for business reception registration information, the recipient, reception time, reception location, number of guests, accompanying persons, reception resource limitations, and actual reception information; and for business procurement filing information, the procurement project name, procurement resource limitations, procurement method, supplier information, procurement officer, approval process nodes, and procurement contract number. During the collection process, a field mapping mechanism ensures that the data returned by the system interface matches the pre-defined fields. An incremental collection strategy is used to acquire only newly added and changed data. The collection frequency is set to a full synchronization of the previous day's data at 2 AM daily. The collected data is categorized and labeled according to "travel / reception / procurement," forming business activity data containing 12 core fields and an average of 320 records per day, ensuring that the data can support subsequent risk assessment and processing.
[0072] The second step involves connecting the resource payment system and the bank's platform through a standardized data interaction interface to collect business resource data.
[0073] Business resource data includes resource delivery records, reimbursement voucher information, account transaction data, resource approval documents, etc.
[0074] Specifically, a standardized data exchange interface can be used to connect the resource payment system and the bank's platform. From the resource payment system, the system can collect data on resource delivery records, including delivery value, recipient, delivery time, resource purpose, delivery approval number, and handler; reimbursement voucher information, including the reimbursing person, reimbursement amount, reason for reimbursement, number of attachments, reimbursement category, review stage, and payment method; and resource approval documents, including approval process number, application amount, approval level, approval opinions at each level, and approval time. From the bank's platform, the system can collect account transaction data, including transaction time, transaction amount, counterparty, transaction summary, account name, transaction type, and resource flow.
[0075] During the data collection process, a data verification mechanism was activated to check the consistency of the values and times of the resource allocation records and account transaction data in the resource payment system, ensuring that the resource transfer trajectory is complete and traceable. Electronic data on resource transfer was constructed, which includes 15 core fields covering the entire process of resource transfer. The data spans nearly 3 years, with an average monthly data volume of 4,500 records, providing basic data support for the subsequent identification of risks associated with abnormal resource transfers.
[0076] The third step involves collecting business approval-related data through the data sharing channel of the operation and supervision platform.
[0077] The business approval-related data includes approval process records, business supervision process-related data, project management process documents, etc.
[0078] Specifically, by relying on the data sharing channel of the operation and supervision platform and adopting a process node decomposition and collection method, comprehensive data related to business approval can be collected. The fields collected for approval process records include the name of the approval item, the applicant, the application time, the acceptance time, the person in charge of each approval stage, the processing time, the approval opinion, the approval result, the completion time, the fee standard, and the resource value. Data related to the business supervision process includes the case number, case type, case handler, case filing time, investigation record, basis for punishment, punishment resource value, relevant personnel in the case, the case approval process, and the case closure time. Project management process documents include project initiation approval documents, project resource restriction information, project leader, project progress nodes, resource allocation progress, and project acceptance information.
[0079] During the data collection process, process nodes are labeled according to the three types of data mentioned above, and the operation objects and timestamps of each node are recorded to form business approval-related data containing 20 core fields and covering the entire process. By sorting out the process sequence, the traceability of data in each link is ensured, providing data support for risk identification.
[0080] The fourth step involves collecting external supervision data by connecting to the supervision platform, monitoring system, and third-party supervision feedback channels via API interfaces.
[0081] External oversight data includes information related to business reports, leads, and oversight and inspection records.
[0082] Specifically, the system connects to the oversight platform and monitoring system via API interfaces, and simultaneously establishes a data receiving port for third-party oversight feedback channels to collect external oversight data. This includes collecting information from the oversight platform such as the reporting time, method, content, target, type of tip, and verification status; using keyword extraction technology, the monitoring system sets core keywords to capture tips from various online platforms in real time, including publication time, platform, content, involved parties, dissemination scope, and popularity; and collecting inspection records from third-party oversight feedback channels, including inspection time, scope, inspectors, description of problems found, resources involved, and rectification requirements.
[0083] During the data collection process, unstructured information is processed using text structuring to extract core information elements. The collected data is then categorized by clue level (general / important / urgent) to form fragmented monitoring data. This ensures that no clues are overlooked and provides diverse data support for identifying potential risks.
[0084] The fifth step involves removing duplicate data and filtering invalid data from the collected data, standardizing the data using a unified data format, and performing time-series calibration of multi-source data through timestamp matching to generate standardized multi-source business-related data.
[0085] Specifically, data cleaning algorithms can be used to process the collected raw data, duplicate data can be removed by field hash comparison, and invalid value filtering rules can be used to remove raw data with missing core fields and logical contradictions. The invalid data filtering ratio should be controlled within 3%. After the data is processed and standardized according to a unified format, for example, the date field in the data can be unified into the format "YYYY-MM-DD HH:MM:SS", the resource value field can be retained to two decimal places, and the time sequence calibration of multi-source data can be achieved by timestamp matching, thereby obtaining multi-source business-related data.
[0086] In this embodiment, by comprehensively collecting multi-dimensional data such as business activity data, business resource data, business approval-related data, and external supervision data, the limitations of traditional manual investigation and periodic audits, which rely solely on partial data, are overcome, achieving full coverage of risk-related data. Standardized processing unifies the format, definitions, and statistical standards of different data types, generating standardized multi-source business-related data. This effectively eliminates the drawbacks of disorganized and difficult-to-integrate raw data, ensuring data consistency and usability. Compared to traditional scattered and fragmented data formats, this provides high-quality, multi-dimensional data support for subsequent in-depth mining of related risk points, improving the systematicness and comprehensiveness of risk assessment from the source and avoiding the omission of risk points due to missing data or inconsistent formats.
[0087] Based on the above embodiments, this application provides an optional method for determining multi-dimensional business characteristics, such as... Figure 2 As shown, the specific steps include:
[0088] S201, extract features from business activity data to obtain abnormal business frequency features, business object association features, and business process deviation features.
[0089] Among them, the abnormal business frequency feature represents the business-related features that occur abnormally in the target business; the business object association feature represents the association relationship between various objects in the target business; and the business process deviation feature represents the degree of deviation between the target business and the corresponding standard business process.
[0090] Optionally, time-series analysis can be used to process business activity data, setting two time windows: 7 days and 30 days. Data on the frequency of business trips, receptions, and purchases by the same entity within different windows can be statistically analyzed, with an abnormal frequency threshold set at ±2 times the average frequency for the same period over the past three years. For example, the abnormally high business frequency characteristic of each department's responsible personnel (6 business trips in 7 days, average 2 trips) and the abnormally low business frequency characteristic of a certain business department (no business purchase records in 30 days, average 8 purchases) can be extracted. By using geocoding matching technology to mark the coordinates of business destinations, the cross-regional dense characteristic of each entity conducting business across 5 provinces within 15 days can be identified, forming abnormal business frequency characteristics.
[0091] By employing association rule mining algorithms and combining them with identity information verification mechanisms, we can mine the relationships between business participants and business objects, such as kinship and interest relationships, and obtain the association features of business objects.
[0092] Establish a standard business process model, compare the actual business process of the target business with the standard model, and identify information such as missing approval steps in a certain business, reversed order of execution before approval in a certain procurement process, and overdue processing steps of an administrative approval exceeding the prescribed time limit by 15 working days, thereby constructing business process deviation characteristics.
[0093] For example, business activity data can be grouped based on the business entity dimension to construct a business time-series sequence for each entity. Specifically, by employing a entity field grouping algorithm, business activity data is processed, using the unique identifier (department code + personnel number) of the business participant entity as the grouping key to aggregate and categorize all business activity records (travel, reception, procurement) of the same entity. Each group of data is sorted in ascending order by timestamp to construct a business time-series sequence for each entity. The sequence field includes core information such as time node, business type, business frequency, business object, and business location. During the grouping process, data integrity checks are performed to ensure that the time-series sequence of each entity has no time breaks, and that multiple business records at the same time node are fully included, ultimately forming a collection of individual business time-series sequences covering all business entities and spanning nearly 3 years.
[0094] This approach allows for the analysis of regular patterns in the frequency of various business activities, identifying the normal frequency range for different time periods. Specifically, by employing a time-series trend fitting method, the regular patterns in the frequency of each business activity are analyzed. A sliding window analysis is performed on the time-series business data of each activity, with a window size of 7 days and a step size of 1 day, to calculate the average frequency within each window. Combining annual, quarterly, and monthly time dimensions, the distribution characteristics of business frequency across different time periods are statistically analyzed, such as the frequency differences between weekdays and weekends, holidays and weekdays, the beginning and end of the year and the middle of the year, and peak and off-peak seasons. A statistical distribution fitting algorithm is used to determine the normal frequency range for different time periods, using the mean frequency ± 2 standard deviations for each time period as the boundary of the normal range. A frequency benchmark table is generated, including the activity identifier, time period type, upper limit of normal frequency, and lower limit of normal frequency. For example, the normal business travel frequency range for a certain activity on weekdays is 1-3 times / day, and the normal procurement frequency range during the peak season of project approval at the beginning of the year is 5-10 times / month. Clearly defining the frequency benchmarks for each time period provides a basis for subsequent anomaly identification.
[0095] Next, the actual business frequency of each entity is compared with the normal business frequency range for the corresponding time period to identify abnormal increases or decreases in frequency that exceed the range, generating abnormal business frequency characteristics and corresponding abnormal time period identifiers. For example, by using a frequency threshold comparison method, the actual business frequency of each entity is compared with the normal frequency range for the corresponding time period in the frequency benchmark table on a time-by-time basis. If the actual frequency of a certain entity exceeds the upper limit of the normal range in a certain time period, it is marked as an abnormal increase in frequency; if it exceeds the lower limit of the normal range and the duration exceeds 3 consecutive windows, it is marked as an abnormal decrease in frequency. Abnormal time period identifiers (abnormal start time, abnormal duration) are added to the identified abnormal situations, generating abnormal business frequency characteristics that include entity identifier, abnormal type, abnormal time period, actual frequency, and normal range.
[0096] Extract information on participating entities and business objects from business activities to construct a subject-object association network. Specifically, by employing information extraction algorithms, extract information on participating entities and business objects for each business activity from standardized business activity data, and remove invalid object information (such as fuzzy records without specific object labels). Using participating entities and business objects as nodes and business activities as edges, construct a subject-object association network. The attributes of the edges include information such as business type, business time, business frequency, and business amount.
[0097] This process involves uncovering indirect links between participating entities and business objects, identifying potential interest connections corresponding to kinship and business partnerships, and generating business object association features. Specifically, an association link mining algorithm traverses the entity-object association network to uncover indirect links between participating entities and business objects, i.e., the association path from entity A to intermediate node to business object B. By connecting to an identity information database, user-related information at each node in the link is compared to identify object associations. The identified association links are verified, and false associations are eliminated, ultimately generating business object association features that include associated entities, associated objects, association links, and association types.
[0098] The process involves retrieving standard process specifications for various business types and comparing them step-by-step with actual business processing records. Missing, reversed, or overdue steps are identified, and deviations are identified to construct a business process deviation feature. Specifically, a structured parsing method is used to retrieve standard process specifications for each business type, breaking them down into a standard process node list that includes the process step name, step order, responsible party, and processing time limit. For example, a standard business reception process might include four steps: "application-approval-implementation-reimbursement," with an approval processing time limit of one working day. Using this standard process node list as a benchmark, each step is compared with actual business processing records. Sequence comparison algorithms are used to determine the completeness and order of steps, and time limit difference calculations are used to determine the reasonableness of processing time limits. Information on marked abnormal steps is summarized and associated with the corresponding business record's subject, time, and object, constructing a business process deviation feature that includes process type, abnormal step, abnormal type, and violation duration / deviation details.
[0099] S202, based on business resource data and target business information, determine resource flow matching characteristics, abnormal resource fluctuation characteristics, and resource approval link characteristics.
[0100] Among them, resource flow matching features characterize the degree of matching between the resource flow direction and the objective of the target business; abnormal resource fluctuation features characterize abnormal fluctuations in resources within the target business; and resource approval link features characterize the approval link-related features of resources within the target business. Business information may include basic business attribute information and business processing flow information of the target business.
[0101] Optionally, by employing a field association matching method, using the business activity number, participants, and timestamp as association keys, structured business resource data and target business information are matched, comparing the resource usage with the business purpose to obtain resource flow matching features. For example, features that do not match the business purpose when the resource flow is "equipment procurement" but the corresponding business record is "meeting reception" can be extracted.
[0102] Using a time-series trend analysis method, we fitted the monthly resource allocation change curves for similar businesses over the past three years, setting the normal fluctuation range to ±15% of the curve mean, and constructed abnormal resource fluctuation characteristics. For example, we identified an abnormal increase in business reception resources in a certain month, which increased by 42% compared to the average, and an abnormal decrease in resource allocation for a certain project, which decreased by 30% month-on-month for three consecutive months.
[0103] By employing a process chain decomposition method, the entire approval chain from resource application to payment is analyzed, a complete list of approval nodes is constructed, and actual approval records are compared to identify resource approval chain characteristics. For example, features such as missing approval nodes (e.g., a voucher lacking department head approval), excessive approval authority (e.g., resource allocation being directly approved by an unauthorized entity), and abnormal approval time (e.g., a resource approval taking 28 working days (standard 7 working days) are extracted.
[0104] S203, based on business approval-related data and external supervision data, determine the risk characteristics of business nodes and the correlation characteristics of supervision feedback.
[0105] Among them, the business node risk characteristics represent the risk status of each node in the target business; the supervision feedback correlation characteristics represent the feedback information of external supervision.
[0106] Optionally, a node extraction algorithm can be used to break down the entire business operation process from business approval-related data, locate key decision nodes (such as project initiation approval and large-scale resource approval), approval nodes (such as administrative approval acceptance and law enforcement penalty decision), and execution nodes (such as project implementation and penalty execution), and construct a node graph.
[0107] Next, a case matching and comparison method was used to associate historical business-related data with the node graph, thereby uncovering high-risk characteristics in the target business, i.e., business node risk characteristics. Then, using the timestamps and involved entities of each node as association dimensions, fragmented external supervision data was matched with the nodes to generate supervision feedback association characteristics.
[0108] After identifying the features in each dimension, the features (abnormal business frequency features, business object association features, business process deviation features, resource flow matching features, abnormal resource fluctuation features, resource approval link features, business node risk features, and supervision feedback association features) can be normalized, and redundant features can be eliminated by feature association degree screening, and multi-dimensional business features can be integrated.
[0109] For example, the above features can be processed using the min-max normalization method, mapping feature values of different dimensions to the [0,1] interval. For instance, the business frequency (values 1-10 times) and resource amount (values 1000-100000 yuan) in the above features can be uniformly normalized. The correlation between each feature is calculated using the Pearson correlation coefficient analysis method, with a correlation threshold of 0.8, and highly redundant features are removed. The remaining effective features are integrated through a feature fusion algorithm to construct multi-dimensional business features. Among them, the multi-dimensional business features can include 3 primary categories, 12 secondary categories, and 45 tertiary features. The primary categories are abnormal business behavior features (abnormal business frequency features, business object association features, and business process deviation features), abnormal resource flow features (resource flow matching features, abnormal resource fluctuation features, and resource approval link features), and operational risk features (business node risk features and supervision feedback association features). The secondary categories include frequency anomalies, association anomalies, process deviations, etc., and the tertiary features are specific abnormal manifestation features.
[0110] In addition, feature type tags and risk association identifiers can be added to multi-dimensional business features. Specifically, a feature classification and labeling method can be used to add tags to various features in the multi-dimensional business features, and feature type tags can be added according to feature attributes. For example, "abnormally high business frequency feature" can be labeled as "time-series behavior type", "resource flow mismatch with business purpose feature" can be labeled as "resource association type", and "node-supervision issue association feature" can be labeled as "supervision type", for a total of 6 feature type tags. Combined with the historical risk case database, risk association identifiers can be added to each feature. The identifier content includes the associated risk type, risk level and case number.
[0111] In this embodiment of the application, by extracting features from multi-source business-related data and integrating the obtained features to form multi-dimensional business features, the limitations of traditional single-dimensional risk assessment are broken. This allows for the capture of hidden risk points that are interconnected between different links, thereby ensuring the accuracy of data risk assessment.
[0112] Based on the above embodiments, this application provides a method for determining optional features of business object association, specifically including the following steps:
[0113] The first step is to label the nodes in the subject-object association network with attributes, including subject identity information, object attribute information, and associated scenario information.
[0114] Specifically, node attribute extraction and annotation methods are used to comprehensively annotate the nodes in the subject-object association network. Subject identity information annotation includes a unique identifier (department code + personnel number), name, department, position, job responsibilities, and permission level. Object attribute information annotation distinguishes between organizational and individual objects. Organizational objects are annotated with the organization name, object identifier, industry, business scope, and business history with the business entity. Individual objects are annotated with name, object identifier, occupation, and affiliated organization. Association scenario information annotation includes business type, business frequency, first cooperation time, most recent cooperation time, and average cooperation amount, ensuring the completeness of attribute information for each node and providing basic attribute support for subsequent association weight calculation and path mining.
[0115] The second step is to construct association weight calculation rules based on node attribute information and assign basic weights to directly associated links.
[0116] Specifically, a multi-dimensional association weight calculation rule is constructed based on node attribute information. This rule covers four core dimensions: business frequency, cooperation duration, transaction amount, and permission association degree. Business frequency weight is assigned based on annual cooperation frequency; transaction amount weight is assigned based on the total annual cooperation amount; and permission association degree weight is assigned based on the degree of association between the entity's permissions and the object's business: direct approval association is assigned a value of 0.2, indirect supervision association is assigned a value of 0.1, and no direct association is assigned a value of 0.05. The weights of the four dimensions of the directly associated links are summed to obtain the basic weight.
[0117] The third step is to mine the indirect connection paths in the network and calculate the cumulative degree of connection of the indirect connection paths.
[0118] Specifically, a depth-first search algorithm is used to mine indirect connection paths in the network, with a maximum path depth of 3 layers (i.e., main body - intermediate node 1 - intermediate node 2 - business object), traversing all indirect connection relationships of main body nodes. For each mined indirect connection path, the cumulative correlation degree is calculated based on the basic weights of the direct connection links in each segment of the path. The cumulative correlation degree is equal to the sum of the products of the basic weights of the direct connection links in each segment.
[0119] For example, attribute annotation is performed on nodes in the subject-object association network to extract permission parameters of participating subjects, business scope parameters of business objects, and subject resource exchange frequency parameters from electronic resource transfer data, generating a node attribute parameter set. Specifically, by employing a node attribute extraction algorithm, attribute annotation is performed on all nodes in the subject-object association network. Permission parameters of participating subjects are extracted from the standardized business dataset, including permission codes, approval limit limits, approval item types, and regulatory scope. Subject resource exchange frequency parameters are extracted from structured business resource data, including total annual exchanges, average monthly exchanges, and high-frequency exchange periods. The three types of parameters are integrated in the format of "unique node identifier - permission parameter - business scope parameter - resource exchange frequency parameter," and nodes with missing parameters are removed to generate a node attribute parameter set.
[0120] Furthermore, association relationship determination rules are constructed based on the node attribute parameter set, using the overlap between permission parameters and business scope parameters as the basis for direct association determination, generating a set of direct association relationships between nodes. Specifically, multi-dimensional association relationship determination rules are constructed based on the node attribute parameter set, with the overlap between permission parameters and business scope parameters as the core basis for direct association determination, while resource exchange frequency parameters are used for auxiliary verification. The overlap calculation adopts a semantic similarity matching algorithm, semantically comparing the approval item type and regulatory scope in permissions with the core business category and service area in the business scope of business objects, setting an overlap of ≥60% as the threshold for direct association determination. Overlap calculation and auxiliary verification are performed on all pairs of nodes, generating a set of direct association relationships between nodes that includes associated node pairs, overlap values, resource exchange verification results, and association determination results.
[0121] Furthermore, based on the set of direct relationships between nodes, and combined with node collaboration records in business approval-related data, indirect relationship paths in the relationship network are mined. The number of intermediate nodes and the type of business interaction between nodes are extracted for each indirect path to generate an indirect relationship path feature set. Specifically, based on the set of direct relationships between nodes, a breadth-first search algorithm is used to mine indirect relationship paths in the relationship network, setting the maximum path depth to 4 layers (subject - intermediate node 1 - intermediate node 2 - intermediate node 3 - business object) to ensure coverage of multi-level collaboration scenarios. During the mining process, node collaboration records in business approval-related data are combined to extract information such as collaboration timestamps, items, and responsibility allocation ratios to filter out indirect relationship paths with actual collaborative behavior. For each mined indirect relationship path, core feature parameters are extracted: the number of intermediate nodes (the total number of nodes in the path excluding the subject and business object) and the type of business interaction between nodes (divided into four categories: approval collaboration, project collaboration, resource collection / payment, and information transmission, assigned values of 1.0, 0.8, 0.9, and 0.5 respectively).
[0122] Furthermore, based on the feature set of indirect association paths, the number of intermediate nodes characterizes the path transmission attenuation level, and the business interaction type parameter characterizes the path association tightness. The cumulative association degree of each indirect association path is generated through a coupled calculation between the path transmission attenuation level and the path association tightness. Specifically, a cumulative association degree calculation model is constructed based on the indirect association path feature set, clarifying the coupled calculation logic between the path transmission attenuation level and the path association tightness. The path transmission attenuation level is determined by the number of intermediate nodes, with attenuation coefficient calculation rules set as follows: 0.9 for 1 intermediate node, 0.7 for 2, and 0.5 for 3, with a smaller attenuation coefficient for more nodes. The path association tightness is obtained by weighted summation of the business interaction type parameters between nodes, with the weights being the importance proportion of each interaction segment in the path (averaged across the path length). Cumulative association degree = Path association tightness × Path transmission attenuation level.
[0123] The fourth step is to set the correlation filtering criteria and filter out indirect correlation paths that meet the cumulative correlation criteria.
[0124] Specifically, by combining historical risk case data to set correlation screening criteria, and by statistically analyzing the cumulative correlation distribution of indirect correlation paths in historical cases, a screening threshold of 0.3 was determined, meaning only indirect correlation paths with a cumulative correlation score ≥ 0.3 were retained. All mined indirect correlation paths were compared using cumulative correlation scores, and paths with a cumulative correlation score below 0.3 were eliminated. Simultaneously, the validity of the selected paths was verified, checking the authenticity of the connections between each node in the path and eliminating false correlation paths (such as incorrect connections due to data entry errors), ensuring that the screened indirect correlation paths have actual correlation significance and providing path data for subsequent correlation identification.
[0125] The fifth step is to perform path parsing on the selected indirect association paths, identify the core association nodes and association types in the paths, and determine the potential association relationships between participating entities and business objects.
[0126] Specifically, a path parsing algorithm is used to perform structured parsing on the selected indirect association paths, extracting all nodes in the paths and the connections between nodes, and locating core association nodes (i.e., intermediate nodes in the path that are strongly associated with both the subject and the object; the criterion for this is that the basic weight of this node with both of its two endpoints is ≥0.5). By connecting to the identity information database, the relationship types between the core association nodes and the subject and the object are verified.
[0127] The sixth step is to associate the identified potential relationships with the corresponding business activity records to generate business object association features that include association type, association degree and corresponding business information.
[0128] Specifically, by using the business activity number, participating entity, and timestamp as association keys, the identified potential associations are linked to the corresponding business activity records in the standardized business activity data. The association type, cumulative association degree, core association node information, and detailed information of the corresponding business activity (business type, business time, business amount, approval process) are integrated to generate standardized business object association features. These feature fields include entity identifier, business object identifier, association path, core association node, association type, cumulative association degree, corresponding business number, business type, business amount, and occurrence time.
[0129] Based on the above embodiments, this application provides an optional method for determining current risk assessment parameters, such as... Figure 3 As shown, the specific steps include:
[0130] S301 uses a graph attention network to determine the feature association matrix based on the relationships between multi-dimensional business features.
[0131] The feature association matrix is used to characterize the relationships between features in a multi-dimensional business feature set. Matrix elements represent the strength of the association between features.
[0132] Optionally, multi-dimensional business features can be input into a graph attention network, and the correlation and transmission relationships between features can be mined by calculating the attention weights of nodes, and only valid correlation links can be retained to generate a feature correlation matrix.
[0133] S302, Based on historical business risk data, process the feature correlation matrix to obtain the target correlation matrix.
[0134] Historical business risk data can be understood as business risk data related to the target business within a historical period. The target correlation matrix can be understood as the correlation matrix obtained after processing.
[0135] Optionally, the matrix elements in the feature association matrix can be updated based on historical business risk data to obtain the target association matrix.
[0136] S303, Based on the target correlation matrix, determine the current risk assessment parameters for the target business.
[0137] Optionally, the target correlation matrix can be input into a gradient boosting tree model. The model's built-in feature importance learning rules can then be used to extract risk correlation features from the target correlation matrix. Subsequently, the current risk assessment parameters for the target business can be calculated based on these risk correlation features. Here, risk correlation features can be understood as business features that present business risks.
[0138] In this embodiment of the application, by processing the feature correlation matrix between multi-dimensional business features based on historical business risk data, the current risk assessment parameters are obtained, which can ensure the accuracy of the determination of the current risk assessment parameters.
[0139] Based on the above embodiments, this application provides an optional method for determining the target association matrix, such as... Figure 4 As shown, the specific steps include:
[0140] S401, extract each feature association link in the feature association matrix, and the feature association type of each feature association link.
[0141] Feature association type can be understood as the type of association between features, such as unidirectional association and bidirectional association.
[0142] Optionally, a matrix element parsing algorithm can be used to extract each feature association link in the feature association matrix, as well as the association link information of each feature association link. The association link information includes feature node pairs, link connection direction, and feature association type.
[0143] S402, Based on historical business risk data, determine the historical association frequency and historical association strength of each characteristic association link.
[0144] Among them, historical association frequency can be understood as the frequency of occurrence of feature association links within a historical period; historical association strength can be understood as the association strength between each feature in the feature association links within a historical period.
[0145] Optionally, a case-matching counting method can be used in conjunction with historical business risk data to count the number of times each related link appears in the historical business risk data, generating historical association frequency; a risk intensity quantification method can be used to determine the historical association intensity by combining the lost resources and impact scope corresponding to the links in the historical business risk data.
[0146] S403, for each feature association link, determine the feature association confidence level of the feature association link based on the feature association type, historical association frequency, and historical association strength of the feature association link.
[0147] Among them, feature association confidence can be understood as the reliability of feature association links.
[0148] Optionally, for each feature-related link, a preset calculation logic can be used to calculate the feature association confidence level of that link based on its feature association type, historical association frequency, and historical association strength. The calculation logic is as follows: Feature association confidence level parameter = (Historical association frequency / Total number of links in historical business risk data) × Historical association strength × Correction coefficient. The correction coefficient is determined based on the feature type corresponding to the feature-related link.
[0149] S404, the feature association matrix and the feature association confidence of each feature association link are fused to obtain the target association matrix.
[0150] Optionally, a matrix fusion reconstruction algorithm can be used to integrate the feature association matrix and the feature association confidence of each feature association link. The row / column feature node structure of the feature association matrix is retained, and the binary values (0 or 1) representing the existence of association in the feature association matrix are replaced with the corresponding feature association confidence. Positions without association links are still filled with 0. During the reconstruction process, matrix integrity is checked to ensure that all extracted feature association links have been filled with confidence and there are no omissions or mismatches. Finally, a target association matrix with confidence is generated.
[0151] In this embodiment of the application, the feature association matrix is processed by combining the feature association confidence of each feature association link to obtain the target association matrix.
[0152] Based on the above embodiments, this application provides an optional method for determining current risk assessment parameters, such as... Figure 5 As shown, the specific steps include:
[0153] S501 uses a gradient boosting tree model to process the target association matrix and obtain risk association features.
[0154] Optionally, a gradient boosting tree model can be constructed first. The model structure includes an input layer, 10 gradient boosting decision tree weak classifier layers, and an output layer. The model has a built-in feature importance learning rule, which represents the importance by calculating the information gain value of each feature association link. The information gain value = feature association confidence of the association link × risk impact weight of the link. The risk impact weight is determined based on the probability of the link causing risk in historical business risk data.
[0155] The target association matrix can be input into a gradient boosting tree model, and 10 weak classifiers can be trained iteratively using gradient descent. The number of iterations is set to 100, and the learning rate is set to 0.1. The feature importance weights are updated in each iteration. After training, the top 15 risk association features with the highest information gain values are extracted, and risk association feature mapping rules containing risk association features, information gain values, feature association confidence, and corresponding risk types are generated.
[0156] S502, determine the current risk assessment parameters of the target business based on the matrix parameters of the links where the risk association features are located in the target association matrix.
[0157] The matrix parameters include feature association confidence and information gain value; the information gain value is determined based on the feature association confidence of the link and the risk probability of the link in historical business risk data.
[0158] Optionally, the current risk assessment parameters for the target business can be calculated by combining the matrix parameters of the links containing the risk-related features in the target association matrix. For example, the feature association confidence and information gain values can be weighted to obtain the current risk assessment parameters.
[0159] For example, the feature association links in the target association matrix are matched with the risk association features in the mapping rules, and the feature association confidence and information gain value of the matched links are extracted. Then, a weighted summation formula is used to calculate the real-time risk assessment value, which is: Risk assessment value = Σ (feature association confidence × information gain value × risk level weight). The risk level weight can be determined by combining historical business risk data, with a high risk level weight of 1.0, a medium risk level weight of 0.7, and a low risk level weight of 0.4.
[0160] In this embodiment of the application, the current risk assessment parameters are calculated by using the matrix parameters of the links where the risk association features are located in the target association matrix, which can ensure the accuracy of the current risk assessment parameter calculation.
[0161] Based on the above embodiments, this application provides an optional method for determining risk assessment results, such as... Figure 6 As shown, the specific steps include:
[0162] S601, based on the current risk assessment parameters and preset risk benchmark parameters, determine the current risk level and risk items in the target business.
[0163] Here, the current risk level can be understood as the risk level of the target business. A risk event can be understood as a business event that causes risk to the target business. The target business contains multiple business events. Preset risk benchmark parameters are determined based on the business domain of the target business and the object hierarchy of the objects related to the target business. For example, by using a domain-hierarchy classification modeling method, the business domain is divided into four categories: procurement management, project approval, business supervision, and business reception. The hierarchy is divided into high-level, mid-level, and low-level to construct 12 domain-hierarchy combination scenarios, and corresponding preset risk benchmark parameters are configured for each scenario.
[0164] Optionally, the current risk assessment parameters can be compared with the corresponding preset risk benchmark parameters to obtain the current risk level; then, risk items in the high-risk range can be located from each business item based on the risk assessment parameters associated with each business item in the target business.
[0165] S602, construct a risk transmission map based on the business characteristics associated with risk events in the multidimensional business characteristics.
[0166] Among them, constructing a risk transmission map can be understood as the risk transmission between business characteristics in risk events.
[0167] Optionally, the business characteristics corresponding to the risk event can be traced in the multi-dimensional business characteristics, and then the responsible entity corresponding to the risk event can be determined by combining the business characteristics; with the risk event as the core node and the responsible entity and business characteristics as the related nodes, a risk transmission map can be constructed.
[0168] For example, an initial graph is constructed using risk events as the core node, responsible entities and business characteristics as first-level related nodes, and the original data corresponding to the business characteristics as second-level related nodes. Directed edges between nodes are used to indicate the direction of propagation, and edge attributes are used to indicate the confidence level of feature association.
[0169] Furthermore, based on the association links and feature association confidence levels in the target association matrix, an association strength transmission algorithm can be used to analyze the transmission association strength between nodes in the initial graph. Here, transmission association strength = feature association confidence level between nodes × risk impact coefficient, where the risk impact coefficient is determined according to the feature type: 1.2 for resource features, 1.0 for process features, and 1.1 for association features. Node links with a transmission association strength ≥ 0.3 are retained, while weakly associated nodes are removed, generating a risk transmission graph. The graph clearly presents the transmission path from core node → first-level node → second-level node, marking the transmission strength and confidence level of each path, intuitively demonstrating the transmission context of risk from features to events, and from subjects to data.
[0170] S603, determine risk propagation information based on preset risk diffusion parameters and risk transmission maps.
[0171] Risk propagation information can be understood as information related to risk propagation. Preset risk diffusion parameters are used to characterize the speed of risk propagation.
[0172] Optionally, preset risk diffusion parameters can be used to calculate the transmission links in the risk transmission map to obtain risk propagation information.
[0173] For example, three types of risk diffusion parameters can be set: approval node strengthening parameters, regulatory parameters, and related entity isolation parameters. Each type of parameter includes three gradient levels (weak / medium / strong). The gradient combinations of these intervention parameters are input into the risk transmission model. Based on the transmission intensity in the risk transmission map, the model sets the risk diffusion rate calculation rules, where diffusion rate = base diffusion rate × (1 - preset risk diffusion parameter), and the base diffusion rate is set to 0.2 / day. This simulates the risk diffusion speed, diffusion range, and convergence trend under different parameter combinations, obtaining risk propagation information.
[0174] S604. Based on the current risk level, risk transmission map, and risk propagation information, determine the risk assessment results for the target business.
[0175] Optionally, the core path in the risk transmission map can be extracted. Then, the corresponding processing strategy can be selected from the processing strategy library by combining the core path, risk propagation information and the current risk level. After that, the risk assessment result of the target business can be generated by combining the processing strategy, the current risk level, the risk transmission map and the risk propagation information.
[0176] For example, by importing the processing strategy, current risk level, risk transmission map, and risk propagation information into the assessment template, the risk assessment results for the target business can be obtained.
[0177] In this embodiment, the risk assessment results are determined by combining the current risk level, risk transmission map, and risk propagation information. Compared with the traditional assessment results that only give qualitative conclusions about the risk, this has more practical guidance value, can clearly inform the key points of prevention and control, the responsible parties, and the specific intervention directions, and significantly improves the initiative and effectiveness of risk prevention and control.
[0178] Based on the above embodiments, this application provides an optional method for determining a risk transmission spectrum, specifically including the following steps:
[0179] The first step is to extract the relationship information between the risk event node and each associated node from the target relationship matrix, and use the relationship information between the risk event node and each associated node as the basis for initial transmission of relationships.
[0180] Optionally, a matrix association extraction algorithm can be used to selectively extract the association information between risk event nodes and each associated node from the target association matrix. The extracted information includes node pairs, association types, and basic association confidence levels. The extracted association information is then organized in the format of "risk event node - associated node - association type - basic confidence level" to form the initial transmission association basis.
[0181] The second step involves introducing a time decay factor and, in conjunction with the duration of the risk event, analyzing the degree of time-related impact of the transmission.
[0182] Optionally, a time decay factor can be introduced to construct a model for calculating the degree of time impact. The time decay coefficient is determined based on the interval between the duration of the risk event and the current time. The formula for calculating the decay coefficient is set as decay coefficient = 1 - (current time - event end time) / 365, with the time unit being days. If the interval exceeds 365 days, the decay coefficient is fixed at 0.3. The degree of time impact is calculated by combining the decay coefficient with the basic correlation confidence level: degree of time impact = basic correlation confidence level × decay coefficient, quantifying the weakening effect of the time dimension on the transmission correlation.
[0183] The third step is to analyze the transmission and impact scope of risk events based on the electronic records corresponding to the responsible entities.
[0184] Optionally, by employing record correlation analysis, electronic records of the responsible entity can be retrieved, including information such as job position, scope of approval authority, collaborating personnel, and areas of responsibility. Through scope mapping analysis, the scope of the risk event's transmission impact can be determined, and the relationship type between each node within the impact scope and the responsible entity can be marked, clarifying the boundaries and levels of the transmission impact.
[0185] The fourth step is to determine the dynamic transmission relationships between nodes by combining the initial transmission correlation basis, the degree of time influence, and the transmission influence range.
[0186] Optionally, a dynamic transmission association calculation model is constructed, integrating the basic association confidence level, time influence level, and transmission influence range weight from the initial transmission association basis to obtain the dynamic transmission association strength. Here, dynamic transmission association strength = basic association confidence level × time influence level × influence range weight. The dynamic transmission association relationship between nodes is determined based on the dynamic transmission association strength; a strength ≥ 0.1 is considered a valid dynamic association, forming a dynamic transmission association set containing node pairs, dynamic association strength, and association level.
[0187] The fifth step involves setting transmission association filtering conditions based on the dynamic transmission relationships between nodes, filtering out key transmission links that meet the conditions, determining the attribute type of each node and the transmission direction of the key transmission links, and constructing a risk transmission map that includes node attributes, transmission direction, and dynamic transmission associations.
[0188] Optionally, key transmission links can be identified by setting transmission association filtering conditions based on a dynamic transmission association set, with a filtering threshold of dynamic transmission association strength ≥ 0.2. A node attribute classification and labeling method is used to determine the attribute type of each node. The transmission direction of key transmission links is labeled using directed edges. By integrating node attributes, transmission direction, and dynamic transmission association strength, a risk transmission map is constructed. The map clearly presents the hierarchical transmission path from the responsible entity to the risk event and then to the risk characteristic, intuitively labeling the dynamic association strength of each link, providing a basis for risk transmission control.
[0189] Figure 7 This is a flowchart illustrating a risk assessment method based on big data analytics in another embodiment. Building upon the above embodiments, this embodiment provides an optional example of a risk assessment method based on big data analytics. (Combined with...) Figure 7 The specific implementation process is as follows:
[0190] S701, in response to a risk assessment request for the target business, obtains multi-source business-related data of the target business from various business-related platforms.
[0191] Among them, multi-source business-related data includes business activity data, business resource data, business approval-related data, and external supervision data.
[0192] S702 extracts features from business activity data to obtain abnormal business frequency features, business object association features, and business process deviation features.
[0193] S703 determines resource flow matching characteristics, abnormal resource fluctuation characteristics, and resource approval link characteristics based on business resource data and target business information.
[0194] S704, based on business approval-related data and external supervision data, determine the risk characteristics of business nodes and the correlation characteristics of supervision feedback.
[0195] S705 uses a graph attention network to process business frequency anomaly characteristics, business object association characteristics, business process deviation characteristics, resource flow matching characteristics, resource anomaly fluctuation characteristics, resource approval link characteristics, business node risk characteristics, and supervision feedback association characteristics to obtain the current risk assessment parameters of the target business.
[0196] Optionally, a feature association matrix is determined based on the correlation between multi-dimensional business features using a graph attention network; each feature association link in the feature association matrix and the feature association type of each feature association link are extracted; the historical association frequency and historical association strength of each feature association link are determined based on historical business risk data; for each feature association link, the feature association confidence level of the feature association link is determined based on the feature association type, historical association frequency, and historical association strength of the feature association link; the feature association matrix and the feature association confidence levels of each feature association link are fused to obtain the target association matrix.
[0197] Furthermore, a gradient boosting tree model is used to process the target association matrix to obtain risk association features. Based on the matrix parameters of the links containing the risk association features in the target association matrix, the current risk assessment parameters of the target business are determined. Among them, the matrix parameters include feature association confidence and information gain value. The information gain value is determined based on the feature association confidence of the link and the risk probability of the link in historical business risk data.
[0198] S706, based on the current risk assessment parameters and preset risk benchmark parameters, determine the current risk level and risk items in the target business.
[0199] The preset risk benchmark parameters are determined based on the business domain of the target business and the object hierarchy of the objects related to the target business.
[0200] S707, construct a risk transmission map based on the business characteristics associated with risk events in the multidimensional business characteristics.
[0201] S708, based on preset risk diffusion parameters and risk transmission maps, determines risk propagation information.
[0202] S709, based on the current risk level, risk transmission map and risk propagation information, determine the risk assessment results of the target business.
[0203] The specific processes of S701-S709 described above can be found in the description of the above method embodiments. Their implementation principles and technical effects are similar, and will not be repeated here.
[0204] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.
[0205] Based on the same inventive concept, this application also provides a risk assessment device based on big data analysis for implementing the risk assessment method based on big data analysis described above. The solution provided by this device is similar to the solution described in the above method; therefore, the specific limitations in one or more embodiments of the risk assessment device based on big data analysis provided below can be found in the limitations of the risk assessment method based on big data analysis described above, and will not be repeated here.
[0206] In one exemplary embodiment, such as Figure 8 As shown, a risk assessment device 1 based on big data analysis is provided, including: a data acquisition module 10, a feature extraction module 20, a parameter acquisition module 30, and a risk assessment module 40, wherein:
[0207] The data acquisition module 10 is used to respond to risk assessment requests for the target business by acquiring multi-source business-related data of the target business from various business-related platforms; wherein, the multi-source business-related data includes business activity data, business resource data, business approval-related data and external supervision data;
[0208] Feature extraction module 20 is used to extract features from multi-source business-related data to obtain multi-dimensional business features of the target business;
[0209] The parameter acquisition module 30 is used to process multi-dimensional business features through a graph attention network to obtain the current risk assessment parameters of the target business.
[0210] The risk assessment module 40 is used to determine the risk assessment result of the target business based on the current risk assessment parameters and the preset risk benchmark parameters.
[0211] In an exemplary embodiment, the multidimensional business features include abnormal business frequency features, business object association features, business process deviation features, resource flow matching features, abnormal resource fluctuation features, resource approval chain features, business node risk features, and supervision feedback association features; the feature extraction module 20 is specifically used for:
[0212] Feature extraction is performed on business activity data to obtain abnormal business frequency features, business object association features, and business process deviation features; based on business resource data and target business information, resource flow matching features, abnormal resource fluctuation features, and resource approval link features are determined; based on business approval-related data and external supervision data, business node risk features and supervision feedback association features are determined.
[0213] In one exemplary embodiment, the parameter acquisition module 30 includes:
[0214] The first processing unit is used to determine the feature association matrix based on the association between multi-dimensional business features through a graph attention network.
[0215] The second processing unit is used to process the feature correlation matrix based on historical business risk data to obtain the target correlation matrix.
[0216] The third processing unit is used to determine the current risk assessment parameters of the target business based on the target correlation matrix.
[0217] In one exemplary embodiment, the second processing unit is specifically used for:
[0218] Extract each feature association link from the feature association matrix, as well as the feature association type of each feature association link; determine the historical association frequency and historical association strength of each feature association link based on historical business risk data; for each feature association link, determine the feature association confidence level of the feature association link based on the feature association type, historical association frequency, and historical association strength; fuse the feature association matrix and the feature association confidence levels of each feature association link to obtain the target association matrix.
[0219] In one exemplary embodiment, the third processing unit is specifically used for:
[0220] A gradient boosting tree model is used to process the target association matrix to obtain risk association features. Based on the matrix parameters of the links containing the risk association features in the target association matrix, the current risk assessment parameters of the target business are determined. The matrix parameters include feature association confidence and information gain value. The information gain value is determined based on the feature association confidence of the link and the risk probability of the link in historical business risk data.
[0221] In one exemplary embodiment, the risk assessment module 40 is specifically used for:
[0222] Based on the current risk assessment parameters and preset risk benchmark parameters, determine the current risk level and risk items in the target business; wherein, the preset risk benchmark parameters are determined according to the business area where the target business is located and the object level of the related processing objects of the target business; construct a risk transmission map based on the business characteristics associated with the risk items in the multi-dimensional business characteristics; determine the risk propagation information based on the preset risk diffusion parameters and the risk transmission map; and determine the risk assessment result of the target business based on the current risk level, the risk transmission map, and the risk propagation information.
[0223] The modules in the aforementioned risk assessment device based on big data analysis can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device, or stored in the memory of a computer device as software, so that the processor can call and execute the corresponding operations of each module.
[0224] In one exemplary embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 9 As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs stored in the non-volatile storage media. The database stores business-related data. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network. When the computer program is executed by the processor, it implements a risk assessment method based on big data analysis.
[0225] Those skilled in the art will understand that Figure 9 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0226] In one embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above method embodiments.
[0227] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon that, when executed by a processor, implements the steps in the above method embodiments.
[0228] In one embodiment, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps in the above method embodiments.
[0229] It should be noted that the data involved in this application (including but not limited to business-related data) are all data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0230] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, artificial intelligence (AI) processors, etc., and are not limited to these.
[0231] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.
[0232] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A risk assessment method based on big data analysis, characterized in that, The method includes: In response to a risk assessment request for a target business, multi-source business-related data for the target business is obtained from various business-related platforms; wherein, the multi-source business-related data includes business activity data, business resource data, business approval-related data, and external supervision data; Feature extraction is performed on the multi-source business-related data to obtain the multi-dimensional business features of the target business; The multidimensional business features are processed using a graph attention network to obtain the current risk assessment parameters of the target business. Based on the current risk assessment parameters and the preset risk benchmark parameters, the risk assessment result of the target business is determined.
2. The method according to claim 1, characterized in that, The multidimensional business characteristics include abnormal business frequency characteristics, business object association characteristics, business process deviation characteristics, resource flow matching characteristics, abnormal resource fluctuation characteristics, resource approval link characteristics, business node risk characteristics, and supervision and feedback association characteristics. The step of extracting features from the multi-source business-related data to obtain multi-dimensional business features of the target business includes: Feature extraction is performed on the business activity data to obtain business frequency anomaly features, business object association features, and business process deviation features; Based on the business resource data and the business information of the target business, determine the resource flow matching characteristics, resource abnormal fluctuation characteristics, and resource approval link characteristics; Based on the business approval-related data and the external supervision data, the risk characteristics of business nodes and the correlation characteristics of supervision feedback are determined.
3. The method according to claim 1, characterized in that, The process of processing the multidimensional business features using a graph attention network to obtain the current risk assessment parameters of the target business includes: Using a graph attention network, a feature association matrix is determined based on the relationships between the multidimensional business features. Based on historical business risk data, the feature correlation matrix is processed to obtain the target correlation matrix; Based on the target correlation matrix, determine the current risk assessment parameters for the target business.
4. The method according to claim 3, characterized in that, The step of processing the feature correlation matrix based on historical business risk data to obtain the target correlation matrix includes: Extract each feature association link from the feature association matrix, and the feature association type of each feature association link; Based on historical business risk data, determine the historical association frequency and historical association strength of each characteristic association link; For each feature association link, the feature association confidence level of the feature association link is determined based on the feature association type, historical association frequency, and historical association strength of the feature association link. The feature association matrix and the feature association confidence of each feature association link are fused to obtain the target association matrix.
5. The method according to claim 3, characterized in that, The step of determining the current risk assessment parameters of the target business based on the target correlation matrix includes: A gradient boosting tree model is used to process the target association matrix to obtain risk association features; Based on the matrix parameters of the links containing the risk association features in the target association matrix, the current risk assessment parameters of the target business are determined; wherein, the matrix parameters include feature association confidence and information gain value; the information gain value is determined based on the feature association confidence of the link and the risk probability of the link in the historical business risk data.
6. The method according to any one of claims 1-5, characterized in that, The step of determining the risk assessment result of the target business based on the current risk assessment parameters and the preset risk benchmark parameters includes: Based on the current risk assessment parameters and the preset risk benchmark parameters, the current risk level and the risk items in the target business are determined; wherein, the preset risk benchmark parameters are determined according to the business domain in which the target business is located and the object hierarchy of the objects associated with the target business. Based on the business characteristics associated with the risk events in the multidimensional business characteristics, a risk transmission map is constructed; Based on the preset risk diffusion parameters and the risk transmission map, the risk propagation information is determined; Based on the current risk level, the risk transmission map, and the risk propagation information, the risk assessment result of the target business is determined.
7. A risk assessment device based on big data analysis, characterized in that, The device includes: The data acquisition module is used to respond to risk assessment requests for the target business by acquiring multi-source business-related data of the target business from various business-related platforms; wherein, the multi-source business-related data includes business activity data, business resource data, business approval-related data, and external supervision data; The feature extraction module is used to extract features from the multi-source business-related data to obtain multi-dimensional business features of the target business. The parameter acquisition module is used to process the multidimensional business features through a graph attention network to obtain the current risk assessment parameters of the target business. The risk assessment module is used to determine the risk assessment result of the target business based on the current risk assessment parameters and preset risk benchmark parameters.
8. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.