Risk assessment method and system for multi-source data fusion
By constructing a risk data system and using multi-source data fusion technology, integrating and linking data from different source systems, building a risk assessment model, and generating a comprehensive risk index, the problems of data dispersion and low utilization rate in existing technologies are solved, enabling accurate identification and quantitative assessment of risks.
Patent Information
- Application Number
- CN202510147635.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-11
- Publication Date
- 2025-11-25
AI Technical Summary
Existing risk assessment methods lack sufficient integration of multi-source data, failing to form a complete risk view. They also lack the ability to accurately identify key individuals, groups, and events, have weak dynamic trend analysis capabilities, lack unified quantitative assessment indicators and algorithms, and have an inadequate early warning mechanism, thus affecting the timeliness and effectiveness of emergency response.
By constructing a risk data system, integrating and linking data from different source systems, building multiple risk assessment models, generating a comprehensive risk index, and using machine learning and natural language processing technologies for data analysis and mining, potential risks can be identified and early warning information can be generated.
It enables efficient association and unified management of multi-source data, improves the ability to identify potential risks and sensitive behaviors, accurately identifies multiple scenarios, generates an accurate comprehensive risk index, and completes quantitative risk assessment.
Smart Images

Figure CN121010230A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of risk data analysis, in particular to a multi-source data fusion risk assessment method and system. BACKGROUND
[0002] In modern social governance, risk assessment is a core task. With the acceleration of urbanization, social contradictions are becoming increasingly diverse and complex, and governments and related organizations need to effectively grasp social dynamics, timely identify potential risks and take measures. However, the existing risk assessment methods have the following technical problems: insufficient multi-source data integration, risk-related data is usually scattered in multiple systems. These data are of various formats and heterogeneous sources, lack of unified integration and correlation mechanism, resulting in low data utilization and inability to form a complete risk view. Risk identification is not accurate, and existing technologies lack the ability to accurately identify key individuals, groups and events. For example, it is difficult to efficiently identify high-risk individuals who frequently complain or predict group events that may trigger large-scale participation, resulting in unreasonable allocation of management resources and limiting the effectiveness of governance. The dynamic trend analysis capability is weak, and the analysis means for the dynamic evolution trend of social events (such as time distribution, regional diffusion, and emotional fluctuation) is insufficient, which cannot accurately predict the development direction and possible subsequent impact of the event. Current risk assessment methods are mostly based on experience and lack of unified quantitative evaluation indicators and algorithms, making it difficult to provide scientific and quantifiable comprehensive evaluation results. The early warning mechanism for potential social risks is not yet perfect, and the existing early warning methods have problems such as response delay and unreasonable risk level division, affecting the timeliness and effectiveness of emergency response. Therefore, an efficient risk assessment method is crucial. SUMMARY
[0003] Therefore, the present application proposes a multi-source data fusion risk assessment method and system, which can perform data analysis based on the constructed risk data system and construct multiple risk assessment models, and then construct a risk assessment algorithm to complete risk assessment. The present application provides the following technical solutions: A multi-source data fusion risk assessment method, the method comprising: acquiring risk-related data and constructing a risk data system; integrating and correlating different source system data in the risk data system to construct a complete data view; performing data analysis and mining on the integrated and correlated data, extracting data features of subject events to obtain development trends, demand analysis and potential correlations of multiple subject events; constructing multiple risk assessment models based on the data analysis results, and finding target behaviors and generating warning information through the multiple risk assessment models; based on the output of the multiple risk assessment models, constructing a weighted calculation risk assessment algorithm to generate a risk comprehensive index and complete quantitative assessment of the risk.
[0004] Optionally, the risk-related data includes a full-volume data list and a range data list, the full-volume data list containing a complete data set of all relevant data items, records or entities, and the range data list including business data of a risk-related industry supervisor.
[0005] Optionally, the method of integrating and correlating different source system data in a risk data system to build a complete data view includes: data cleaning, data standardization and data matching of different source system data; obtaining association keys in different source system data, and integrating and correlating different source system data with the same association keys to form a data set representing the behavior, activity and event of the subject corresponding to the current association key in different systems; integrating the full-volume data list and forming a unified data view, extracting keywords in the data view as identifiers, and correlating different source system data; data cleaning and unifying the range data list, extracting key elements as retrieval elements, and correlating different source system data; based on the integrated and correlated different source system data, building a data view with three dimensions of matter, person and enterprise as the theme.
[0006] Optionally, the method of data analysis and mining of integrated and correlated data, extracting data features of subject events, to obtain the development trend, demand analysis and potential correlation of multiple subject events includes: data cleaning and standardization of data, specifically, data denoising and deduplication, data structuring and standardization, and missing value filling processing; extracting basic attributes of subject events, extracting dynamic features of data from event dimension and spatial dimension; extracting keywords and theme content of subject events, and identifying data content tendency and key information combined with machine learning technology to obtain content features of data; capturing the development trend of subject events through statistical models according to the dynamic features and content features of data; identifying the demand analysis of subject events according to the content features of data with natural language processing and sentiment analysis technology; according to the dynamic features of data, constructing the correlation network between subject events through graph network analysis technology, and identifying key nodes and paths in the network to obtain deep correlation and mutual influence between multiple subject events, thereby obtaining the potential correlation of multiple subject events.
[0007] Optionally, the method of building multiple risk assessment models based on data analysis results, and finding target behavior and generating warning information through multiple risk assessment models includes: identifying multiple preset scenes and collecting data corresponding to the preset scenes from the data analysis results; feature extraction and index construction of the collected data; building corresponding assessment models according to different preset scenes; identifying the output results of each assessment model based on multiple assessment models; finding target behavior and generating warning information through the output results.
[0008] Optionally, the risk assessment algorithm of the weighted calculation is: ; Wherein, is a risk comprehensive index, is the number of operations of the risk assessment model, , , , is the operation result of each model in the first , , , , is the calculation weight of each model in the first operation.
[0009] Optionally, after constructing the risk assessment algorithm of the weighted calculation, it further includes: encapsulating the risk assessment algorithm into an algorithm tool, and integrating into an information system.
[0010] The application further discloses a multi-source data fusion risk assessment system, comprising: a data processing module, used for acquiring risk related data and constructing a risk data system; a data view construction module, used for integrating and associating different source system data in the risk data system to construct a complete data view; a data analysis module, used for performing data analysis and mining on the integrated and associated data, extracting data characteristics of subject events, to obtain development trends, demand analysis and potential correlations of a plurality of subject events; a model construction module, used for constructing a plurality of risk assessment models based on the data analysis results, and finding target behaviors and generating early warning information through the plurality of risk assessment models; and an index calculation module, used for constructing a risk assessment algorithm of weighted calculation based on outputs of the plurality of risk assessment models, to generate a risk comprehensive index and complete quantitative assessment of the risk.
[0011] The application further discloses a computer readable storage medium, the storage medium stores a computer program, and the computer program is executed by a processor to realize the risk assessment method.
[0012] The application further discloses an electronic device, comprising a memory, a processor and a computer program stored in the memory and executable on the processor, and the processor realizes the risk assessment method when executing the program.
[0013] According to the technical scheme of the present application, by constructing a risk data system, comprehensively integrating multi-system data, efficient association and unified management of data from different sources are realized, the problems of data dispersion and low utilization rate in the prior art are solved, a complete data view is formed, based on the risk data system, a plurality of risk assessment models suitable for different application scenarios are constructed, which can accurately find multiple scenarios, and the identification ability of potential risks and sensitive behaviors is significantly improved. Further, based on the output of the plurality of assessment models, a weighted calculation risk assessment algorithm is constructed, which can accurately generate a risk comprehensive index and complete quantitative assessment of the risk. BRIEF DESCRIPTION OF DRAWINGS
[0014] For the purpose of illustration and not limitation, the present application will now be described in conjunction with embodiments of the present application and the accompanying drawings, in which: Figure 1 is a flowchart of a risk assessment method in an embodiment of the present application; Figure 2 is a structural schematic diagram of a risk assessment system in an embodiment of the present application; Figure 3 is a structural schematic diagram of an electronic device in an embodiment of the present application. DETAILED DESCRIPTION
[0015] In order to enable persons skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by persons skilled in the art without creative labor should fall within the scope of protection of the present application.
[0016] Moreover, in addition to being used to represent the positional or locational relationship, the above-mentioned part of the terms can also be used to represent other meanings, for example, the term "upper" can also be used to represent a certain dependent relationship or connection relationship in some cases. For persons skilled in the art, the specific meanings of these terms in the present application can be understood according to the specific circumstances. In addition, the meaning of the term "a plurality of" should be two and more than two.
[0017] It should be noted that the features in the embodiments of the present application and the embodiments can be combined with each other without conflict. The embodiments of the present application will be described in detail below with reference to the accompanying drawings.
[0018] Reference Figure 1 The embodiments of the present application disclose a risk assessment method, which comprises: S100: Obtain risk-related data and build a risk data system. The risk-related data is used to reflect the information data generated by various aspects of the current society, which is composed of a risk data set defined as , wherein, , , are dimensionless. In this embodiment, is a full data list, which refers to a complete data set containing all relevant data items, records or entities, and is the key to ensuring the comprehensiveness and accuracy of model analysis. is a range data list, which refers to the business data of the industry supervisory department highly related to risk, and is a supplement to the full data list .
[0019] The range data list includes 10 data subsets and 25 data items. After obtaining the full data list and the range data list , the data characteristic values in the full data list are extracted, and the corresponding partial data in the range data list are extracted, and finally combined into a complete data set . The above data characteristic value extraction can take the following steps: 1. Data preprocessing: remove duplicates, missing values and outliers by data cleaning, unify the format of similar data from various channels such as date, numerical unit, etc., and use z-score standardization to reduce the impact of dimension. 2. Basic attribute extraction: extract the dynamic characteristics of data from the event dimension and spatial dimension, such as the time, place and number of participants of the event.
[0020] 3. Content feature extraction: use natural language processing (NLP) technology to extract keywords and theme content of the main event, and combine sentiment analysis, text classification, etc. to identify the tendency and key information of the data content, such as identifying events related to a specific theme.
[0021] The above extraction of corresponding partial data in the range data list can take the following steps: 1. Range data matching: Two range data matching methods are adopted, the first one is based on keyword or rule screening, according to the definition of each category in the range data list, the relevant keyword matching rule is designed for each data item. For example, for the "real estate development delivery" data item, "real estate development", "delivery" and "building" can be designed as keywords. The second one is based on NLP model classification, through the preprocessing of word segmentation, stop word removal and vectorization on the descriptive text such as complaint content and case description, the pre-trained BERT model is fine-tuned to train the classifier to classify the full data, and the unlabeled data is automatically allocated to the corresponding category. The F1 score is used to evaluate the classification effect in the classification process.
[0022] 2. Data integration: The extracted data is integrated to form a structured data set. This data set will include data from multiple source systems, and these data have been screened and integrated according to the evaluation requirements, and the clustering and association rule mining is used to reveal the potential relationship between the data.
[0023] S200: The data of different source systems in the risk data system is integrated and associated to construct a complete data view. By constructing the data view, the potential relationship and pattern between the data of different source systems can be revealed, which can better support the construction of risk assessment model. Specifically: S210: Data cleaning, data standardization and data matching are performed on different source system data; S220: Obtain the association key in different source system data, integrate and associate the different source system data with the same association key to form a data set representing the behavior, activity and event of the subject corresponding to the current association key in different systems; S230: The full data list is integrated and a unified data view is formed, the keywords in the data view are extracted as identifiers, and the different source system data are associated; S240: Data cleaning and unification are performed on the range data list, and key elements are extracted as retrieval elements to associate different source system data; S250: Based on the integrated and associated different source system data, a data view is constructed with three dimensions of matter, person and enterprise as the theme.
[0024] The embodiment exemplarily gives a data management method of different source systems. For data integration, first, data cleaning, standardization and data matching steps are performed on the data to avoid errors caused by inconsistent data formats or duplicate data, effectively improving data quality and reliability. Further, the association key is confirmed, and the data connection is completed, thereby forming a comprehensive data set. For the data in the full data list, the data in the full data list is integrated through time, place, type and other data dimensions to form a unified data view. Further, through the fields of certificate number, mobile phone number and content keyword as key identifiers, the data of different source systems is associated. Further, for the range data list association, the embodiment performs data governance and cleaning on the range data list based on the exemplary full data above, and extracts key elements as retrieval elements.
[0025] After the data integration and association in the exemplary full data list and range data list above, an analysis view with three dimensions of matters, persons and enterprises is exemplarily constructed. In the matter dimension, the occurrence frequency, trend and distribution of various events, problems and complaints can be analyzed. In the person dimension, based on information such as certificate number and mobile phone number, the characteristics, needs and problems of different groups of people are analyzed to support personalized services. In the person dimension, the behavior, preferences and characteristics of specific individuals or groups can be focused on. Through the association of the identity card number, the activities of a person in multiple systems, multiple levels and multiple regions can be tracked, and his behavior and social relationships in different fields can be understood. At the same time, the differences and similarities between different groups of people can be analyzed to provide targeted suggestions for policy making and public services. In the enterprise dimension, information such as key enterprises, consumer complaints and labor arbitration is focused on to analyze the operation status, risk points and market trends of enterprises. In the enterprise dimension, the operation, market performance and business status of enterprises can be focused on. Through the integration of enterprise-related data, the competitive position, market share and consumer evaluation of enterprises in different fields can be understood. At the same time, the impact of enterprises on society and the environment can be analyzed to evaluate the comprehensive value and sustainable development ability of enterprises. By constructing an analysis view with three dimensions of matters, persons and enterprises as the theme, a comprehensive and in-depth data insight can be obtained. It helps to discover the potential relationship and pattern between data, and reveals the rules and trends hidden behind the data.
[0026] S300: Perform data analysis and mining on the integrated and correlated data to extract data characteristics of the main events, thereby obtaining the development trends, demand analysis, and potential correlations of multiple main events. After obtaining the integrated and correlated data, the event development trajectory and key points of the main events can be obtained. Simultaneously, in-depth analysis and mining of the data through data analysis techniques can reveal potential event development trends, patterns, and correlations. This implementation method provides specific steps: S310: Perform data cleaning and standardization. Specifically, this involves denoising and deduplicating the data, structuring and standardizing the data, and imputing missing values. This step removes duplicate records, invalid data, and noise, and processes text data through word segmentation and stop word removal to unify the format. For numerical or categorical data, it ensures consistency in units, codes, and fields. Additionally, imputing missing values helps prevent errors in the analysis results.
[0027] S320: Extract the basic attributes of the main event and extract the dynamic features of the data from the event dimension and spatial dimension.
[0028] Basic attribute information is recorded when a major event occurs, such as event type, time, location, and participating themes. This basic attribute data is used for event clustering analysis. Specifically, event clustering analysis integrates multi-dimensional information such as time, region, channel, type, and scale to perform time-series clustering of events. This method can identify clustering patterns of events across different dimensions, such as which events occur frequently at specific times, locations, or channels. Clustering analysis reveals the inherent connections and distribution characteristics between events, providing a strong basis for predicting future event clustering trends and patterns. By collecting and analyzing this data, a comprehensive understanding of a company's employment situation, market operations, and potential risks can be obtained. This analysis helps companies formulate more precise market strategies, optimize resource allocation, and improve operational efficiency. Enterprise event correlation analysis extracts characteristic information of relevant companies from current hot events, more accurately linking events with companies. By understanding public perceptions and attitudes towards companies, brands, or events, government managers can better communicate with companies, and companies can adjust public relations strategies in a timely manner, strengthen brand image management, and respond to potential crises.
[0029] S330: Extract keywords and thematic content of the main event, and combine machine learning technology to identify data content trends and key information in order to obtain the content characteristics of the data.
[0030] S340: Based on the dynamic and content characteristics of the data, capture the development trend of the main event through statistical models.
[0031] In the dynamic development process of event situation, external factors such as policy changes, economic environment, population distribution and hot issues in surrounding cities are introduced to assess their impact on event development. Through multiple regression analysis or causal reasoning analysis, the relationship between these external factors and event development can be quantified, so as to more accurately predict the development trend of the event. This analysis is of great significance for formulating response strategies, adjusting resource allocation and making long-term planning. For example, risk prediction and evaluation analysis, by combining historical data and external information with statistical models, using risk assessment tools and system prediction analysis algorithms to predict the risks of specific events such as decision-making, policy-making, engineering planning, etc. By assessing the likelihood and impact of various risk factors, decision support can be provided for risk event prevention. This analysis helps governments and enterprises to discover potential risks in advance and take corresponding measures to prevent and respond.
[0032] S350: According to the content characteristics of the data, natural language processing and sentiment analysis techniques are used to identify the demand analysis of the subject event.
[0033] In event sentiment public opinion analysis, based on the extracted keywords and theme content of the subject event, natural language processing and sentiment analysis techniques are used to analyze the sentiment tendency of public comments and discussions on specific events or topics on social forums, government websites and other online systems. This method can reveal the public's views and attitudes towards related events or enterprises, providing a reference for developing public relations strategies. By understanding the public's sentiment tendency, we can better understand the voice of the public and better exercise the government's management functions. Through content feature analysis, the demands of the event subject can be deeply understood, and the portrait analysis of the subject's demands not only focuses on the basic demands of the subject, but also further analyzes the subject's emotional tendency, values and other factors. By building a comprehensive portrait of the subject, we can deeply understand the subject's needs and expectations, and provide support for developing personalized response strategies.
[0034] S360: According to the dynamic characteristics of the data, through graph network analysis techniques, the association network between subject events is constructed, and the key nodes and paths in the network are identified to obtain the deep association and mutual influence between multiple subject events, thereby obtaining the potential association of multiple subject events.
[0035] Through event fusion correlation analysis by graph network analysis technology, the correlation network between events is constructed, and the key nodes and paths in the network are identified. This analysis can reveal the deep correlation and mutual influence between events, and provide guidance for optimizing resource allocation and decision path. For example, in the field of public security, by analyzing the correlation between different events, potential risk points can be found to support the prevention of accidents and disasters. Similarly, for the correlation between event subjects, such as character relationships, character relationship correlation analysis uses social network analysis technology, combined with the time and regional characteristics of similar events, to analyze the potential relationship between characters. By correlating with civil affairs data, the event-related groups can be accurately identified, providing strong support for the prevention of mass incidents. This analysis is of great significance in the fields of public security, social stability and crisis management.
[0036] S400: Based on the data analysis result, a plurality of risk assessment models are constructed, and the target behavior is found and the early warning information is generated through the plurality of risk assessment models. Specifically: a plurality of preset scenes are identified, and data corresponding to the preset scenes is collected from the data analysis result; the collected data is feature extracted and index constructed; according to different preset scenes, corresponding risk assessment models are constructed; based on the plurality of risk assessment models, the output results of each risk assessment model are identified; the target behavior is found and the early warning information is generated through the output results.
[0037] The present embodiment gives an exemplary construction method and application of multiple risk models based on actual application. Since the assessment of risk involves multiple scenarios, after completing data analysis, risk analysis models are established according to the preset multiple scenarios, so that they can actively perform intelligent analysis according to specific application scenarios, find key sensitive behaviors such as hot issues, mass incidents, and overactive behaviors, and timely predict and prevent. Exemplarily, the analysis models of multiple preset scenes include key personnel risk analysis model, extreme event risk analysis model, over-level behavior risk analysis model, mass incident risk analysis model, hot event risk analysis model, sudden event risk analysis model, multi-cross event risk analysis model, same character correlation risk analysis model, same event correlation risk analysis model, major event risk assessment model, special event risk monitoring model, and key enterprise risk analysis model. Specifically: A risk analysis model for key personnel, which aims to identify and analyze key personnel in the city in terms of livelihood appeals. By weighting the corresponding indicators, using clustering analysis and anomaly detection algorithms, potential high-risk individuals are identified, and early warning and disposal recommendations are provided to the system. The construction method of the model includes: data collection; feature extraction and index construction; problem complexity assessment; supervision and management situation identification: identify cases that are supervised and managed by relevant departments as key governance, and individuals related to them; threat information identification: use natural language processing technology to identify the threat information published by individuals and assess their potential risks; sensitive problem involvement: analyze the sensitivity of the problems involved by individuals; key breakthrough and help and rescue status: identify whether the individual is a key breakthrough object or a help and rescue object; comparison of key personnel lists from other departments: compare the key personnel lists shared by other departments to identify cross-key personnel. Weighted assignment and model construction: indicator weighting: according to the importance and influence of each indicator, assign appropriate weights; clustering analysis: use clustering algorithms to divide individuals with similar characteristics into different groups for subsequent risk assessment and disposal; anomaly detection: identify abnormal individuals in the group through anomaly detection algorithms, which often have higher risks.
[0038] Extreme event risk analysis model: The extreme event risk analysis model is used to predict and identify mass events that may cause serious consequences. By collecting data on historical extreme events, such as the type of event, time of occurrence, location, number of participants, etc., a prediction model is established using time series analysis and machine learning algorithms to detect potential extreme event risks in advance and provide preventive measures for the system. The specific method is as follows: Algorithm construction process and method: Sample data: 2 million event data are obtained as sample data, and event details are used as raw data for data analysis; Feature selection: analyze the threatening behavior and reflection content of the sample data. Threatening behavior: mainly extract extreme behavior vocabulary; Reflection content: reflects the potential emotionalization of residents and the extreme behavior they may commit in the future, and digs out the key information implied, which can be used as an important indicator to measure the prediction; Data cleaning: there are a lot of invalid data in the original data, so the data needs to be preprocessed to obtain clean and effective training text data, so that it has the ability to effectively express information and does not have too much interference data. The data cleaning process not only prepares for the later algorithm establishment, but also understands the data content and meaning well in this process, and understands the relevant business knowledge; Data standardization: the cleaned sentence text is segmented by a segmentation tool, and the text is split into multiple words; The clean data is segmented by a segmentation tool to obtain a word list; The list is connected into a complete sentence by spaces, and the different data categories and basic probabilities are obtained according to the historical sample statistical analysis and the values are assigned to the data as labels to form a standardized data carrier; Algorithm rule: the extreme event risk analysis model obtains historical extreme behavior risk data, and the model trained by the original data statistics of the threatening and non-threatening probabilities of various events can narrow the range of extreme behavior. The keywords of extreme behavior are extracted and matched from the predicted work orders to screen the participants of the work orders who may have extreme behavior risks; Algorithm tuning and testing: new event information is obtained, special characters and garbage data in the content are removed, and word segmentation processing is performed. The processed data is transmitted to the algorithm for prediction, and the prediction output result is compared with the actual situation. According to the test result, the algorithm model is tuned, and the tuned model is continuously tested. This project will select 400,000 data for model testing until the extreme risk prediction model output accuracy reaches 70%. For the application of this model, the embodiment is exemplified as follows: A new incoming work order is given a probability of possible threatening based on the algorithm model. For those who may threaten, extreme vocabulary matching is performed, and for those who do not threaten, they are directly discarded. The final result is: the user may threaten, proving that the user may commit extreme behavior; The extreme behavior vocabulary contained in it is directly extracted to accurately locate the extreme behavior he may commit.The predicted extreme event work order content details are displayed, including event location, event details, date of occurrence, access type, event graph (for the connection between key elements in the event work order, showing people, things, places, objects, time, participating units, etc. in the event), predicted reasons, processing status, etc. This can enable work staff to handle work content at the work order level.
[0039] Overstep behavior risk analysis model: This model focuses on identifying and analyzing the overstep access behavior of potential risk personnel. By collecting information such as channel complaint records, appeal content, and handling methods of potential risk personnel, using graph analysis and classification algorithms, the model finds out the rules and characteristics of overstep access and provides disposal strategies for overstep behavior. The construction method of the model is as follows: algorithm construction process and method: obtain 2 million event data as sample data, and event details as raw data for data analysis; feature selection: analyze the problem attribution and reflection content of sample data. Problem attribution: as an upward consideration, it is considered that the ability and speed of different regions and different departments to handle problems may affect the residents' upward factors; reflection content: directly expresses the residents' appeal and reflects the residents' potential emotionalization and future possible extreme behavior, and digs out the key information that can be used as an important indicator for prediction; data cleaning: there are a large number of invalid data in the original data, so the data needs to be preprocessed to obtain clean and effective training text data, so that it has the ability to effectively express information and has no excessive interference data information carrier, which specifically includes removing special characters, garbage data, invalid text, empty data, etc.The data cleaning process not only prepares for later algorithm development but also facilitates a thorough understanding of the data content and its meaning, providing a solid grasp of relevant business knowledge. Data standardization involves segmenting the cleaned text into words using a word segmentation tool, creating a word list, and then connecting these words with spaces to form complete sentences. Historical sample statistical analysis is used to determine the categories and base probabilities of different accessed data, assigning values to these data entries as labels to form standardized data carriers. Algorithm rules: The risk analysis model for escalating behavior primarily categorizes historical work orders into multiple progressively higher levels. The number of times the same type of event is processed at different administrative levels is statistically analyzed to obtain the percentage of each category processed at each level. This percentage allows us to determine the probability of such events at each administrative level. For example, for category A events, the percentage processed at the first level (highest level) is 20%, at the second level it is 40%, and at the third level it is... The processing rate at the first level is 30%, and at the fourth level it's 10%. Therefore, for a new work order in category A, the probability of it going to the first-level department is 20%, to the second-level department is 40%, to the third-level department is 30%, and to the fourth-level department is 10%. After giving these probabilities, we consider eight influencing factors (labels for existing work orders) to adjust the original probabilities. For example, the probability of a category A event going to the first-level department might be 20% after classification, but if all eight influencing factors are unfavorable, the probability of going to the first-level department increases significantly. If only one influencing factor has a negative impact while the others are positive, the probability of going to the first-level department decreases. The probability with the highest probability is taken as the risk probability of the event going to the administrative level. The calculation process for the above probabilities is as follows: Input data processing: including work order text, historical statistical data, and eight influencing factors; Initial probability calculation: based on the processing proportion of different categories of work orders at each administrative level in historical data; Probability correction: adjusting the original probabilities through weighted average based on the eight influencing factors; Risk prediction: selecting the administrative level corresponding to the highest probability as the prediction result based on the adjusted probabilities. Specifically, the model code examples involved in the above calculations are as follows: import pandas as pd import numpy as np # 1. Initial probability calculation def calculate_initial_probabilities(history_data, event_category): """ Calculate initial probabilities based on historical data.
[0040] :param history_data: DataFrame containing historical ticket data :param event_category: Category of the current ticket :return: Dictionary of initial probabilities, keys are administrative levels, values are corresponding probabilities """ category_data = history_data[history_data['category'] == event_category] total_count = category_data['count'].sum() probabilities = category_data.groupby('level')['count'].sum() / total_count return probabilities.to_dict() # 2. Probability Adjustment def adjust_probabilities(initial_probs, factors): """ Adjust initial probabilities based on eight influencing factors.
[0041] :param initial_probs: Dictionary of initial probabilities :param factors: Dictionary of weights for eight influencing factors, keys are factors, values are influence weights (-1~1) :return: Dictionary of adjusted probabilities """ # Assume each factor has an impact on each level, here we do a simple weighted sum adjusted_probs = initial_probs.copy() total_factor_weight = sum(abs(weight) for weight in factors.values()) for level in adjusted_probs.keys(): adjusted_probs[level] = initial_probs[level] * (1 + factors[level] / total_factor_weight) if adjusted_probs[level] < 0: adjusted_probs[level] = 0 if adjusted_probs[level] > 1: adjusted_probs[level] = 1 return adjusted_probs Adjust probabilities for each level factor_adjustment = sum(factors[factor] * (1 if factor in level else -1) for factor in factors) adjusted_probs[level] += factor_adjustment / total_factor_weight # Ensure probabilities sum to 1 total_prob = sum(adjusted_probs.values()) for level in adjusted_probs: adjusted_probs[level] = max(0, adjusted_probs[level]) / total_prob # Ensure non-negative and normalize return adjusted_probs # 3. Predicting the risk level def predict_risk(adjusted_probs): """ Predict the risk level based on the adjusted probabilities.
[0042] :param adjusted_probs: Dictionary of adjusted probabilities :return: The administrative level with the highest probability and the probability of that level """ predicted_level = max(adjusted_probs, key=adjusted_probs.get) predicted_prob = adjusted_probs[predicted_level] return predicted_level, predicted_prob # Example data history_data = pd.DataFrame({ 'category': ['salary','salary','salary','salary', 'construction','construction','construction','construction'], 'level': ['national','province','city','district', 'national','province','city','district'], 'count': [20, 40, 30, 10, 10, 30, 50, 10] }) new_event = { 'category':'salary', 'factors': { 'initial_case': 0.2, 'group_visit': 0.1, 'legal_involvement': 0.3, 'threats': 0.4, 'case_status': -0.2, # Completion status 'evaluation': 0.1 # Evaluation index } } # Execute the prediction process initial_probs = calculate_initial_probabilities(history_data, new_event['category']) adjusted_probs = adjust_probabilities(initial_probs, new_event['factors']) predicted_level, predicted_prob = predict_risk(adjusted_probs) # Output the results print(f"Initial Probabilities: {initial_probs}") print(f"Adjusted Probabilities: {adjusted_probs}") print(f"Predicted Level: {predicted_level}, Probability: {predicted_prob:.2f}") # Output example Initial Probabilities: {'national': 0.2,'province': 0.4,'city': 0.3,'district': 0.1} Adjusted Probabilities: {'national': 0.25,'province': 0.35,'city':0.3,'district': 0.1} Predicted Level: province, Probability: 0.35.
[0043] Furthermore, algorithm tuning and testing: extract a single-layer neural network model, then process the document vectors through the hidden layer for multi-classification. Input the profile, remove special characters and garbage data from the content, and perform word segmentation processing. The processed data is transmitted to the algorithm for prediction, and the predicted output is compared with the actual situation. According to the test results, the algorithm model is tuned, and the tuned model is continuously tested.
[0044] Crowd event risk analysis model: This model is used to predict and identify potential crowd events that may trigger large-scale gatherings. By collecting and analyzing data from historical crowd events, such as the causes, development process, and number of participants, and using existing multi-channel appeal data, a person relationship network analysis and prediction model is established to discover potential crowd event risks and provide early warning and response measures for the system. The specific construction method of the model is as follows: Algorithm construction process and method: Obtain 2 million event data (de-sensitized data such as personnel information, responsible department information, and work order information) as sample data, and use event details as the original data for data analysis; feature selection: analyze the reflection content of the sample data, which directly expresses the residents' access appeal and reflects the residents' potential emotionalization and future possible radical behavior, and excavates the key information that can be used as an important indicator for prediction; data cleaning: there is a lot of invalid data in the original data, so the data needs to be preprocessed to obtain clean and effective training text data, so that it has the ability to effectively express information and has no excessive interference data information carrier, which specifically includes removing special characters, garbage data, invalid text, and empty data. The data cleaning process not only prepares for the later algorithm establishment, but also understands the data content and meaning well in this process, and understands the relevant business knowledge; data standardization: the cleaned sentence text is segmented by a segmentation tool, and the text is split into multiple words; the clean data is segmented by a segmentation tool to obtain a word list; the list is connected into a complete sentence by spaces, and the categories and basic probabilities of different access data are obtained according to historical sample statistical analysis and the values are assigned to the data as labels to form a standardized data carrier; algorithm rule: the crowd risk prediction analysis takes the work order reflection content as the main body, classifies and counts the historical work orders, and classifies the content through rules to form a crowd event prediction model, the classification rule of the crowd event is to classify according to the actual number of events, including 5, 20, 50, 100, 200, and 500 six levels, and after obtaining the probability, the actual number of events of this type of event in all channels is combined to predict the probability of the occurrence of the collective event; algorithm tuning and testing: according to the historical data, the event occurrence identification of the problem location, the age, education level, and frequency identification of the personnel, and the event classification identification are used as the prediction input vector to predict the probability of the new event developing into a crowd event, the prediction output result is compared with the actual situation, the algorithm model is tuned according to the test result, and the tuned model is continuously tested. This project will select 400,000 data to test the model until the accuracy rate of the crowd risk prediction model output result reaches 70%.
[0045] Hotspot event risk analysis model: This model aims to track and analyze current social hot issues in real time. By collecting real-time data from various channels, using text mining and topic modeling algorithms, the key information of hot events is extracted to provide public opinion analysis and coping strategies for the system. The specific construction method of the model is as follows: data collection and processing; data cleaning: remove noise, duplicates, invalid information, etc. in the data to ensure the accuracy and effectiveness of the data; data preprocessing: perform word segmentation, stop word removal, part-of-speech tagging, etc. on the cleaned data for subsequent text mining and topic modeling analysis; text mining and topic modeling: text mining: use natural language processing techniques to deeply mine the preprocessed text and extract key information such as event name, time, location, characters, opinions, etc.; topic modeling: use LDA (Latent Dirichlet Allocation) and other topic modeling algorithms to classify and cluster the text, identify the current social hot events and topics; hotspot event identification and analysis: hotspot event identification: according to the results of topic modeling, combined with the dynamic changes of real-time data, identify the current social widely concerned hot events; event analysis: in-depth analysis of the identified hot events, including event development trend, influence evaluation, social emotional tendency, regional coverage, daily occurrence quantity, etc. to provide comprehensive situation analysis for the system.
[0046] Emergency risk analysis model: the specific model construction method is as follows: algorithm construction process and method: obtain 2 million relevant event data as sample data, and take event details as original data for data analysis; feature selection: analyze the problem attribution and reflection content of sample data. Problem attribution: the reason for being considered as uplink is that the ability and speed of different regional departments to handle problems may affect the uplink of residents; reaction content: analyze the reflection content of sample data, which directly expresses the access appeal of residents, reflects the potential emotionalization and future possible radical behavior of residents, and excavates the key information that can be used as an important index for prediction; data cleaning: there are a large amount of invalid data in the original data, so the data needs to be preprocessed to obtain clean and effective training text data, so as to make it an information carrier that can effectively express information and has no excessive interference data, which specifically includes removing special characters, garbage data, invalid text, empty data and the like. The process of data cleaning not only prepares for the establishment of the later algorithm, but also can well understand the data content and meaning and understand the relevant business knowledge in this process; data standardization: the sentence text after cleaning is segmented by a segmentation tool, and the text is split into multiple words; the clean data is segmented by a segmentation tool to obtain a word list; the list is connected into a complete sentence by spaces, and the categories and basic probabilities of different access data are obtained according to historical sample statistical analysis and the values are assigned to the data as labels to form a standardized data carrier; algorithm rule: the emergency risk prediction analysis takes the work order reflection content as the main body, classifies and counts the historical work orders, and comprehensively calculates the indicators such as the reaction appeal quantity change curve (dispersion degree) of the event, the group involved in the appeal, the region covered by the appeal, and the domain involved in the appeal. Algorithm optimization and testing: according to the historical data, the event occurrence identification of problem attribution, the age, education, frequency identification and event classification identification of personnel are taken as the prediction input vector, so as to predict the probability of emergency event, compare the prediction output result with the actual situation, optimize the algorithm model according to the test result, continue to test the optimized model, and the project will select 400,000 data to test the model until the accuracy rate of the output result of the emergency risk prediction model reaches 70%.
[0047] Multi-cross event risk analysis model. By integrating data from different sources, using correlation analysis and network model algorithms, the correlations and mutual influences between events are found out, and cross-domain and cross-regional collaborative disposal suggestions are provided for the system. The specific construction method of the model is as follows: data collection and integration: multi-source data collection; data cleaning and preprocessing: remove noise, duplicates and invalid information in the data, clean and preprocess the data for subsequent analysis and modeling; data integration: integrate the cleaned and preprocessed data into a unified platform to form a multi-cross event data set. Correlation analysis and network model construction: event correlation analysis: use correlation analysis algorithms to identify potential connections between different events, including temporal correlation, spatial correlation, content correlation, etc.; network model construction: based on the results of event correlation analysis, construct an event network model, taking events as nodes and the correlations between events as edges to form a complex event network; key node and path identification: identify cross-domain and cross-regional key nodes and paths in the network model, which are often the key to event propagation and diffusion.
[0048] Same person correlation risk analysis model: This model aims to identify and analyze the characters who frequently appear in events in different channels. By collecting and analyzing information such as the records of these characters' appeals, using knowledge graph algorithms and link analysis techniques, potential connections of the same characters in different channels are found out, providing data support for character profiling and key personnel management. The specific construction method of this model is as follows: data collection and preprocessing: data collection: collect data related to characters from multiple sources, including but not limited to name, identity, appeal records, activity trajectory, social media interaction, etc.; data cleaning: clean the collected data to remove duplicates, noise and invalid information, ensuring the accuracy and consistency of the data; data standardization: standardize data from different sources for subsequent analysis and modeling. Character identification and correlation analysis: character identification: use natural language processing (NLP) techniques and text mining algorithms to accurately identify the characters involved in the text data and extract key information; character information integration: integrate character information from different sources to build a complete character information library; correlation analysis: based on knowledge graph algorithms and link analysis techniques, conduct correlation analysis on the identified characters to find out their potential connections and relationship networks. Character relationship network construction: network model construction: visualize the results of correlation analysis to construct a complex character relationship network graph; network feature extraction: extract key features such as centrality, density, clustering coefficient from the character relationship network for subsequent analysis and decision-making; character profiling: based on the constructed character relationship network and related data, conduct multi-dimensional profiling analysis of the characters, including personal characteristics, behavior patterns, social relationships, etc., and tag the character profiles with corresponding external data to improve the detail and accuracy of the character profiles.
[0049] Same event correlation risk analysis model: This model is used to analyze the same event of different channel data that are similar or related. By collecting and analyzing the data of these events, such as the type of event, the time of occurrence, the location, etc., using clustering analysis and similarity calculation algorithms, the correlation and similarity between events are found out, and then the events are combined, providing reference and reference for the system to handle events. The specific construction method of the model is as follows: Data collection and preprocessing: data collection; data cleaning: cleaning the collected data to remove duplicates, noise and invalid information; data standardization: standardizing the data from different sources to ensure the format and unit of event information are consistent. Event identification and feature extraction: event identification: use natural language processing (NLP) technology to identify events from text data and extract key information such as type, time, location, etc.; feature extraction: extract multiple features of events for subsequent event correlation analysis. Event correlation analysis: similarity calculation: use similarity calculation algorithms (such as cosine similarity, Jaccard similarity, etc.) to calculate the similarity between different events; clustering analysis: based on the similarity calculation results, use clustering algorithms (such as K-means, hierarchical clustering, etc.) to cluster similar events; correlation and combination: combine similar events after clustering to form a more complete and accurate event description.
[0050] Major event risk assessment model: This model is used to assess events involving significant interests or impacts. Import the risk assessment data of major political and legal events, analyze the data of these events, such as the impact range of the event, the number of people involved, and the degree of social concern, and use risk assessment and decision support algorithms to provide assessment reports and decision recommendations for the system. The specific construction method of the model is as follows: evaluation index: the evaluation model will analyze the event based on a series of evaluation indexes. These indexes include but are not limited to: the direct and indirect impact range of the event; the number of stakeholders and people involved; the urgency and long-term nature of the event; the compliance with laws and regulations; social public opinion and media attention. Risk assessment: by using risk assessment algorithms, the model will analyze the data and identify potential risk points. Risk assessment may involve quantitative analysis of the probability and potential impact of the event, resulting in a risk level. Decision support: based on the risk assessment results, the model will provide decision support suggestions.
[0051] Special event risk monitoring model: The special event monitoring model is used for multi-channel and multi-dimensional continuous monitoring and analysis of events related to a specific theme or field. By setting keywords or themes, real-time collection and analysis of relevant data is performed, and time series analysis and visualization techniques are used to show the development trend and changes of the event, providing real-time monitoring and early warning functions for the system. The modeling process is as follows: Set keywords or themes: According to user needs, set the keywords or themes that need to be monitored. Data collection: Collect data related to keywords or themes from multiple channels in real time. Data preprocessing: Clean, de-duplicate, classify and other preprocessing operations are performed on the collected data to ensure data quality. Data analysis: Use time series analysis, sentiment analysis and other algorithms to analyze the data and extract useful information. Visualization: Display the analysis results in the form of charts, images and other forms to help users understand the development trend and changes of the event. Early warning function: When the event reaches a certain condition, the system triggers an early warning according to the preset threshold or rules, reminding users to pay attention and handle it in time.
[0052] Key enterprise risk analysis model: Based on the data of key enterprises and non-governmental organizations in the city, the data between multiple systems is associated to form an enterprise portrait, including basic information and operating risk information of the enterprise. Monitor related risk areas such as complaints, enterprise registration changes, enterprise penalties, enterprise-related public opinion information, equity structure, social security payment, etc. When multiple channel complaints involving the enterprise are found, timely warning prompts are given, and specific risk details can also be searched for designated units. The specific construction method of the model is as follows: Data collection and integration: Data collection; Data integration: Clean, de-duplicate, standardize the collected data to form a unified data format, including enterprise basic information, operating risk information, public opinion information, etc. Enterprise portrait model construction: Basic information analysis: Sort out the basic information of the enterprise, including registration time, business scope, industry, equity structure, etc., to form the basic portrait of the enterprise; Operating risk analysis: Based on complaint records, penalty information, change records, etc., analyze the operating risk status of the enterprise, such as complaint frequency, penalty type, change reason, etc.; Public opinion analysis: Use natural language processing technology to analyze the sentiment and theme clustering of enterprise-related public opinion information to evaluate the social reputation and potential risks of the enterprise. Risk monitoring and early warning: Risk area setting: According to business needs, set the risk areas that need to be monitored, such as complaints, enterprise registration changes, enterprise penalties, enterprise-related public opinion, etc.; Risk score calculation: Based on the set risk areas and corresponding data indicators, calculate the risk score for each enterprise, with a higher score indicating a higher risk; Early warning: When the risk score of an enterprise exceeds the preset threshold, the system triggers an early warning and pushes the relevant information to the relevant regulatory departments.
[0053] In each of the above models, the model output layer can use the Softmax activation function, and the hidden layer can use the ReLU activation function. Taking the overstepping behavior risk analysis model as an example, the operation process is as follows: 1. Data collection and preprocessing: Collect 2 million event data, remove invalid data and clean up, including removing special characters, junk data, invalid text, etc. 2. Feature extraction and standardization: Extract features from the data, including question attribution, content reflection, and emotional expression information. Use a word segmentation tool to segment the text and obtain standardized text data. 3. Model construction and training: Use a neural network model for multi-classification processing. Input the preprocessed feature data, pass through the hidden layer (ReLU activation function), and output the probability of each class through the output layer (Softmax activation function). 4. Probability adjustment and correction: According to historical work order data statistics of different categories in different administrative levels, and combined with the influence factors, the probability is corrected to obtain the final overstepping risk prediction. 5. Prediction and risk assessment: For each new work order, calculate its overstepping risk based on historical data and provide prediction and disposal suggestions for relevant administrative departments.
[0054] The model code involved in the above operation is as follows: Data preprocessing and feature extraction import pandas as pd from sklearn.model_selection import train_test_split from sklearn.preprocessing import StandardScaler from sklearn.feature_extraction.text import CountVectorizer import numpy as np data = pd.read_csv('incident_data.csv') # Data cleaning, removing invalid text data_clean = data.dropna(subset=['incident_detail']) # Delete invalid data `data_clean['incident_detail'] = data_clean['incident_detail'].str.replace(r'[^a-zA-Z0-9\s]','', regex=True)` # Remove special characters #Word segmentation vectorizer = CountVectorizer(stop_words='english') X = vectorizer.fit_transform(data_clean['incident_detail']).toarray() #Word frequency matrix #Feature selection (location of the problem, content reflected, etc.) X_additional = data_clean[['region','complaint_type']] # Assuming other features are in these two columns X_final = np.hstack((X, X_additional)) y = data_clean['visit_type'] X_train, X_test, y_train, y_test = train_test_split(X_final, y, test_size=0.2, random_state=42) #Data Standardization scaler = StandardScaler() X_train = scaler.fit_transform(X_train) X_test = scaler.transform(X_test) Neural network model from tensorflow.keras.models import Sequential from tensorflow.keras.layers import Dense from tensorflow.keras.optimizers import Adam #Building a neural network model model = Sequential() model.add(Dense(128, input_dim=X_train.shape[1], activation='relu')) # Hidden layer, ReLU activation model.add(Dense(3, activation='softmax')) # Output layer, Softmax activation (3 categories) #Compilation Model model.compile(loss='sparse_categorical_crossentropy', optimizer=Adam(), metrics=['accuracy']) #Training the model history = model.fit(X_train, y_train, epochs=10, batch_size=32,validation_data=(X_test, y_test)) #Test Model loss, accuracy = model.evaluate(X_test, y_test) print(f'Model Accuracy: {accuracy*100:.2f}%') #Predicting New Data new_data = np.array([[...], [...], ...]) # New input data new_data_scaled = scaler.transform(new_data) # Standardize the new data predictions = model.predict(new_data_scaled) print(f'Predictions: {predictions}') Iteration function and model tuning def tune_model(X_train, y_train, X_test, y_test, epochs_range=[10,20], batch_size_range=[16, 64]): best_accuracy = 0 best_params = {'epochs': 10,'batch_size': 32} for epochs in epochs_range: for batch_size in batch_size_range: model = Sequential() model.add(Dense(128, input_dim=X_train.shape[1],activation='relu')) model.add(Dense(3, activation='softmax')) model.compile(loss='sparse_categorical_crossentropy',optimizer=Adam(), metrics=['accuracy']) history = model.fit(X_train, y_train, epochs=epochs,batch_size=batch_size, validation_data=(X_test, y_test)) accuracy = history.history['val_accuracy'][-1] if accuracy > best_accuracy: best_accuracy = accuracy best_params = {'epochs': epochs,'batch_size': batch_size} print(f"Best Params: {best_params} with Accuracy: {best_accuracy*100:.2f}%") return best_params, best_accuracy.
[0055] S500: Based on the outputs of multiple risk assessment models, a weighted risk assessment algorithm is constructed to generate a comprehensive risk index and complete the quantitative assessment of risk.
[0056] Specifically, the weighted risk assessment algorithm is as follows: ; wherein, is a risk comprehensive index, is the number of operations of the risk assessment model, , , , are the operation results of each model in the first , , , , are the calculation weights of each model in the first operation.
[0057] For the setting of the calculation weight, in the embodiments of the present application, firstly, the weight setting can be dynamically adjusted, and in the actual application process, it is continuously adjusted according to the dynamic changes of real-time data and event influence, and the independent modeling of different scenes ensures the reliability and independence of the input results, and the goal of the weighting process is to effectively integrate these results; secondly, the setting of the risk assessment algorithm of the initial weighting calculation is derived from the following two points: (1) Policy demand The policy demand mainly comes from the practical demand of social governance and the management focus of current social stability work. Specifically, it includes: clearly stipulating the hierarchical management and responsibility department handling mechanism of the access event, emphasizing the priority warning and key governance of major overstep events, especially the major access matters involving social stability, which must be handled and resolved in time to avoid the problem from expanding or upgrading to overstep events. According to the regulations, overstep behavior prediction and group event monitoring are important inputs of the algorithm model. It is required to evaluate major social risk events in advance and develop emergency plans. According to the law, the weight of major social contradictions and sudden group events is set as the core content.
[0058] (2) Social governance actual demand: overstep management demand: overstep events often involve sensitive issues and have a greater impact on social stability, so they need higher algorithm weight. Historical data statistics show that more than 70% of overstep events are concentrated in specific regions or industries, which provides data support for weight design. 2. Dynamic monitoring of key groups: for key personnel who have been visited many times, involve group events or have complex historical records, a higher evaluation weight needs to be set.
[0059] Finally, the calculation results are combined with actual experience to adjust the model weight in reverse for continuous optimization and iteration. Therefore, in general, how to set the weight can be summarized as follows: set the initial weight according to the policy and governance actual demand, and continuously improve and optimize it according to the actual work.
[0060] After constructing the risk assessment algorithm of the weighted calculation, the risk assessment algorithm is packaged into an algorithm tool and integrated into an information system.
[0061] Reference Figure 2 The embodiment further discloses a risk assessment system, comprising: A data processing module 21 is configured to acquire risk-related data and construct a risk data system.
[0062] A data view construction module 22 is configured to integrate and associate different source system data in the risk data system to construct a complete data view, and is further configured to perform data cleaning, data standardization and data matching on the different source system data, acquire an association key in the different source system data, integrate and associate different source system data with the same association key to form a data set representing behaviors, activities and events of a subject corresponding to the current association key in different systems, integrate the full-amount data list and form a unified data view, extract keywords in the data view as identifiers, associate different source system data, perform data cleaning and unification on the range data list, extract key elements as retrieval elements, and associate different source system data. Based on the integrated and associated different source system data, a data view with three dimensions of matters, persons and enterprises as themes is constructed.
[0063] A data analysis module 23 is configured to perform data analysis and mining on the integrated and associated data, extract data features of subject events, and obtain development trends, demand analysis and potential associations of multiple subject events. The data analysis module 23 is further configured to perform data cleaning and standardization on the data, specifically, denoising and deduplication of the data, structuring and standardization of the data, and filling processing of missing values, extract basic attributes of subject events, extract dynamic features of data from event dimensions and spatial dimensions, extract keywords and theme contents of subject events, and identify data content tendencies and key information in combination with machine learning technology to obtain content features of the data, capture development trends of subject events through a statistical model according to dynamic features and content features of the data, identify demand analysis of subject events according to content features of the data in combination with natural language processing and sentiment analysis technology, and construct an association network between subject events through graph network analysis technology according to dynamic features of the data, and identify key nodes and paths in the network to obtain deep associations and mutual influences between multiple subject events, thereby obtaining potential associations of the multiple subject events.
[0064] The model construction module 24 is configured to construct a plurality of risk assessment models based on the data analysis result, and find the target behavior and generate the early warning information through the plurality of risk assessment models; and is further configured to identify a plurality of preset scenes, and collect data corresponding to the preset scenes from the data analysis result; perform feature extraction and index construction on the collected data; construct a corresponding assessment model according to different preset scenes; identify the output result of each assessment model based on the plurality of assessment models; and find the target behavior and generate the early warning information through the output result.
[0065] The index calculation module 25 is configured to construct a weighted calculation risk assessment algorithm based on the output of the plurality of risk assessment models, to generate a risk comprehensive index, and complete quantitative assessment of the risk; and is further configured to encapsulate the risk assessment algorithm into an algorithm tool, and integrate the algorithm tool into an information system.
[0066] Figure 3 An electronic device entity structure schematic diagram provided by the embodiment of the present application is shown in FIG. 1. Figure 3 As shown in FIG. 1, the electronic device 50 includes a processor 501, a memory 502 and a bus 503. The processor 501 and the memory 502 can communicate with each other through the bus 503; the processor 501 is configured to invoke program instructions in the memory 502, to execute the method provided by the above-mentioned method embodiments.
[0067] The embodiment of the present application provides a non-transitory computer readable storage medium, which stores computer instructions, and the computer instructions make the computer execute the method provided by the above-mentioned method embodiments.
[0068] Those skilled in the art can understand that all or part of the steps of the above-mentioned method embodiments can be completed by program instruction related hardware, and the foregoing program can be stored in a computer readable storage medium, and the program is executed to perform the steps of the above-mentioned method embodiments; and the foregoing storage medium includes ROM, RAM, magnetic disc or optical disc and various storage media that can store program codes.
[0069] The device embodiments described above are only schematic, and the units shown as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0070] Those skilled in the art can clearly understand the implementation of the embodiments by the description of the above embodiments. The embodiments can be implemented by means of software and necessary universal hardware platforms, and of course, can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that makes a contribution, can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, etc., and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the method of each embodiment or some parts of the embodiment.
[0071] The above specific embodiments do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can occur depending on design requirements and other factors. Any modification, equivalent replacement, and improvement made within the spirit and principles of the present application shall fall within the protection scope of the present application.
Claims
1. A risk assessment method of multi-source data fusion, characterized in that, The method comprises: acquiring risk-related data and constructing a risk data system; integrating and correlating different source system data in the risk data system to construct a complete data view; performing data analysis and mining on the integrated and correlated data, extracting data features of subject events, to obtain development trends, demand analysis and potential correlations of multiple subject events; constructing multiple risk assessment models based on the data analysis results, and finding target behaviors and generating early warning information through the multiple risk assessment models; based on the output of the multiple risk assessment models, constructing a weighted risk assessment algorithm to generate a risk comprehensive index, completing quantitative assessment of the risk.
2. The risk assessment method of claim 1, wherein, The risk-related data includes a full data list and a range data list, the full data list contains complete data sets of all related data items, records or entities, and the range data list includes business data of risk-related industry authorities.
3. The risk assessment method of claim 2, wherein, The method of integrating and correlating different source system data in the risk data system to construct a complete data view comprises: data cleaning, data standardization and data matching of different source system data; acquiring association keys in different source system data, integrating and correlating different source system data with the same association keys to form a data set representing the behavior, activity and event of the subject corresponding to the current association key in different systems; integrating the full data list and forming a unified data view, extracting keywords in the data view as identifiers to associate different source system data; data cleaning and unification of the range data list, extracting key elements as retrieval elements to associate different source system data; based on the integrated and correlated different source system data, constructing a data view with three dimensions of matter, person and enterprise as the theme.
4. The risk assessment method of claim 1, wherein, The method of performing data analysis and mining on the integrated and correlated data, extracting data features of subject events, to obtain development trends, demand analysis and potential correlations of multiple subject events comprises: data cleaning and standardization, specifically, data denoising and deduplication, data structuring and standardization, and missing value filling; extracting basic attributes of subject events, extracting dynamic features of data from event and spatial dimensions; extracting keywords and theme content of subject events, and identifying data content tendency and key information by machine learning technology to obtain content features of data; capturing the development trend of subject events through statistical models according to the dynamic features and content features of data; identifying demand analysis of subject events according to the content features of data with natural language processing and sentiment analysis technology; according to the dynamic features of data, constructing the correlation network between subject events through graph network analysis technology, and identifying key nodes and paths in the network to obtain deep correlations and mutual influences between multiple subject events, thereby obtaining potential correlations of multiple subject events.
5. The risk assessment method of claim 1, wherein, The method of constructing multiple risk assessment models based on data analysis results, and finding target behaviors and generating early warning information through the multiple risk assessment models comprises: Identify a plurality of preset scenes, and collect data corresponding to the preset scenes from the data analysis results; Perform feature extraction and index construction on the collected data; Construct corresponding evaluation models according to different preset scenes; Identify the output results of each evaluation model based on the plurality of evaluation models; Find the target behavior through the output results and generate early warning information.
6. The risk assessment method of claim 1, wherein, The risk evaluation algorithm of the weighted calculation is: ; wherein, is a risk composite index, is the number of operations of the risk assessment model, , , , are the operation results of each model at the th operation, respectively, , , , are the calculation weights of each model at the th operation, respectively.
7. The risk assessment method of claim 1, wherein, After the risk evaluation algorithm of the weighted calculation is constructed, the method further comprises: Encapsulate the risk evaluation algorithm into an algorithm tool and integrate it into an information system.
8. A multi-source data fusion risk assessment system, characterized in that, It comprises: A data processing module for acquiring risk-related data and constructing a risk data system; A data view construction module for integrating and correlating different source system data in the risk data system to construct a complete data view; A data analysis module for performing data analysis and mining on the integrated and correlated data, extracting data features of subject events, and obtaining development trends, demand analysis and potential correlations of a plurality of subject events; A model construction module for constructing a plurality of risk evaluation models based on the data analysis results, and finding target behavior through the plurality of risk evaluation models and generating early warning information; An index calculation module for constructing a risk evaluation algorithm of weighted calculation based on the outputs of the plurality of risk evaluation models to generate a risk comprehensive index and complete quantitative evaluation of the risk.
9. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is executed by the processor to implement the method of any one of claims 1-7.
10. An electronic device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, The processor executes the program to implement the method of any one of claims 1-7.