Knowledge graph generation method and system for science and technology project risk control
The dynamic knowledge map is constructed through multi-source heterogeneous data conversion and risk quantitative evaluation model, which solves the problem of insufficient dynamic identification of risk control in scientific and technological projects in the existing technology, and realizes real-time and traceability of risk assessment.
Patent Information
- Application Number
- CN202510774422.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Existing risk control methods for scientific and technological projects are difficult to cope with the dynamic risk identification needs of complex projects, resulting in risk assessment lag and insufficient decision-making support capabilities.
By obtaining multi-source heterogeneous data, including structured indicator data, unstructured text data and timing behavior data, the heterogeneous data fusion mechanism is used to convert it into standard vector sequences and semantic vector sequences, input the risk quantitative evaluation model to generate risk entity feature matrix and association intensity matrix, construct a dynamic knowledge graph, and identify potential risk propagation paths.
It realizes the integrity of risk characteristics and complementarity of data characterization, improves the intelligence level of risk assessment and decision-making support efficiency, and can update in real time and maintain the traceability of multi-dimensional risk characteristics.
Smart Images

Figure CN120296180A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of knowledge graphs, and in particular, to a method and system for generating a knowledge graph for risk control of scientific and technological projects. Background Art
[0002] With the in-depth digital transformation, the risk control technology of scientific and technological projects has gradually become the core means to ensure scientific research safety and investment benefits. The current mainstream risk control methods for scientific and technological projects usually adopt structured data analysis models, construct risk assessment systems by extracting project indicators, or rely on expert experience to manually label risks for unstructured documents, and then generate risk analysis reports. However, traditional methods are difficult to meet the dynamic risk identification requirements of complex scientific and technological projects, and are prone to cause lag in risk assessment and insufficient decision-making support capabilities. Summary of the Invention
[0003] The present invention provides a method and system for generating a knowledge graph for risk control of scientific and technological projects.
[0004] In a first aspect, an embodiment of the present invention provides a method for generating a knowledge graph for risk control of scientific and technological projects, the method comprising: obtaining a multi-source heterogeneous data set of a target scientific and technological project, the multi-source heterogeneous data set including structured index data, unstructured text data, and time-series behavior data; through a heterogeneous data fusion mechanism, converting the structured index data into a standard vector sequence, converting the unstructured text data into a semantic vector sequence, and converting the time-series behavior data into a behavior pattern vector sequence; inputting the standard vector sequence, the semantic vector sequence, and the behavior pattern vector sequence into a risk quantification assessment model to generate a risk entity feature matrix and a risk association intensity matrix; determining the node distribution topology of the knowledge graph according to the entity feature vectors in the risk entity feature matrix, and determining the entity relationship topology of the knowledge graph according to the association intensity values in the risk association intensity matrix; generating a dynamic knowledge graph of the target scientific and technological project based on the node distribution topology and the entity relationship topology, and identifying potential risk propagation paths in the dynamic knowledge graph.
[0005] In a second aspect, an embodiment of the present invention provides a computer system, comprising: a memory, in which a computer program is stored; a processor, configured to load the computer program to implement the method for generating a knowledge graph for risk control of scientific and technological projects as described above.
[0006] The knowledge graph generation method based on the risk control of scientific and technological projects provided by the present invention unifies structured indicators, unstructured texts, and time-series behavior data into computable standardized vector sequences through a collaborative fusion mechanism of multi-source heterogeneous data, effectively eliminating the information island effect caused by the analysis of a single data source in traditional methods, and significantly improving the integrity of risk characteristics and the complementarity of data representation; uses a risk quantification evaluation model to jointly model the multi-modal vector sequences, and synchronously generates a risk entity feature matrix and a risk association strength matrix through a cross-attention mechanism and bidirectional interaction calculation, realizing the collaborative optimization of entity attributes and association relationships, and overcoming the feature matching errors caused by step-by-step modeling; based on the dynamic knowledge graph generation mechanism, performs spatio-temporal alignment and joint encoding on the node distribution topology and the entity relationship topology, and constructs a dynamically interpretable dynamic graph structure through node cluster optimization driven by density gradient and edge connection pruning by adaptive threshold segmentation, enabling the risk conduction path analysis to accurately capture the propagation laws of potential risks in the time and space dimensions; through the design of the present invention, realizes the deep coupling of knowledge graph generation and risk conduction analysis, and uses the dynamic edge weight allocation and path integrity verification mechanism to make the risk identification results have both the dynamic characteristics of real-time update and the traceability of multi-dimensional risk characteristics, comprehensively improving the intelligent level and decision-making support efficiency of scientific and technological project risk assessment. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 is a flowchart of a knowledge graph generation method for risk control of scientific and technological projects provided by an embodiment of the present invention; Figure 2 is a schematic diagram of the composition of a computer system provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0008] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present invention.
[0009] Please refer to Figure 1 , Figure 1 is a flowchart of a knowledge graph generation method for risk control of scientific and technological projects provided by an embodiment of the present invention. The knowledge graph generation method for risk control of scientific and technological projects can be executed by a computer system, and may include the following steps: Step S100: Obtain a multi-source heterogeneous data set of a target scientific and technological project. The multi-source heterogeneous data set includes structured indicator data, unstructured text data, and time-series behavior data.
[0010] In the embodiments of the present invention, the target scientific and technological project is a specific scientific and technological project that requires risk control and the construction of a knowledge graph. The multi-source heterogeneous data set indicates that the data comes from different sources and has different structures and formats. Structured indicator data refers to data with a clear structure and format, which can be represented in tabular form, such as the budget data and progress data of a project. These data can be organized and stored according to preset rules, facilitating statistics and analysis. Unstructured text data refers to text information without a fixed structure, such as project documents, reports, meeting records, etc. These text data may contain various descriptions and analyses of project risks, but due to their inconsistent formats, special processing is required to extract useful information. Temporal behavior data is behavior records related to time, such as the operation behaviors of project team members and the running states of systems. These data are arranged in chronological order, reflecting the behavior characteristics of the project at different time points.
[0011] Step S200: Through the heterogeneous data fusion mechanism, convert the structured indicator data into a standard vector sequence, convert the unstructured text data into a semantic vector sequence, and convert the temporal behavior data into a behavior pattern vector sequence.
[0012] The heterogeneous data fusion mechanism is used to integrate and convert data with different structures and formats. The purpose is to convert multi-source heterogeneous data into a unified vector representation form for subsequent analysis and processing. The standard vector sequence is a vector representation obtained after converting the structured indicator data, with a unified dimension and measurement standard, facilitating comparison and analysis in subsequent calculations. The semantic vector sequence is a vector obtained by processing the unstructured text data, which can reflect the semantic information of the text and helps to mine the risk factors contained in the text. The behavior pattern vector sequence is converted from the temporal behavior data, reflecting the behavior pattern characteristics of the project.
[0013] As an implementation manner, in step S200, through the heterogeneous data fusion mechanism, convert the structured indicator data into a standard vector sequence, convert the unstructured text data into a semantic vector sequence, and convert the temporal behavior data into a behavior pattern vector sequence. Specifically, it may include the following steps: Step S210: Perform abnormal state detection on the structured indicator data, generate a data correction mask matrix according to the detection results, and use the data correction mask matrix to perform interpolation reconstruction on the original indicator data, and output the corrected structured indicator data.
[0014] Anomaly state detection is the process of identifying and marking outliers in structured metric data. Outliers may be caused by data entry errors, system failures, etc. These outliers will affect the subsequent data analysis and processing results, so detection and correction are required. The data correction mask matrix is a binary matrix used to mark which data points in the original metric data are abnormal and which are normal. Interpolation reconstruction is the process of filling and correcting abnormal data points based on the information of normal data points.
[0015] When performing anomaly state detection, a statistical model-based method can be adopted. Specifically, first obtain the historical metric data of similar scientific and technological projects, calculate the mean vector and covariance matrix of each metric dimension, and generate a multi-dimensional normal distribution model based on the mean vector and covariance matrix. The historical metric data of similar scientific and technological projects can be obtained from the project database. By performing statistical analysis on these data, the mean and covariance of each metric dimension are obtained. The multi-dimensional normal distribution model is an existing probability distribution model that can describe the distribution of data in multiple dimensions. Then, input the structured metric data of the current project into the multi-dimensional normal distribution model and calculate the probability density value of each data point. When the probability density value is lower than the preset quantile of the distribution model, it indicates that this data point may be an outlier, and an anomaly marking sequence is generated. The preset quantile can be set according to the actual situation. For example, it can be set to 0.05, that is, when the probability density value of a data point is lower than 0.05, it is considered that this data point is an outlier.
[0016] According to the position information of the outliers in the anomaly marking sequence, a binary mask matrix is generated. In the binary mask matrix, the corresponding positions of the outliers take the value of zero, and the corresponding positions of the normal points take the value of one. In this way, the outliers in the original metric data can be clearly marked through the binary mask matrix. To eliminate isolated anomaly marking points, a morphological closing operation is performed on the binary mask matrix. The morphological closing operation can connect adjacent outliers to form a connected region. Through the closing operation, some accidentally occurring outliers can be prevented from being misjudged as real outliers. After generating the connected region correction mask matrix, based on the abnormal region marked by this matrix, locate the abnormal data points in the original metric data and extract the metric values of the adjacent data points within the time window before and after the abnormal data points. The size of the time window can be set according to the actual situation. For example, it can be set to 3 data points before and after each.
[0017] Input the metric values of adjacent data points into a moving average filter to generate an interpolation filling sequence. The moving average filter can smooth the input data and generate new values by calculating the average of adjacent data points. Modify the zero-value positions of the mask matrix according to the connected regions, replace the original abnormal data points with the interpolation filling sequence, and finally output the replaced structured metric data as the corrected structured metric data. In this way, the outliers in the structured metric data can be effectively corrected, improving the quality and reliability of the data.
[0018] As an implementation, in step S210, perform an abnormal state detection on the structured metric data, generate a data correction mask matrix according to the detection result, perform interpolation reconstruction on the original metric data using the data correction mask matrix, and output the corrected structured metric data. Specifically, it may include the following steps: Step S211: Obtain the historical metric data of similar scientific and technological projects, calculate the mean vector and covariance matrix of each metric dimension, and generate a multi-dimensional normal distribution model based on the mean vector and covariance matrix.
[0019] The historical metric data of similar scientific and technological projects refers to the metric data recorded by other scientific and technological projects with similar properties, scales, and business fields as the current target scientific and technological project over a period of time in the past. The mean vector is a vector composed of the averages of all data points in each metric dimension. For a data set with n metric dimensions, the mean vector can be represented as an n-dimensional vector, where each element is the average of all data points in the corresponding metric dimension. The method for calculating the mean vector is to add up all the data points in each metric dimension and then divide by the total number of data points. The covariance matrix is an n×n matrix used to describe the correlation between each metric dimension. The elements of the covariance matrix represent the covariance between the i-th metric dimension and the j-th metric dimension. The calculation formula for covariance is: ; where, and represent the values of the k-th data point in the i-th and j-th metric dimensions respectively, and are the means of the i-th and j-th metric dimensions respectively.
[0020] The multi-dimensional normal distribution model can be used to describe the joint distribution of multiple variables. In the embodiments of the present invention, a multi-dimensional normal distribution model is generated using the calculated mean vector and covariance matrix. The probability density function of the multi-dimensional normal distribution can be expressed as: ; where x is an n-dimensional vector representing the value of the data point; is the mean vector; is the covariance matrix; is the determinant of the covariance matrix; It is the inverse matrix of the covariance matrix. Through this multi-dimensional normal distribution model, the probability density of the structured index data of the current project can be calculated to determine whether the data points are abnormal.
[0021] Step S212: Input the structured index data of the current project into the multi-dimensional normal distribution model, calculate the probability density value of each data point, and generate an abnormal marking sequence when the probability density value is lower than the preset quantile of the distribution model.
[0022] The structured index data of the current project refers to the index data with clear structure and format collected in the target scientific and technological project, such as the progress data and cost data of the project. Input these data into the previously generated multi-dimensional normal distribution model, and calculate the probability density value of each data point according to the probability density function of the multi-dimensional normal distribution. The probability density value represents the likelihood of the data point appearing in the multi-dimensional normal distribution. For an n-dimensional data point x, its probability density value can be calculated by substituting x into the probability density function f(x) of the multi-dimensional normal distribution. The preset quantile of the distribution model is a preset threshold for determining whether the data point is abnormal. The quantile is the numerical point that divides the probability distribution into several equal parts.
[0023] When the calculated probability density value of the data point is lower than the preset quantile, mark the data point as an abnormal point. By traversing all the data points, an abnormal marking sequence is generated. The abnormal marking sequence is a sequence with the same length as the structured index data, where each element corresponds to a data point. If the data point is an abnormal point, it is marked as 1, otherwise it is marked as 0.
[0024] Step S213: Generate a binary mask matrix according to the position information of the abnormal points in the abnormal marking sequence. In the binary mask matrix, the positions corresponding to the abnormal points take zero values, and the positions corresponding to the normal points take one values.
[0025] The abnormal marking sequence records the information of whether each data point in the structured index data is an abnormal point. According to the position information of the abnormal points in this sequence, a binary mask matrix can be generated. The binary mask matrix is a matrix with the same dimension as the structured index data, and its function is to mark the abnormal and normal states of the data points through 0 and 1.
[0026] Suppose the structured index data is an m×n matrix, where m represents the number of data points and n represents the number of index dimensions. The abnormal marking sequence is a one-dimensional sequence with a length of m. For each element in the abnormal marking sequence, if its value is 1, it means the data point is an abnormal point, and the corresponding positions in the row and column of the binary mask matrix take zero values; if its value is 0, it means the data point is a normal point, and the corresponding positions in the row and column of the binary mask matrix take one values.
[0027] Step S214: Perform morphological closing operation on the binary mask matrix to eliminate isolated abnormal marked points and generate a connected region corrected mask matrix.
[0028] In the embodiments of the present invention, the binary mask matrix is regarded as a binary image, where 0 represents abnormal points (black pixels) and 1 represents normal points (white pixels). The process of morphological closing operation includes two basic operations: dilation and erosion. The dilation operation expands the white area (normal points) in the binary mask matrix outward so that adjacent white areas are connected; the erosion operation shrinks the white area inward to remove some small isolated white areas. By first performing the dilation operation and then the erosion operation, the isolated abnormal marked points in the binary mask matrix can be eliminated, and adjacent abnormal points can be connected into connected regions.
[0029] Specifically, for the dilation operation, a structuring element (such as a 3×3 square) can be defined, and the structuring element is slid on the binary mask matrix. When the structuring element completely overlaps a certain position in the matrix, the pixel value at that position is set to 1. For the erosion operation, when the structuring element is completely contained within the white area, the pixel value at that position is set to 1.
[0030] Through the morphological closing operation, the original isolated abnormal marked points are eliminated, and a connected region corrected mask matrix is generated. This matrix can more accurately reflect the true abnormal regions in the structured index data, providing more reliable information for subsequent interpolation and reconstruction.
[0031] Step S215: Based on the abnormal regions marked by the connected region corrected mask matrix, locate the abnormal data points in the original index data and extract the index values of the adjacent data points within the time windows before and after the abnormal data points.
[0032] The connected region corrected mask matrix marks the position information of the abnormal regions in the structured index data. According to this matrix, the abnormal data points can be accurately located in the original index data. The original index data is the structured index data collected before any processing. The time window is a preset range used to determine the number of adjacent data points before and after the abnormal data point. For example, the time window can be set to 3 data points before and after. For each abnormal data point, the index values of the adjacent data points within its time window before and after are extracted.
[0033] Suppose the original indicator data is a time - series data, containing indicator values at multiple time points. For an abnormal data point at the $i$-th time point, according to the setting of the time window, extract the indicator values of the adjacent data points from the $(i - 3)$-th to the $(i - 1)$-th time points and from the $(i + 1)$-th to the $(i + 3)$-th time points. By extracting the indicator values of the adjacent data points within the time window before and after the abnormal data point, the information of these normal data points can be used to interpolate and reconstruct the abnormal data point, thus correcting the anomaly.
[0034] Step S216: Input the indicator values of the adjacent data points into a moving average filter to generate an interpolation filling sequence. According to the connected region, correct the zero - value positions of the mask matrix, and replace the original abnormal data point with the interpolation filling sequence.
[0035] The moving average filter is used to smooth the time - series data. The basic principle is to calculate the average value of the data within a fixed - length window and use this average value as the new value at the center position of the window. Input the indicator values of the adjacent data points extracted above into the moving average filter. When calculating the interpolation filling sequence, start from the first element of the sequence and calculate the average value of the data within each window in turn. The zero - value positions of the connected - region - corrected mask matrix correspond to the abnormal data points in the original indicator data. According to these zero - value positions, replace the values in the interpolation filling sequence with the original abnormal data points. In this way, the abnormal data points are corrected using the information of adjacent normal data points, improving the quality of the structured indicator data.
[0036] Step S217: Output the replaced structured indicator data as the corrected structured indicator data.
[0037] After completing the replacement operation of the abnormal data points in the original indicator data, output the replaced structured indicator data as the corrected structured indicator data. The corrected structured indicator data removes the outliers and can more accurately reflect the actual situation of the target scientific and technological project, providing a reliable data basis for subsequent analysis and processing.
[0038] Step S220: Input the corrected structured indicator data into an indicator encoder, and map the indicator values with different dimensions to a unified metric space through a feature - dimension alignment algorithm to generate a standard vector sequence.
[0039] The index encoder is used to encode structured index data, aiming to convert the corrected structured index data into a vector representation form that is convenient for processing and analysis. The feature dimension alignment algorithm is a method for solving the problem of inconsistent dimensional quantities of different index values. It can map the index values of different dimensions into a unified metric space, making the various indexes comparable. The index values of different dimensions mean that in the structured index data, the measurement units and scales adopted by each index are different. The feature dimension alignment algorithm can be implemented in various ways, such as the standardization method. Standardization is to convert the value of each index into a value with a mean of 0 and a standard deviation of 1. Specifically, for each index dimension, first calculate its mean and standard deviation, and then standardize each data point. After feature dimension alignment, the transformed values of each data point on each index dimension are combined to form a vector. In this way, the corrected structured index data is converted into a standard vector sequence, where each vector corresponds to a data point, and the dimension of the vector is equal to the number of index dimensions.
[0040] Step S230: Perform semantic segment cutting on the unstructured text data, identify the text paragraphs containing risk factor descriptions, input the text paragraphs into the pre-trained semantic encoder to generate primary semantic vectors, and perform spatial transformation on the primary semantic vectors through the context-aware network to output a sequence of semantic vectors.
[0041] Semantic segment cutting is to divide the unstructured text data according to semantics and structure, and split it into multiple meaningful segments. The unstructured text data consists of a large amount of text without a fixed format and structure, such as project documents, reports, etc. Through semantic segment cutting, these text data can be decomposed into smaller and more easily processed units.
[0042] The sentence segmentation algorithm and paragraph division algorithm in natural language processing can be used for semantic segment cutting. For example, based on punctuation marks and grammar rules, the text is segmented into sentences, and then according to the semantic relevance between sentences, the relevant sentences are combined into paragraphs.
[0043] Identifying the text paragraphs containing risk factor descriptions is to screen out those paragraphs containing information related to project risks from the segmented text paragraphs. This can be achieved through methods such as keyword matching and semantic analysis. For example, some keywords related to risks are predefined, such as "risk", "crisis", "uncertainty", etc., and then the occurrence of these keywords is searched in each text paragraph. If a certain paragraph contains these keywords and its context semantics are related to project risks, then it is considered that the paragraph contains risk factor descriptions.
[0044] A pre-trained semantic encoder is a model that has been pre-trained on large-scale text data and can convert an input text passage into a vector representation. For example, the BERT (Bidirectional Encoder Representations from Transformers) model is a pre-trained semantic encoder. Inputting a text passage containing risk factor descriptions into the pre-trained semantic encoder, the encoder extracts and transforms the features of the text to generate a primary semantic vector. The primary semantic vector is a vector representation of the text passage and can reflect the semantic information of the text.
[0045] A context-aware network is a neural network model that can consider text context information and can perform a spatial transformation on the primary semantic vector to further mine the semantic information and context relationships in the text. For example, a general LSTM (Long Short-Term Memory) network or GRU (Gated Recurrent Unit) network can be used as a context-aware network. Inputting the primary semantic vector into the context-aware network, the network adjusts and transforms the primary semantic vector according to the context information of the text and outputs a sequence of semantic vectors.
[0046] Step S240: Perform event window partitioning on the time-series behavior data, extract the type distribution features and time interval features of the behavior operations within the window, and perform an orthogonal projection transformation on the type distribution features and the time interval features to generate a sequence of behavior pattern vectors.
[0047] Event window partitioning divides the time-series behavior data into multiple windows of a fixed length in chronological order. The time-series behavior data is time-related behavior records, such as the operation behaviors of project team members, the running status of the system, etc. Through event window partitioning, continuous time-series data can be segmented into multiple discrete windows, facilitating the analysis of the behavior data within each window.
[0048] The time interval feature is the time interval information between adjacent behavior operations within an event window. Statistical quantities such as the average value and standard deviation of the time intervals between adjacent behavior operations can be calculated as the time interval features.
[0049] In a non-essential optional embodiment, the orthogonal projection transformation is a method of fusing and transforming the type distribution features and the time interval features. It can project these two different types of features into a new vector space to generate a sequence of behavior pattern vectors. The set of basis vectors for the orthogonal projection transformation is determined by the singular value decomposition result of the covariance matrix of the standard vector sequence and the semantic vector sequence.
[0050] Specifically, first calculate the covariance matrix of the standard vector sequence and the semantic vector sequence, and then perform singular value decomposition on this covariance matrix to obtain a set of basis vectors. Project the feature vector composed of the type distribution feature and the time interval feature onto these basis vectors to obtain the behavior pattern vector sequence.
[0051] Step S300: Input the standard vector sequence, the semantic vector sequence, and the behavior pattern vector sequence into the risk quantification and assessment model to generate a risk entity feature matrix and a risk association strength matrix.
[0052] The risk quantification and assessment model is a model used to quantitatively analyze and evaluate the risks of scientific and technological projects. It can comprehensively consider the information contained in the standard vector sequence, the semantic vector sequence, and the behavior pattern vector sequence, and mine the risk entities and risk association relationships in the project.
[0053] The standard vector sequence is a vector sequence obtained by encoding and aligning the feature dimensions of the corrected structured index data, which reflects the structured index characteristics of the project. The semantic vector sequence is a vector sequence obtained by processing the unstructured text data and contains the risk semantic information in the text. The behavior pattern vector sequence is a vector sequence obtained by processing and transforming the time-series behavior data, which reflects the behavior pattern characteristics of the project.
[0054] After inputting these three vector sequences into the risk quantification and assessment model, the model analyzes and processes them to generate a risk entity feature matrix and a risk association strength matrix. The risk entity feature matrix is a matrix where each row corresponds to a risk entity, each column corresponds to a feature dimension, and the matrix element represents the value of the risk entity on that feature dimension. A risk entity can be a certain task, a certain resource, a certain stage, etc. in the project. The risk association strength matrix is a matrix used to represent the association strength between different risk entities, and the matrix element represents the degree of association between two risk entities.
[0055] For example, assume that the risk quantification and assessment model is a neural network-based model, which includes an input layer, a hidden layer, and an output layer. Concatenate the standard vector sequence, the semantic vector sequence, and the behavior pattern vector sequence together as the input to the input layer. After feature extraction and transformation by the hidden layer, the output layer outputs the risk entity feature matrix and the risk association strength matrix respectively.
[0056] As an implementation, in step S300, input the standard vector sequence, semantic vector sequence, and behavior pattern vector sequence into the risk quantification and assessment model to generate a risk entity feature matrix and a risk association strength matrix, which can specifically include the following steps: Step S310: Perform cross-attention calculation on the standard vector sequence and the semantic vector sequence to generate a cross-attention weight matrix of metrics and semantics. Each element of the cross-attention weight matrix reflects the association strength between the metric vector and the semantic vector.
[0057] The standard vector sequence is a vector sequence converted from structured metric data, and the semantic vector sequence is a vector sequence converted from unstructured text data. The cross-attention weight matrix is a two-dimensional matrix, the number of rows of which is equal to the number of vectors in the standard vector sequence, and the number of columns is equal to the number of vectors in the semantic vector sequence. Each element in the matrix represents the association strength between a metric vector and a semantic vector.
[0058] The process of cross-attention calculation can be achieved through the following steps: First, define queries (Query), keys (Key), and values (Value) for the standard vector sequence and the semantic vector sequence respectively. For example, the standard vector sequence can be used as the query, and the semantic vector sequence can be used as the key and value. Then, calculate the similarity between the query vector and the key vector. The similarity calculation method can use the dot product operation. Next, perform normalization processing on the calculated similarity values, such as the softmax function. The role of the softmax function is to convert the similarity values into a probability distribution so that the sum of all elements is 1. For each query vector q i , calculate the softmax value of its similarity with all key vectors: ; where n is the number of vectors in the semantic vector sequence. attn ij is the element in the i-th row and j-th column of the cross-attention weight matrix, which reflects the association strength between the metric vector q i and the semantic vector k j .
[0059] Step S320: Based on the cross-attention weight matrix, perform weighted aggregation on the semantic vector sequence to generate a semantically enhanced metric feature sequence, and multiply the semantically enhanced metric feature sequence element-wise with the standard vector sequence to generate a cross-fusion feature tensor.
[0060] Performing weighted aggregation on the semantic vector sequence based on the cross-attention weight matrix means performing weighted summation on each vector in the semantic vector sequence according to the weight values in the cross-attention weight matrix. Each element in the cross-attention weight matrix represents the association strength between the metric vector and the semantic vector. Through weighted aggregation, the information in the semantic vector sequence can be fused into the metric vector to generate a semantically enhanced metric feature sequence.
[0061] Multiply the calculated semantically enhanced metric feature sequence element by element with the standard vector sequence to generate a cross-fusion feature tensor. Element-by-element multiplication means multiplying each element of each vector in the semantically enhanced metric feature sequence by the corresponding vector in the standard vector sequence at the corresponding position.
[0062] Step S330: Expand the cross-fusion feature tensor into a two-dimensional matrix along the feature dimension, and perform a time alignment operation with the behavior pattern vector sequence. Adjust the row and column distribution of the two-dimensional matrix according to the timestamp information in the behavior pattern vector sequence to generate a time-aligned feature matrix.
[0063] Expanding the cross-fusion feature tensor into a two-dimensional matrix along the feature dimension means converting the three-dimensional cross-fusion feature tensor into a two-dimensional matrix. The cross-fusion feature tensor has three dimensions, for example: sample dimension, time dimension, and feature dimension. By expanding along the feature dimension, the feature information in the tensor can be flattened to obtain a two-dimensional matrix.
[0064] The time alignment operation means matching and aligning the expanded two-dimensional matrix with the behavior pattern vector sequence in time. The behavior pattern vector sequence contains timestamp information. According to this timestamp information, the row and column distribution of the two-dimensional matrix is adjusted so that the data in the two-dimensional matrix corresponds to the data in the behavior pattern vector sequence in time. After the time alignment operation, a time-aligned feature matrix is generated. The data in this matrix is consistent with the behavior pattern vector sequence in time, facilitating subsequent further analysis and processing.
[0065] Step S340: Perform an outer product operation on the time-aligned feature matrix and the behavior pattern vector sequence to generate an initial interaction matrix, and perform channel attention weighting on the initial interaction matrix to calculate the contribution degree of each channel dimension to the risk association, generating a channel weight vector.
[0066] The outer product operation is used to calculate the interaction information between two matrices. In the embodiments of the present invention, an outer product operation is performed on the time-aligned feature matrix and the behavior pattern vector sequence to obtain an initial interaction matrix. The initial interaction matrix reflects the interaction relationship between the time-aligned feature matrix and the behavior pattern vector sequence. Channel attention weighting is a method for calculating the contribution degree of each channel dimension to the risk association. The channel dimension is the third dimension of the initial interaction matrix. By channel attention weighting, a weight can be assigned to each channel dimension to reflect the importance degree of this channel dimension to the risk association.
[0067] Specifically, first, perform global average pooling operation on each channel dimension of the initial interaction matrix to obtain the average eigenvalue of each channel. Then, input these average eigenvalues into a fully connected layer, and after being processed by an activation function (such as the Sigmoid function), obtain the channel weight vector. The length of the channel weight vector is equal to the number of channel dimensions of the initial interaction matrix, and each element represents the weight of the corresponding channel dimension.
[0068] In one implementation, in step S340, perform an outer product operation on the time-aligned feature matrix and the behavior pattern vector sequence to generate the initial interaction matrix, which may specifically include the following steps: Step S341: According to the row and column distribution of the time-aligned feature matrix, divide the time-aligned feature matrix into multiple feature sub-blocks, and each feature sub-block corresponds to the behavior pattern vector in the same time window in the behavior pattern vector sequence.
[0069] The time-aligned feature matrix is the matrix obtained after the time alignment operation, and its row and column distribution reflects the data information of different time points and different feature dimensions. According to the row and column distribution of the time-aligned feature matrix, divide it into multiple feature sub-blocks. Each feature sub-block is a subset of the time-aligned feature matrix, and it corresponds to the behavior pattern vector in the same time window in the behavior pattern vector sequence.
[0070] For example, assume that the dimension of the time-aligned feature matrix is m×n, the behavior pattern vector sequence is divided into k windows according to time windows, and the length of each window is w. Then, the time-aligned feature matrix can be divided according to the length of the time window to obtain k feature sub-blocks, and the dimension of each feature sub-block is w×n. Each feature sub-block corresponds to the behavior pattern vector within one time window in the behavior pattern vector sequence.
[0071] Step S342: Perform an outer product operation on each feature sub-block and the behavior pattern vector of the corresponding time window to generate a sub-block interaction matrix.
[0072] Perform an outer product operation on each of the divided feature sub-blocks and the behavior pattern vector of the corresponding time window respectively. The process of the outer product operation is similar to the outer product operation described above. For each element in the feature sub-block and each element in the behavior pattern vector of the corresponding time window, calculate their product to obtain the sub-block interaction matrix.
[0073] Step S343: Perform an edge padding operation on the sub-block interaction matrix to make its dimension consistent with that of the time-aligned feature matrix, and generate the padded sub-block interaction matrix.
[0074] The edge padding operation is to add additional elements to the edges of the sub-block interaction matrix to make its dimension consistent with that of the time-aligned feature matrix. Since the sub-block interaction matrix is obtained by the outer product operation of the feature sub-blocks and the behavior pattern vectors, its dimension may be different from that of the time-aligned feature matrix. Through the edge padding operation, the dimension of the sub-block interaction matrix can be extended to be the same as that of the time-aligned feature matrix, facilitating subsequent splicing operations. Possible edge padding methods include zero padding, that is, adding zero elements to the edges of the sub-block interaction matrix. For example, assuming that the dimension of the sub-block interaction matrix is w×n×p and the dimension of the time-aligned feature matrix is m×n, zero elements can be added above and below and on the left and right of the sub-block interaction matrix to make its dimension become m×n×p.
[0075] Step S344: Splice the padded sub-block interaction matrices in the order of time windows to generate an interaction tensor with a complete time dimension.
[0076] Splice the sub-block interaction matrices after the edge padding operation in the order of time windows. Since each sub-block interaction matrix corresponds to a time window, splicing in time order can obtain an interaction tensor with a complete time dimension.
[0077] Step S345: Perform channel dimension compression on the interaction tensor, and map the multi-dimensional channels to a single correlation strength channel through a summation operation to generate an initial interaction matrix.
[0078] Channel dimension compression is to compress multiple channel dimensions of the interaction tensor into one channel dimension. By performing a summation operation on multiple channel elements at each position of the interaction tensor, the multi-dimensional channels are mapped to a single correlation strength channel to obtain the initial interaction matrix.
[0079] Step S346: Perform a residual connection between the initial interaction matrix and the time-aligned feature matrix, retain the original time-aligned feature information, and generate an enhanced initial interaction matrix.
[0080] Residual connection is used to solve the problems of gradient disappearance and model degradation. In the embodiments of the present invention, a residual connection is performed between the initial interaction matrix and the time-aligned feature matrix, that is, the elements at their corresponding positions are added to obtain an enhanced initial interaction matrix. The enhanced initial interaction matrix contains both the interaction information in the initial interaction matrix and the information in the original time-aligned feature matrix. This can avoid losing important feature information during the calculation process and improve the performance of the model.
[0081] Step S350: Perform a Hadamard product operation on the channel weight vector and the initial interaction matrix to generate a weighted interaction matrix, and input the weighted interaction matrix into a bidirectional recurrent neural network to capture historical dependence features through the forward propagation path and capture potential correlation features through the backward propagation path to generate a bidirectional fusion feature matrix.
[0082] The Hadamard product operation multiplies the elements at the corresponding positions of two matrices. In the embodiments of the present invention, the channel weight vector and the initial interaction matrix are subjected to the Hadamard product operation to obtain a weighted interaction matrix. The channel weight vector reflects the contribution degree of each channel dimension to the risk association. Through the Hadamard product operation, a weight can be assigned to each element in the initial interaction matrix to highlight important channel information.
[0083] A Bidirectional Recurrent Neural Network (BRNN) is a neural network model that can consider both past and future information of sequential data. It consists of a forward propagation path and a backward propagation path. The forward propagation path starts from the beginning position of the sequence and processes each element in the sequence in turn to capture historical dependence features; the backward propagation path starts from the end position of the sequence and processes each element in the sequence backward to capture potential association features. The weighted interaction matrix is input into the bidirectional recurrent neural network, and the network processes the sequential data in the weighted interaction matrix. During the forward propagation process, the network predicts the current element based on the information of the previous elements; during the backward propagation process, the network predicts the current element based on the information of the subsequent elements. Finally, the outputs of the forward propagation and the backward propagation are combined to generate a bidirectional fusion feature matrix.
[0084] Step S360: Perform a strided convolution operation on the bidirectional fusion feature matrix to extract multi-scale spatio-temporal association patterns, and add the weighted interaction matrix and the strided convolution output through a skip connection to generate an association strength matrix.
[0085] The strided convolution operation can skip some elements during the convolution process, thereby expanding the receptive field of the convolution kernel and extracting multi-scale spatio-temporal association patterns. The bidirectional fusion feature matrix contains historical dependence features and potential association features. Through the strided convolution operation, the spatio-temporal association information at different scales in the matrix can be further mined.
[0086] The process of the strided convolution operation is as follows: Define a convolution kernel, and the size and stride of the convolution kernel are preset. Then slide the convolution kernel on the bidirectional fusion feature matrix, and the stride of each slide is determined by the stride parameter. At each sliding position, calculate the sum of the element products of the convolution kernel and the corresponding area of the matrix to obtain the convolution output. The skip connection is a method of connecting the outputs of different layers. In the embodiments of the present invention, the weighted interaction matrix and the strided convolution output are added through a skip connection. In this way, the original information in the weighted interaction matrix can be fused with the multi-scale spatio-temporal association information extracted by the strided convolution to generate an association strength matrix.
[0087] Step S370: Input the association strength matrix into the feature distillation network, extract local association patterns through the separable convolutional layer, and compress redundant features through the pooling layer to output the risk association strength matrix.
[0088] The feature distillation network is a network model for extracting and compressing features, which can further process the association strength matrix, extract features from it, and at the same time compress redundant information.
[0089] The separable convolutional layer is a convolutional layer that decomposes the traditional convolution operation into two steps: depthwise separable convolution and pointwise convolution. The depthwise separable convolution performs convolution operations on each input channel separately to extract local association patterns; the pointwise convolution then performs a linear combination on the output of the depthwise separable convolution to generate the final output. Through the separable convolutional layer, local association patterns in the association strength matrix can be effectively extracted while reducing the computational amount. The pooling layer can reduce the number of features and remove redundant information. Pooling operations include, for example, max pooling and average pooling. In the embodiments of the present invention, the output of the separable convolutional layer is processed by the pooling layer to compress redundant features.
[0090] Input the association strength matrix into the feature distillation network. First, extract local association patterns through the separable convolutional layer, then compress redundant features through the pooling layer, and finally output the risk association strength matrix. The risk association strength matrix is a matrix after further processing and compression, which more concisely represents the association strength between different risk entities.
[0091] Step S380: Input the cross-fusion feature tensor into the feature encoder, expand the feature receptive field through the dilated convolutional layer, and filter the output result of the dilated convolutional layer based on the feature selection weights generated from the risk association strength matrix to generate the risk entity feature matrix.
[0092] Input the cross-fusion feature tensor into the feature encoder for further processing and feature extraction. The dilated convolutional layer is a convolutional layer that inserts holes between the elements of the convolutional kernel to expand the receptive field of the convolutional kernel. Through the dilated convolutional layer, wider feature information can be extracted without increasing the size of the convolutional kernel. For example, the size of a traditional convolutional kernel is 3×3, and the dilated convolutional layer can insert holes between the elements of the convolutional kernel to expand its receptive field to 5×5 or larger. The feature selection weights generated based on the risk association strength matrix assign a weight to each feature in the cross-fusion feature tensor according to the element values in the risk association strength matrix. These weights reflect the importance of each feature to the risk entity. Through the feature selection weights, the output result of the dilated convolutional layer can be filtered to retain important features and remove unimportant features.
[0093] Specifically, the feature selection weights are calculated according to the risk association intensity matrix. Normalization processing can be used to convert the element values in the risk association intensity matrix into feature selection weights. Then, the output result of the dilated convolutional layer is multiplied element by element with the feature selection weights to obtain the filtered features. Finally, the filtered features are combined into a risk entity feature matrix.
[0094] Step S400: Determine the node distribution topology of the knowledge graph according to the entity feature vectors in the risk entity feature matrix, and determine the entity relationship topology of the knowledge graph according to the association intensity values in the risk association intensity matrix.
[0095] In the embodiment of the present invention, the node distribution topology of the knowledge graph is determined according to the entity feature vectors in the risk entity feature matrix, and the entity relationship topology of the knowledge graph is determined according to the association intensity values in the risk association intensity matrix.
[0096] The entity feature vectors in the risk entity feature matrix are the feature descriptions of each risk entity, which contain various attributes and feature information of the risk entity. By processing and analyzing these entity feature vectors, the distribution of nodes in the knowledge graph can be determined. The node distribution topology reflects the positions and layouts of the nodes in the knowledge graph, which can help us intuitively understand the relative relationships between risk entities. The association intensity values in the risk association intensity matrix represent the association degrees between different risk entities. According to these association intensity values, the relationships between entities in the knowledge graph can be determined, including the types and intensities of the relationships. The entity relationship topology reflects the connection methods and relationship intensities between entities in the knowledge graph, which can help us analyze the spread and influence of risks between different entities.
[0097] As an implementation manner, in step S400, to determine the node distribution topology of the knowledge graph according to the entity feature vectors in the risk entity feature matrix, it may specifically include the following steps: Step S410: Perform nonlinear dimensionality reduction processing on each entity feature vector in the risk entity feature matrix to generate low-dimensional embedding vectors, and map the low-dimensional embedding vectors to the topological space coordinate system.
[0098] The purpose of non-linear dimensionality reduction is to reduce the dimensionality of data while preserving the main features of the data, so as to facilitate subsequent visualization and analysis. In the embodiments of the present invention, non-linear dimensionality reduction processing is performed on each entity feature vector in the risk entity feature matrix. Non-linear dimensionality reduction methods include, for example, t-distributed stochastic neighbor embedding (t-SNE) and autoencoder (Autoencoder), etc. Taking t-SNE as an example, it constructs probability distributions between data points in high-dimensional space and low-dimensional space respectively, and then finds the optimal embedding of high-dimensional data in low-dimensional space by minimizing the difference between these two distributions. Specifically, in high-dimensional space, Gaussian distribution is used to measure the similarity between data points; in low-dimensional space, t-distribution is used to measure the similarity. Through continuous iterative optimization, similar data points in high-dimensional space also remain similar in low-dimensional space. The generated low-dimensional embedding vectors are mapped to the topological space coordinate system. The topological space coordinate system can be a two-dimensional plane coordinate system or a three-dimensional space coordinate system.
[0099] Step S420: Based on the distribution density of the embedding vectors in the topological space coordinate system, divide the space into grid cells, and calculate the density gradient value of the entity feature vectors in each space grid cell.
[0100] After mapping the low-dimensional embedding vectors to the topological space coordinate system, the space is divided based on the distribution density of these embedding vectors. The space grid cell is to divide the topological space into several small regions, and each region is a space grid cell. The division method can be selected according to the actual situation. For example, regular rectangular grids or hexagonal grids can be used.
[0101] Calculate the density of the entity feature vectors in each space grid cell. The density can be defined as the ratio of the number of entity feature vectors in the grid cell to the volume of the grid cell. For example, on a two-dimensional plane, the grid cell is a rectangle, and its density is the number of entity feature vectors in the rectangle divided by the area of the rectangle.
[0102] The density gradient value is a physical quantity that describes the density change rate. For each space grid cell, calculate its density gradient value. The density gradient can be approximately calculated by the finite difference method. Taking the two-dimensional plane as an example, for a grid cell (i,j), its density gradient in the x direction can be calculated by dividing the density difference between this cell and its adjacent cell in the x direction by the distance between the adjacent cells; the same is true in the y direction. In this way, the density gradient vector of each space grid cell can be obtained, which reflects the change trend of the density around the cell.
[0103] Step S430: Perform a clustering operation on the space grid cells according to the density gradient values to generate an initial node cluster set, and each node cluster in the initial node cluster set corresponds to a local high-density region in the topological space.
[0104] The clustering operation is a process of dividing data points into different groups or clusters according to their similarity. In the embodiments of the present invention, clustering is performed on the spatial grid cells according to their density gradient values. A density-based clustering algorithm such as the DBSCAN (Density-Based Spatial Clustering of Applications with Noise) algorithm can be used.
[0105] The DBSCAN algorithm divides clusters based on the density of data points. It defines two important parameters: the neighborhood radius and the minimum number of points. For a spatial grid cell, if the number of grid cells within its neighborhood is greater than or equal to the minimum number of points, then this cell is a core point; if a cell is not a core point but is within the neighborhood of a core point, then this cell is a border point; otherwise, it is a noise point.
[0106] Starting from a core point, continuously expand the core points and border points within its neighborhood to form a cluster. Repeat this process until all core points are assigned to a certain cluster, and finally obtain the initial set of node clusters. Each node cluster corresponds to a local high-density region in the topological space, which means that the entity feature vectors within this region have high similarity and may represent a class of similar risk entities.
[0107] Step S440: Extract the boundary overlapping regions between adjacent node clusters, calculate the feature similarity of the entity feature vectors within the boundary overlapping regions, and fuse the node clusters whose feature similarity exceeds the inter-cluster merging threshold to generate an optimized distribution of node clusters.
[0108] The boundary overlapping region between adjacent node clusters is the part where two adjacent node clusters overlap in the topological space. Extract the entity feature vectors within this overlapping region.
[0109] Calculate the feature similarity of the entity feature vectors within the boundary overlapping regions. The feature similarity can be calculated using various methods such as cosine similarity, Euclidean distance, etc.
[0110] Set an inter-cluster merging threshold. When the average feature similarity of the entity feature vectors within the boundary overlapping regions of two adjacent node clusters exceeds this threshold, it is considered that these two node clusters have high similarity and they are fused. The fusion process can be to merge all the entity feature vectors of the two node clusters into a new node cluster.
[0111] In this way, the initial set of node clusters is optimized, some unnecessary partitions are removed, and an optimized distribution of node clusters is obtained, making the node distribution of the knowledge graph more reasonable.
[0112] Step S450: According to the optimized node cluster distribution, calculate the spatial coordinates of the center points of each node cluster, perform weighted interpolation on the spatial coordinates and the risk level scores in the risk entity feature matrix, and generate a node distribution topology.
[0113] For each optimized node cluster, calculate the spatial coordinates of its center point. The calculation method of the center point can be selected according to specific circumstances. For example, calculate the average value of all entity feature vectors in the node cluster in the topological space.
[0114] The risk entity feature matrix contains the risk level scores of each risk entity. Perform weighted interpolation on the spatial coordinates of the node cluster center point and the risk level scores in the risk entity feature matrix. The purpose of weighted interpolation is to comprehensively consider the spatial position and risk level information of the nodes, so that the node distribution topology can more comprehensively reflect the risk situation. For example, linear weighting can be used for interpolation, and the weights can be determined according to the distance from the entity feature vector to the center point. The closer to the center point, the greater the weight. Obtain a comprehensive risk level value through weighted summation, and combine it with the spatial coordinates of the center point to generate a node distribution topology.
[0115] Step S460: Perform back-projection matching between the center points of the node clusters in the node distribution topology and the entity feature vectors in the risk entity feature matrix, establish a one-to-many mapping relationship from entities to nodes, and generate a final node distribution topology structure.
[0116] Back-projection matching is the process of matching the center points of the node clusters in the node distribution topology with the entity feature vectors in the risk entity feature matrix. Since in the previous steps, the entity feature vectors have been processed such as dimensionality reduction and clustering to obtain the center points of the node clusters, now the original entity feature vectors need to be remapped to these nodes.
[0117] The matching can be performed by calculating the similarity between the entity feature vector and the center point of the node cluster. For example, use the Euclidean distance to measure the similarity, and assign the entity feature vector to the node corresponding to the center point of the node cluster with the closest distance. In this way, a one-to-many mapping relationship from entities to nodes is established, that is, multiple entity feature vectors can be mapped to the same node. Through this mapping relationship, a final node distribution topology structure is generated, which clearly shows the corresponding relationship between risk entities and knowledge graph nodes, providing a basis for subsequent knowledge graph construction and risk analysis.
[0118] As an implementation manner, in step S400, according to the association strength values in the risk association strength matrix, determine the entity relationship topology of the knowledge graph, which may specifically include the following steps: Step S470: Perform a symmetrization process on the risk association strength matrix, set the diagonal elements of the matrix to zero and take the maximum value of the upper and lower triangular matrices to generate an undirected association strength matrix.
[0119] The risk association strength matrix is a square matrix, and its elements represent the association strength between different risk entities. In practical applications, risk associations are often undirected, that is, the association strength between entity A and entity B should be the same as that between entity B and entity A. Therefore, it is necessary to symmetrize the risk association strength matrix.
[0120] First, set the diagonal elements of the matrix to zero. The diagonal elements represent the association strength of an entity with itself. In a knowledge graph, this self - association usually has no practical significance, so it is set to zero. Then, take the maximum value of the upper and lower triangular matrices. After such processing, a symmetric matrix, that is, an undirected association strength matrix, is obtained.
[0121] Step S480: Input the undirected association strength matrix into the graph structure generation network, and extract the distribution quantiles of the association strength values through the adaptive threshold segmentation algorithm to generate a dynamic association threshold interval.
[0122] Input the undirected association strength matrix into the graph structure generation network, and the network will analyze and process the association strength values in the matrix. The adaptive threshold segmentation algorithm is an algorithm that can automatically determine the threshold according to the data distribution. In the embodiments of the present invention, the distribution quantiles of the association strength values are extracted through this algorithm. The distribution quantile is a numerical point that divides the data distribution into several equal parts. For example, the median is the 50% quantile.
[0123] The method of quantile regression can be used to determine the distribution quantiles. Quantile regression is a regression method for estimating conditional quantiles, which can calculate the values corresponding to different quantiles according to the data distribution.
[0124] According to the calculated distribution quantiles, a dynamic association threshold interval is generated. For example, the 25% quantile and the 75% quantile can be used as the upper and lower limits of the threshold interval. The dynamically generated association threshold interval can be adjusted according to the actual distribution of the association strength values, and can more accurately reflect the characteristics of the data.
[0125] Step S490: Divide the association strength values into a strong - association interval, a weak - association interval, and a noise interval according to the dynamic association threshold interval, and retain the matrix elements in the strong - association interval to generate an initial edge connection set.
[0126] According to the dynamic correlation threshold interval, the correlation strength values in the undirected correlation strength matrix are divided into three intervals: a strong correlation interval, a weak correlation interval, and a noise interval. The strong correlation interval is the interval where the correlation strength value is greater than the upper limit of the dynamic correlation threshold interval. The correlation strength in this interval is strong, indicating a close connection between two risk entities. The weak correlation interval is the interval where the correlation strength value is within the dynamic correlation threshold interval, and its correlation strength is relatively weak. The noise interval is the interval where the correlation strength value is less than the lower limit of the dynamic correlation threshold interval. The correlations in this interval may be caused by noise or random factors and are considered to have no practical significance.
[0127] Retain the matrix elements in the strong correlation interval, and use the entity pairs corresponding to these elements as the initial edge connection set. The initial edge connection set represents the strong correlation relationships between entities in the knowledge graph and provides a basis for subsequent graph structure construction.
[0128] Step S4100: Perform path connectivity verification on each edge in the initial edge connection set. When there are multiple association paths between two entity nodes, calculate the cumulative correlation strength of each path and retain the path with the maximum value.
[0129] Path connectivity verification is to check whether the edges in the initial edge connection set can form effective path connections between two entity nodes. For each edge in the initial edge connection set, start from one entity node and search along the direction of the edge to see if it can reach the other entity node. When there are multiple association paths between two entity nodes, calculate the cumulative correlation strength of each path. The cumulative correlation strength is the product of the correlation strengths of all the edges on the path (or other appropriate cumulative methods). Retain the path with the maximum cumulative correlation strength. This can remove some redundant paths and make the edge connections in the knowledge graph more concise and effective. Through path connectivity verification and path selection, ensure that the association relationships between entities in the knowledge graph are the most representative and strongest.
[0130] Step S4110: Generate a sparse correlation matrix based on the retained edge connection set, and perform subgraph partitioning operations on the sparse correlation matrix to detect fully connected subgraphs and bridge edge structures.
[0131] Generate a sparse correlation matrix based on the retained edge connection set. The sparse correlation matrix is a matrix with the same dimension as the undirected correlation strength matrix, but most of its elements are zero, and only the elements corresponding to the retained edge connections are non-zero.
[0132] The sub - graph partitioning operation is the process of partitioning the graph represented by the sparse incidence matrix into multiple sub - graphs. Algorithms in graph theory, such as the Kosaraju algorithm or the Tarjan algorithm, can be used for sub - graph partitioning. A fully - connected sub - graph is a sub - graph in which there is an edge connection between any two nodes. A bridging edge is an edge that connects different sub - graphs and plays a key role in the connectivity of the graph. By analyzing the results of sub - graph partitioning, the fully - connected sub - graph and the bridging - edge structure are detected.
[0133] For example, for a graph composed of multiple nodes and edges, after sub - graph partitioning, several independent sub - graphs may be obtained. Some of these sub - graphs are fully - connected sub - graphs, and the edges connecting these sub - graphs are bridging edges. Detecting the fully - connected sub - graph and the bridging - edge structure helps to understand the overall structure and connectivity of the knowledge graph.
[0134] Step S4120: Compress the fully - connected sub - graph into a super - node, calculate the association strength of the bridging edges between the super - nodes, and generate a hierarchical entity - relationship topology.
[0135] Compress the detected fully - connected sub - graph into a super - node. A super - node is an abstract representation that combines multiple nodes into one node, representing all the nodes in the fully - connected sub - graph. By compressing the fully - connected sub - graph into a super - node, the structure of the knowledge graph can be simplified and the number of nodes can be reduced. Calculate the association strength of the bridging edges between the super - nodes. The association strength of the bridging edges can be calculated based on the association strength between the nodes in the fully - connected sub - graphs connected by the bridging edges. For example, the average value of the association strength between all node pairs in the two fully - connected sub - graphs connected by the bridging edge can be taken as the association strength of the bridging edge between the super - nodes.
[0136] Based on the compressed super - nodes and the calculated association strength of the bridging edges, generate a hierarchical entity - relationship topology. The hierarchical entity - relationship topology shows the entity relationships at different levels in the knowledge graph, reflecting the macroscopic associations between entities from the perspective of super - nodes, and providing a clearer structure for subsequent risk analysis and propagation path identification.
[0137] Step S4130: Perform time - window slicing on the hierarchical entity - relationship topology according to the temporal attributes in the risk - entity feature matrix to generate a dynamic entity - relationship topology structure with temporal dependence.
[0138] The risk - entity feature matrix may contain temporal attributes, such as the time when the risk occurs, the duration of the risk, etc. Perform time - window slicing on the hierarchical entity - relationship topology according to these temporal attributes.
[0139] Time - window slicing is to divide the hierarchical entity - relationship topology into multiple time windows in chronological order, and each time window corresponds to a time period. Within each time window, analyze the relationships between entities and generate the entity - relationship topology within that time window.
[0140] For example, assume that the temporal attribute in the risk entity feature matrix is the time of risk occurrence. The time range is divided into multiple time periods. For each time period, entities within that time period and the association relationships between them are extracted to generate the corresponding entity relationship topology.
[0141] Through time window slicing, a dynamic entity relationship topology structure with temporal dependence relationships is generated. This structure can reflect the changes in entity relationships in the knowledge graph over time, and is helpful for analyzing the spread and evolution of risks at different time points.
[0142] Step S4140: Align the dynamic entity relationship topology structure and the node distribution topology spatially, and generate a weighted edge connection topology according to the mapping relationship from entities to nodes.
[0143] Spatial alignment is to match and align the dynamic entity relationship topology structure and the node distribution topology spatially. Since the dynamic entity relationship topology structure describes the relationships between entities, and the node distribution topology describes the distribution of nodes, the two can be combined through spatial alignment.
[0144] According to the mapping relationship from entities to nodes, the edge connections between entities in the dynamic entity relationship topology structure are mapped to the connections between nodes in the node distribution topology. At the same time, weights are assigned to these edge connections, and the weights can be determined according to the association strength between entities.
[0145] In this way, a weighted edge connection topology is generated. This topology structure combines node distribution and entity relationship information, and the edge connections have weights, which can more accurately reflect the association strength and structure between entities in the knowledge graph.
[0146] Step S500: Based on the node distribution topology and the entity relationship topology, generate a dynamic knowledge graph of the target science and technology project, and identify potential risk propagation paths in the dynamic knowledge graph.
[0147] A dynamic knowledge graph is a graph structure that can reflect the changes in knowledge over time. It combines the information of the node distribution topology and the entity relationship topology, and shows the distribution of risk entities in the target science and technology project and the changes in the association relationships between them over time.
[0148] The process of generating a dynamic knowledge graph based on the node distribution topology and the entity relationship topology is to integrate the nodes in the node distribution topology and the edge connections in the entity relationship topology into a graph structure. The position and attributes of the nodes are determined by the node distribution topology, and the connections and weights of the edges are determined by the entity relationship topology. At the same time, considering the dynamics of the entity relationship topology, a time dimension is added to the graph structure so that the knowledge graph can reflect the situation at different time points.
[0149] Identifying potential risk propagation paths in a dynamic knowledge graph is a process of analyzing how risks spread from one node to other nodes in the knowledge graph.
[0150] As an implementation, in step S500, based on the node distribution topology and the entity relationship topology, a dynamic knowledge graph of the target science and technology project is generated, which may specifically include the following steps: Step S510: Extract the cosine similarity of the feature vectors of each entity in the node distribution topology to generate a topology similarity matrix, and perform channel concatenation on the topology similarity matrix and the risk association intensity matrix to form a multi-modal association tensor.
[0151] Perform channel concatenation on the generated topology similarity matrix and the risk association intensity matrix. Channel concatenation is to splice the two matrices in the channel dimension to form a multi-modal association tensor. Suppose the dimension of the topology similarity matrix is n×n, and the dimension of the risk association intensity matrix is also n×n. Then the dimension of the multi-modal association tensor formed after channel concatenation is n×n×2, where the third dimension represents the channel. Through channel concatenation, two different types of information, namely topology similarity and risk association intensity, are integrated into one tensor, providing richer features for subsequent graph convolutional network processing.
[0152] Step S520: Input the multi-modal association tensor into a graph convolutional network, and aggregate the risk features of adjacent nodes through the neighborhood message passing mechanism to generate a node enhanced feature matrix.
[0153] A graph convolutional network (GCN) is a neural network model specifically designed for processing graph-structured data. Inputting the multi-modal association tensor into the graph convolutional network, the network will extract and update the features of the nodes and edges in the graph.
[0154] The neighborhood message passing mechanism is one of the core mechanisms of the graph convolutional network. It passes and aggregates the feature information of adjacent nodes to the current node through the edges connecting the nodes. Specifically, for each node in the graph, it collects the features of its adjacent nodes, performs a weighted sum of these features according to the edge weights, and then merges and transforms the sum result with the features of the current node to obtain the updated node features.
[0155] In the embodiment of the present invention, through the neighborhood message passing mechanism, the graph convolutional network aggregates the risk features of adjacent nodes. After multiple iterations of message passing and feature update, the network can learn the complex relationships and risk propagation patterns between nodes. Finally, a node enhanced feature matrix is generated, and each element in this matrix represents the enhanced node features.
[0156] Step S530: Calculate the directional derivative of each node's eigenvector in the node enhancement feature matrix with respect to the associated edges in the entity relationship topology to generate an edge weight gradient matrix.
[0157] The directional derivative is the rate of change of a vector in a certain direction. In the embodiments of the present invention, according to the eigenvector of each node in the node enhancement feature matrix, the directional derivative of it with respect to the associated edges in the entity relationship topology is calculated. For example, the finite difference method can be used to approximately calculate the directional derivative. Calculate for all the associated edges in the entity relationship topology to obtain the edge weight gradient matrix. Each element of the edge weight gradient matrix represents the directional derivative of the corresponding associated edge, which reflects the change of the node features in the edge direction and provides a basis for subsequent edge weight adjustment.
[0158] Step S540: Perform element-wise multiplication of the edge weight gradient matrix and the risk association strength matrix to generate a dynamic edge weight distribution matrix, and perform quantile normalization processing on the dynamic edge weight distribution matrix to map it to the probability distribution space.
[0159] Element-wise multiplication means multiplying the elements in the corresponding positions of two matrices. Perform element-wise multiplication of the edge weight gradient matrix and the risk association strength matrix to obtain the dynamic edge weight distribution matrix. The dynamic edge weight distribution matrix combines the directional derivative information of the edges and the risk association strength information, and more comprehensively reflects the importance and dynamic changes of the edges.
[0160] Quantile normalization processing is a method to map the elements in the dynamic edge weight distribution matrix to the probability distribution space. First, calculate the quantiles of the elements in the dynamic edge weight distribution matrix, such as the median, the 25th percentile, and the 75th percentile. Then, normalize the elements in the matrix according to these quantiles so that the values of the elements are between 0 and 1 and conform to the characteristics of the probability distribution.
[0161] Through quantile normalization processing, the elements in the dynamic edge weight distribution matrix are converted into probability values, which is convenient for subsequent edge pruning and risk analysis.
[0162] Step S550: Based on the normalized dynamic edge weight distribution matrix, adaptively prune the edge connections in the entity relationship topology, and retain the associated edges whose gradient cumulative amount exceeds the average similarity between nodes.
[0163] Adaptive pruning is a process of screening and removing the edge connections in the entity relationship topology according to the importance of the edges. Based on the normalized dynamic edge weight distribution matrix, the gradient cumulative amount of each edge is calculated. The gradient cumulative amount can be defined as the cumulative value of the dynamic edge weight of the edge within a preset time window. The average similarity between nodes is the average cosine similarity between all node pairs in the node distribution topology. The gradient cumulative amount of each edge is compared with the average similarity between nodes, and the associated edges with the gradient cumulative amount exceeding the average similarity between nodes are retained, while other associated edges are removed. Through adaptive pruning, some unnecessary edge connections are removed, making the knowledge graph more concise and clear, while retaining the most important risk association information.
[0164] Step S560: Align the node enhanced feature matrix with the pruned entity relationship topology in space and time, and adjust the temporal dependence of the node features according to the timestamps of the risk association intensity matrix.
[0165] Spatial-temporal alignment is to match and align the node enhanced feature matrix and the pruned entity relationship topology in space and time. Since the node enhanced feature matrix reflects the feature information of the nodes, and the pruned entity relationship topology reflects the connection relationship between the nodes, the two can be combined through spatial-temporal alignment.
[0166] Adjust the temporal dependence of the node features according to the timestamps of the risk association intensity matrix. The timestamps in the risk association intensity matrix record the risk association situations at different time points. According to these timestamps, the node features in the node enhanced feature matrix are adjusted so that the node features can reflect the risk states at different time points.
[0167] For example, for a certain node, at different time points, its feature values may change according to the change of the risk association intensity. By adjusting the temporal dependence of the node features, the dynamic knowledge graph can more accurately reflect the evolution of risks over time.
[0168] Step S570: Jointly encode the spatially and temporally aligned node features and associated edges through a graph structure encoder to generate the graph embedding representation of the dynamic knowledge graph, where the node embedding vector contains the risk propagation potential coefficient and the edge embedding vector contains the risk attenuation factor.
[0169] The graph structure encoder is a model used to process graph data and convert it into a low-dimensional vector representation. The spatially and temporally aligned node features contain the risk state information of the nodes at different time points, and the associated edges reflect the connection relationship between the nodes and the risk propagation path.
[0170] Taking the Graph Neural Network (GNN) as an example of a graph structure encoder, it can encode nodes and edges through a message passing mechanism. The message passing mechanism is the core of GNN, and its basic idea is that nodes update their own feature representations by exchanging information with adjacent nodes. Specifically, for each node, it collects the feature information of its adjacent nodes, weights these information according to the weights of the edges, then combines and transforms the summation result with its own features to obtain the updated node features.
[0171] During the encoding process, first, the node features and the information of associated edges after spatio-temporal alignment are input into the input layer of the GNN. The input layer converts this information into a format suitable for model processing. Then, in the hidden layer, through multiple iterative message passing processes, the GNN allows nodes to continuously absorb the information of adjacent nodes, thereby learning the global features of nodes in the graph structure. Each message passing can be regarded as an aggregation and propagation of information, enabling the features of nodes to gradually contain more graph structure information. For example, for node i, its update formula in the k-th message passing can be expressed as: ; where is the feature representation of node i after the k-th iteration, N(i) is the set of adjacent nodes of node i, is the normalization coefficient of edge (i, j), and are learnable weight matrices, is the activation function.
[0172] For associated edges, they can also be encoded during the message passing process. The information of edges can affect the transmission weights of information between nodes. For example, the risk attenuation factor of an edge can be used as part of the weight to adjust the intensity of information transmission between nodes.
[0173] Finally, in the output layer, the GNN maps the learned feature representations of nodes and edges to a low-dimensional space to generate the graph embedding representation of the dynamic knowledge graph.
[0174] The node embedding vector contains a risk propagation potential energy coefficient because in the scenario of risk analysis, each node has different potential influences during the risk propagation process. The risk propagation potential energy coefficient reflects the ability of a node to act as a risk source or an intermediate node in risk propagation. During the encoding process, the graph structure encoder encodes this potential influence into the node embedding vector by learning the features of the node and its position in the graph. For example, if a node is connected to multiple high-risk nodes and has high resources or influence itself, then its risk propagation potential energy coefficient will be high. The edge embedding vector contains a risk attenuation factor because when risk propagates through an edge, it will be attenuated by various factors. The risk attenuation factor describes the degree of attenuation of risk when propagating on an edge connection. In the graph structure encoder, the risk attenuation factor is encoded into the edge embedding vector by learning the features of the edge (such as the type, weight, etc.) and its role in the graph. For example, for an edge connecting two nodes that are far apart or weakly associated, its risk attenuation factor may be large, meaning that the risk will decay rapidly when propagating on this edge.
[0175] Through the joint encoding of the graph structure encoder, the generated graph embedding representation can comprehensively reflect the features of nodes and edges in the dynamic knowledge graph and the relationships between them, providing an effective feature representation for subsequent risk propagation path identification and risk analysis.
[0176] As an implementation, in step S500, to identify the potential risk propagation paths in the dynamic knowledge graph, it can specifically include the following steps: Step S580: Extract the risk propagation potential energy coefficient from the node attributes of the dynamic knowledge graph and extract the risk attenuation factor from the edge attributes.
[0177] The node attributes of the dynamic knowledge graph contain various information of the nodes, and the risk propagation potential energy coefficient is an important parameter in the node embedding vector generated by the graph structure encoder in the previous steps. By accessing the node attributes, the risk propagation potential energy coefficient of each node is extracted.
[0178] The edge attributes also contain relevant information of the edges, and the risk attenuation factor is a parameter in the edge embedding vector. The risk attenuation factor of each edge is extracted from the edge attributes.
[0179] Step S590: Use the nodes with a risk propagation potential energy coefficient greater than the average value of the same type of risks in the node attributes as candidate risk source nodes, and generate the reverse propagation path priority according to the reverse order of the risk attenuation factors of the edge attributes.
[0180] Calculate the average value of the same type of risks in the computing node attributes. The same type of risks refer to the risks faced by nodes with similar risk characteristics. For each node, compare the magnitude of its risk propagation potential energy coefficient with the average value of the same type of risks. Nodes with a risk propagation potential energy coefficient greater than the average value of the same type of risks are regarded as candidate risk source nodes.
[0181] Arrange the edges in reverse order according to the risk attenuation factor of the edge attributes. The smaller the risk attenuation factor, the smaller the attenuation of the risk during its propagation along this edge, and the more conducive it is to the propagation of the risk. Therefore, arrange the edges with smaller risk attenuation factors in the front to generate the priority of the reverse propagation path.
[0182] Step S5100: Perform a breadth-first traversal on the candidate risk source nodes along the priority of the reverse propagation path, and record the node sequence passed during the traversal and the risk attenuation factors of the corresponding edges.
[0183] In the embodiment of the present invention, perform a breadth-first traversal on the candidate risk source nodes along the priority of the reverse propagation path. Starting from each candidate risk source node, select the next node to be visited according to the priority of the reverse propagation path. During the traversal, record the passed node sequence and the risk attenuation factors of the corresponding edges. For example, starting from the candidate risk source node A, according to the priority of the reverse propagation path, first visit the node B, and record the node sequence [A, B] and the risk attenuation factor d of the edge (A, B). AB ; then continue to visit the next node from the node B, and so on, until all reachable nodes are traversed.
[0184] Step S5110: Perform a conduction path integrity verification on each node sequence. When the risk propagation potential energy coefficients of adjacent nodes in the sequence satisfy the exponential decay law, calculate the cumulative conduction probability of this conduction path.
[0185] The conduction path integrity verification is a process of checking whether the node sequence forms a complete risk conduction path. For each node sequence, check whether the risk propagation potential energy coefficients of adjacent nodes in the sequence satisfy the exponential decay law. The exponential decay law can be expressed as , where p i and p i+1 are the risk propagation potential energy coefficients of adjacent nodes, is the attenuation coefficient. When the risk propagation potential energy coefficients of adjacent nodes satisfy the exponential decay law, calculate the cumulative conduction probability of this conduction path. The cumulative conduction probability can be calculated according to the risk attenuation factors of each edge on the path.
[0186] Step S5120: Perform a convolution operation on the cumulative conduction probability and the risk propagation potential energy coefficient of the node sequence to generate a path comprehensive risk assessment value.
[0187] The comprehensive risk assessment value of the path comprehensively considers the risk propagation potential of nodes and the conduction probability of the path, and more comprehensively evaluates the possibility of risk propagation and the degree of influence on this path.
[0188] Step S5130: Compare the comprehensive risk assessment value of the path with the benchmark value of similar historical risk cases, and screen out the abnormal conduction paths whose assessment values deviate from the standard deviation range of the benchmark value.
[0189] Similar historical risk cases are risk cases in historical projects with similar characteristics and risk situations to the current target science and technology project. For these similar historical risk cases, calculate the benchmark value and standard deviation of their comprehensive risk assessment values of the path.
[0190] Compare the comprehensive risk assessment value of each conduction path with the benchmark value of similar historical risk cases. When the assessment value deviates from the standard deviation range of the benchmark value, consider this conduction path as an abnormal conduction path.
[0191] Step S5140: Perform two-way risk diffusion simulation on the abnormal conduction path to capture the risk amplification nodes in the forward propagation path and the key inhibition nodes in the reverse blocking path.
[0192] Two-way risk diffusion simulation is a simulation process that simultaneously conducts forward risk diffusion and reverse risk blocking. For the abnormal conduction path, conduct forward risk diffusion simulation, starting from the risk source node, simulate the propagation process of risk on the path, and observe which nodes can amplify the risk. These nodes are the risk amplification nodes. At the same time, conduct reverse risk blocking simulation, starting from the end node of the path, reverse simulate the risk blocking process, and find out which nodes can effectively inhibit the propagation of risk. These nodes are the key inhibition nodes.
[0193] Step S5150: Generate a multi-dimensional risk conduction map containing path weights and intervention priorities according to the spatial distribution density of risk amplification nodes and key inhibition nodes.
[0194] Calculate the spatial distribution density of risk amplification nodes and key inhibition nodes. The spatial distribution density can be defined as the ratio of the number of risk amplification nodes or key inhibition nodes to the area or volume of this spatial range within a predefined spatial range.
[0195] According to the spatial distribution density of risk amplification nodes and key inhibition nodes, assign path weights to each abnormal conduction path. The path weights can be determined according to the size of the spatial distribution density. The larger the spatial distribution density, the greater the path weight.
[0196] At the same time, determine the intervention priority according to the path weight and the severity of the risk. The higher the path weight and the more severe the risk of the path, the higher its intervention priority.
[0197] Integrate path weights and intervention priority information into a graph to generate a multi-dimensional risk conduction graph that includes path weights and intervention priorities. This graph can visually display the importance and intervention order of abnormal conduction paths, providing strong support for risk control and management.
[0198] It should be noted that in the various operation links involved above, if there is a fusion calculation of variables with different dimensions, or a fusion calculation of variables with different dimensions, those skilled in the art can perform preprocessing such as dimension elimination and dimension alignment according to conventional techniques in the art (such as normalization, linear interpolation, truncation, etc.), and then perform subsequent corresponding processing. Since this belongs to the conventional technical means in the art, it is not described in detail in the embodiments of the present invention.
[0199] As an implementation method, the method provided in the embodiments of the present invention further includes a dynamic knowledge graph update mechanism, which specifically includes the following steps: Step S600: Real-time monitor the newly added heterogeneous data of the target science and technology project. When it is detected that the numerical change of the structured index data exceeds the set sensitivity or a new risk keyword appears in the unstructured text data, trigger a graph update event.
[0200] Real-time monitoring of the newly added heterogeneous data of the target science and technology project means continuously collecting and analyzing the new structured index data, unstructured text data, and time-series behavior data generated in the target science and technology project. The numerical change of the structured index data may reflect the actual progress and changes in the risk status of the project. The appearance of a new risk keyword in the unstructured text data may imply the emergence of new risk factors. The set sensitivity is a preset threshold used to determine whether the numerical change of the structured index data is significant. For example, for the budget index of a project, if its numerical change exceeds the set percentage (such as 5%), it is considered that the numerical change exceeds the set sensitivity. A new risk keyword is a keyword that appears in the unstructured text data and was not previously recognized as risk-related. New risk keywords can be identified through keyword matching or natural language processing techniques. When it is detected that the numerical change of the structured index data exceeds the set sensitivity or a new risk keyword appears in the unstructured text data, trigger a graph update event. The triggering of the graph update event means that the dynamic knowledge graph needs to be updated to reflect the latest risk situation of the project.
[0201] Step S700: Perform incremental feature extraction on the newly added heterogeneous data to generate an incremental entity feature vector and an incremental association strength matrix.
[0202] Incremental feature extraction means extracting useful feature information from newly added heterogeneous data. For the newly added structured metric data, it is necessary to perform preprocessing and feature extraction on it, such as removing outliers, normalizing, etc., and then extract its feature vectors. For the newly added unstructured text data, semantic analysis and feature extraction are required, such as identifying entities and relationships therein and extracting semantic features of the text. For the newly added time-series behavior data, it is necessary to analyze its behavior patterns and time features and extract corresponding feature vectors.
[0203] Generating incremental entity feature vectors is to integrate the feature vectors extracted from the newly added heterogeneous data into feature vectors representing the newly added entities. The incremental association strength matrix is a matrix that reflects the association strength between newly added entities and existing entities as well as between newly added entities.
[0204] As an implementation manner, in step S700, perform incremental feature extraction on the newly added heterogeneous data to generate incremental entity feature vectors and an incremental association strength matrix, which may specifically include the following steps: Step S710: Perform moving average filtering on the newly added structured metric data to eliminate sudden noise interference.
[0205] Through the moving average filtering process, the newly added structured metric data can be made smoother, reducing the impact of noise on subsequent feature extraction and analysis.
[0206] Step S720: Perform named entity recognition on the newly added unstructured text data to extract new risk entities not registered in the existing knowledge graph.
[0207] Step S730: Perform operation mode clustering analysis on the newly added time-series behavior data to identify new mode categories whose distances from the cluster centers of the existing behavior mode vector sequences exceed a set radius.
[0208] Operation mode clustering analysis is a process of classifying and clustering operation modes in the newly added time-series behavior data. Clustering algorithms such as the K-Means algorithm or the DBSCAN algorithm can be used to cluster the newly added time-series behavior data.
[0209] First, calculate the cluster centers of the existing behavior mode vector sequences. Then, for each operation mode vector in the newly added time-series behavior data, calculate its distance from the existing cluster centers. Set a radius threshold. When the distance of an operation mode vector from the existing cluster center exceeds this radius, it is considered that this operation mode belongs to a new mode category.
[0210] Step S740: Map the new risk entities and new mode categories to the feature space to generate incremental entity feature vectors.
[0211] Mapping new risk entities and new mode categories to the feature space is a process of converting them into vector representations. For new risk entities, features can be extracted and converted into vectors based on their attributes and relevant information. For new mode categories, feature vectors can be generated according to the characteristics of their operation modes, such as operation frequency, operation time interval, etc. Combine the feature vectors of new risk entities and new mode categories to generate incremental entity feature vectors.
[0212] Step S750: Calculate the association strength between the new risk entity and the existing entities, as well as the association degree between the new mode category and the existing behavior patterns, and generate an incremental association strength matrix.
[0213] Multiple methods can be used to calculate the association strength between the new risk entity and the existing entities, such as methods based on feature similarity or rules. For example, for the feature vectors of the new risk entity and the existing entities, calculate the cosine similarity between them, and use the similarity value as the association strength. Calculating the association degree between the new mode category and the existing behavior patterns can be achieved by comparing their operation mode characteristics. For example, calculate the similarity of features such as operation frequency and operation time interval between the new mode category and the existing behavior patterns, and use the similarity value as the association degree.
[0214] Combine the association strength between the new risk entity and the existing entities and the association degree between the new mode category and the existing behavior patterns into a matrix to generate an incremental association strength matrix.
[0215] Step S800: Perform similarity matching between the incremental entity feature vector and the existing node distribution topology. When the similarity is lower than the matching threshold, create a new node; otherwise, update the feature vector of the existing node.
[0216] Perform similarity matching between the incremental entity feature vector and the node feature vectors in the existing node distribution topology. Similarity matching can use methods such as cosine similarity and Euclidean distance.
[0217] Set a matching threshold. When the similarity between the incremental entity feature vector and the existing node feature vector is lower than this threshold, it is considered that the incremental entity represents a new risk entity, and a new node is created to represent this entity. The feature vector of the new node is the incremental entity feature vector.
[0218] When the similarity is higher than the matching threshold, it indicates that the incremental entity has a high similarity with the existing node, and update the feature vector of the existing node. The weighted average method can be used to perform weighted averaging on the incremental entity feature vector and the existing node feature vector to obtain the updated node feature vector.
[0219] Step S900: Perform matrix superposition on the incremental association strength matrix and the existing entity relationship topology to generate an updated edge weight distribution.
[0220] Matrix superposition is to add the elements at the corresponding positions of the incremental association strength matrix and the association strength matrix in the existing entity relationship topology. The association strength matrix in the existing entity relationship topology represents the association strength between existing entities, and the incremental association strength matrix represents the association strength between newly added entities and existing entities as well as between newly added entities.
[0221] Through matrix superposition, the newly added association strength information is integrated into the existing entity relationship topology to generate an updated edge weight distribution. The updated edge weight distribution more comprehensively reflects the association relationship between entities in the knowledge graph.
[0222] Step S1000: According to the updated node distribution and edge weight distribution, re - execute the risk conduction path analysis and update the set of potential risk propagation paths.
[0223] According to the updated node distribution and edge weight distribution, re - execute the risk conduction path analysis process described in the previous steps. This includes steps such as identifying candidate risk source nodes, generating the priority of the reverse propagation path, performing breadth - first traversal, calculating the cumulative conduction probability, and generating the comprehensive risk assessment value of the path. Through re - analysis, a new set of potential risk propagation paths is obtained. Updating the set of potential risk propagation paths can make the dynamic knowledge graph more accurately reflect the latest risk status of the target scientific and technological project, providing more timely and effective information for risk control and management.
[0224] As an implementation manner, the method provided by the embodiments of the present invention further includes a process for generating a risk mitigation strategy, which specifically includes the following steps: Step S1100: After identifying the potential risk propagation path, extract the set of influence factors of the key nodes in the path.
[0225] After identifying the potential risk propagation path, determine the key nodes in the path. Key nodes are nodes that play an important role in the risk propagation process, such as risk amplification nodes or key suppression nodes. Extract the set of influence factors of the key nodes. Influence factors are various factors related to the key nodes, and these factors will affect the role of the nodes in risk propagation. For example, for a project task node, its influence factors may include the completion time of the task, resource requirements, technical difficulty, etc.
[0226] By analyzing the attributes and relevant information of the key nodes, extract the set of their influence factors. These sets of influence factors will be used as the input for subsequent optimization of the risk mitigation strategy.
[0227] Step S1200: Construct a risk mitigation strategy optimization model. The risk mitigation strategy optimization model takes minimizing the risk level score of the key nodes as the objective function and the project resource constraint as the boundary condition.
[0228] The risk mitigation strategy optimization model is a mathematical model used to find the optimal risk mitigation strategy. The goal of this model is to minimize the risk level score of critical nodes, that is, by reasonably allocating resources and taking corresponding measures, the risk degree of critical nodes is reduced.
[0229] The objective function can be expressed as the sum of the risk level scores of critical nodes. The project resource constraint means that when implementing the risk mitigation strategy, it is restricted by the available project resources. Resource constraints can include limitations in aspects such as human resources, material resources, and financial resources. For example, the total project budget cannot exceed the preset amount, and the available human resources are limited, etc.
[0230] Incorporate the project resource constraint as a boundary condition into the risk mitigation strategy optimization model to ensure that the optimized risk mitigation strategy is within the scope allowed by the project resources.
[0231] Step S1300: Solve the risk mitigation strategy optimization model through the gradient descent algorithm to obtain the optimal resource allocation plan for each critical node.
[0232] By solving through the gradient descent algorithm, the optimal resource allocation plan for each critical node is obtained. This plan can minimize the risk level score of critical nodes on the premise of meeting the project resource constraints.
[0233] Step S1400: Convert the optimal resource allocation plan into a natural language description and associate it with the corresponding node in the knowledge graph.
[0234] Converting the optimal resource allocation plan into a natural language description is for the convenience of project managers to understand and implement. Associate the natural language description of the optimal resource allocation plan with the corresponding node in the knowledge graph. In the knowledge graph, each critical node has its corresponding attributes and relationships. Associate the optimal resource allocation plan as an attribute of the node to this node. In this way, when viewing the knowledge graph, the risk mitigation strategy of each critical node can be intuitively understood.
[0235] Step S1500: Recalculate the comprehensive risk value of the potential risk propagation path according to the change in the risk level of the nodes after resource allocation.
[0236] After allocating resources to critical nodes according to the optimal resource allocation plan, the risk level of critical nodes will change. Because the allocation of resources may improve the situation of the nodes and reduce their risk degree. Recalculate the comprehensive risk value of the potential risk propagation path. The calculation method of the comprehensive risk value is the same as described in the previous steps, that is, considering the risk propagation potential energy coefficient of each node on the path and the risk attenuation factor of the edge, calculating the cumulative conduction probability, and performing a convolution operation on the cumulative conduction probability and the risk propagation potential energy coefficient of the node sequence.
[0237] By recalculating the comprehensive risk value of potential risk propagation paths, the effectiveness of risk mitigation strategies can be evaluated, providing a basis for further risk control and management.
[0238] As an implementation, the risk mitigation strategy optimization model includes: defining a decision variable matrix, where each element represents the type and quantity of resources allocated to a node.
[0239] The decision variable matrix is an important part of the risk mitigation strategy optimization model, which is used to represent the resource situation allocated to each key node. Each row of the matrix corresponds to a key node, and each column corresponds to a type of resource. Each element in the matrix represents the quantity of the resource type allocated to the node.
[0240] Construct a resource benefit function to calculate the attenuation coefficient of each resource type for the node risk level score.
[0241] The resource benefit function is a function used to describe the relationship between resource input and the reduction of the node risk level score. Different resource types have different degrees of influence on the node risk level score. By constructing the resource benefit function, the attenuation coefficient of each resource type for the node risk level score can be calculated.
[0242] Through the analysis and modeling of historical data, the specific form and parameters of the resource benefit function can be determined, so as to calculate the attenuation coefficient of each resource type for the node risk level score.
[0243] Establish a resource constraint inequality group to ensure that the total resource consumption does not exceed the project budget and complies with the resource allocation rules.
[0244] The resource constraint inequality group is to ensure that during the resource allocation process, the total resource consumption does not exceed the project budget and complies with the resource allocation rules. The project budget is the total available resources of the project, and the resource allocation rules are some restrictive conditions for resource allocation, such as the allocation ratio of certain resources cannot exceed the preset range.
[0245] Introduce a risk conduction inhibition factor, which is inversely proportional to the edge conduction probability in the knowledge graph.
[0246] The risk conduction inhibition factor is a parameter used to measure the degree of risk inhibition on edge connections. It is inversely proportional to the edge conduction probability in the knowledge graph, that is, the larger the risk conduction inhibition factor, the smaller the edge conduction probability, and the lower the possibility of risk propagation on this edge.
[0247] Introducing the risk conduction inhibition factor can more accurately describe the risk propagation situation in the knowledge graph, and considering the risk conduction inhibition factor in the risk mitigation strategy optimization model can better control the risk propagation.
[0248] By adjusting the value of the risk conduction inhibition factor, the edge conduction probability can be changed, thereby affecting the risk propagation path and the comprehensive risk value.
[0249] The constraint conditions are integrated into the objective function by the Lagrange multiplier method to form an unconstrained optimization problem.
[0250] In the risk mitigation strategy optimization model, the objective function is to minimize the risk level score of key nodes, and there is a system of resource constraint inequalities. By introducing Lagrange multipliers, the system of resource constraint inequalities is transformed into equality constraints and integrated into the objective function to form an unconstrained optimization problem. Specifically, for each resource constraint inequality, a Lagrange multiplier is introduced, multiplied by the inequality constraint, and added to the objective function.
[0251] By solving this unconstrained optimization problem, the optimal solution that satisfies the resource constraint conditions can be obtained. The adaptive learning rate optimization algorithm is used to iteratively update the decision variable matrix until the objective function converges to a local optimal solution.
[0252] Please refer to Figure 2 , Figure 2 FIG. 16 is a schematic structural diagram of a computer system provided by an embodiment of the present invention. The computer system at least includes a processor 101, a communication interface 102, and a memory 103. Among them, the processor 101, the communication interface 102, and the memory 103 can be connected through a bus or other means. Among them, the processor 101 (or the central processing unit (CPU)) is the computing core and control core of the computer system, which can parse various instructions in the computer system and process various data in the computer system. The communication interface 102 may optionally include a standard wired interface, a wireless interface (such as WI-FI, a mobile communication interface, etc.), and can be used to send and receive data under the control of the processor 101; the communication interface 102 can also be used for the transmission and interaction of internal data in the computer system. The memory 103 (Memory) is a memory device in the computer system, used to store programs and data. It can be understood that the memory 103 here can include both the built-in memory of the computer system and, of course, the extended memory supported by the computer system. The memory 103 provides a storage space, and the operating system of the computer system is stored in this storage space, which is not limited in the present invention.
[0253] In one embodiment, the processor 101 runs the computer program in the memory 103 to execute the knowledge graph generation method for science and technology project risk control provided above in the embodiments of the present invention.
Claims
1. A method for generating a knowledge graph for risk control of scientific and technological projects, characterized in that, The method includes: obtaining a multi-source heterogeneous data set of a target science and technology project, where the multi-source heterogeneous data set includes structured index data, unstructured text data, and time-series behavior data; through a heterogeneous data fusion mechanism, converting the structured index data into a standard vector sequence, converting the unstructured text data into a semantic vector sequence, and converting the time-series behavior data into a behavior pattern vector sequence; inputting the standard vector sequence, semantic vector sequence, and behavior pattern vector sequence into a risk quantification assessment model to generate a risk entity feature matrix and a risk association intensity matrix; determining the node distribution topology of a knowledge graph according to the entity feature vectors in the risk entity feature matrix, and determining the entity relationship topology of the knowledge graph according to the association intensity values in the risk association intensity matrix; generating a dynamic knowledge graph of the target science and technology project based on the node distribution topology and the entity relationship topology, and identifying potential risk propagation paths in the dynamic knowledge graph.
2. The method according to claim 1, wherein The step of converting the structured index data into a standard vector sequence, converting the unstructured text data into a semantic vector sequence, and converting the time-series behavior data into a behavior pattern vector sequence through the heterogeneous data fusion mechanism includes: performing abnormal state detection on the structured index data, generating a data correction mask matrix according to the detection results, using the data correction mask matrix to perform interpolation reconstruction on the original index data, and outputting the corrected structured index data; inputting the corrected structured index data into an index encoder, and mapping the index values with different dimensions to a unified metric space through a feature dimension alignment algorithm to generate a standard vector sequence; performing semantic segment cutting on the unstructured text data, identifying text paragraphs containing risk factor descriptions, inputting the text paragraphs into a pre-trained semantic encoder to generate primary semantic vectors, and performing spatial transformation on the primary semantic vectors through a context-aware network to output a semantic vector sequence; performing event window division on the time-series behavior data, extracting the type distribution features and time interval features of the behavior operations within the window, and performing orthogonal projection transformation on the type distribution features and the time interval features to generate a behavior pattern vector sequence.
3. The method according to claim 2, wherein Performing abnormal state detection on the structured indicator data, generating a data correction mask matrix according to the detection results, and performing interpolation reconstruction on the original indicator data by using the data correction mask matrix to output the corrected structured indicator data, including: obtaining the historical indicator data of similar scientific and technological projects, calculating the mean vector and covariance matrix of each indicator dimension, and generating a multi-dimensional normal distribution model based on the mean vector and covariance matrix; inputting the structured indicator data of the current project into the multi-dimensional normal distribution model, calculating the probability density value of each data point, and generating an abnormal marking sequence when the probability density value is lower than the preset quantile of the distribution model; generating a binary mask matrix according to the position information of the abnormal points in the abnormal marking sequence, where the positions corresponding to the abnormal points in the binary mask matrix take zero values, and the positions corresponding to the normal points take one values; performing a morphological closing operation on the binary mask matrix to eliminate isolated abnormal marking points and generate a connected region correction mask matrix; positioning abnormal data points in the original indicator data based on the abnormal regions marked by the connected region correction mask matrix, and extracting the indicator values of adjacent data points within the time window before and after the abnormal data points; inputting the indicator values of the adjacent data points into a moving average filter to generate an interpolation filling sequence, and replacing the original abnormal data points with the interpolation filling sequence according to the zero value positions of the connected region correction mask matrix; outputting the replaced structured indicator data as the corrected structured indicator data.
4. The method according to claim 2, characterized in that, Inputting the standard vector sequence, semantic vector sequence, and behavior pattern vector sequence into a risk quantification and assessment model to generate a risk entity feature matrix and a risk association intensity matrix includes: performing cross-attention calculation on the standard vector sequence and the semantic vector sequence to generate a cross-attention weight matrix between indicators and semantics, where each element of the cross-attention weight matrix reflects the association intensity between the indicator vector and the semantic vector; performing weighted aggregation on the semantic vector sequence based on the cross-attention weight matrix to generate a semantically enhanced indicator feature sequence, and multiplying the semantically enhanced indicator feature sequence element-wise with the standard vector sequence to generate a cross-fusion feature tensor; unfolding the cross-fusion feature tensor along the feature dimension into a two-dimensional matrix, and performing a time alignment operation with the behavior pattern vector sequence, adjusting the row and column distribution of the two-dimensional matrix according to the timestamp information of the behavior pattern vector sequence to generate a time-aligned feature matrix; performing an outer product operation on the time-aligned feature matrix and the behavior pattern vector sequence to generate an initial interaction matrix, and performing channel attention weighting on the initial interaction matrix to calculate the contribution degree of each channel dimension to risk association to generate a channel weight vector; performing a Hadamard product operation on the channel weight vector and the initial interaction matrix to generate a weighted interaction matrix, and inputting the weighted interaction matrix into a bidirectional recurrent neural network to capture historical dependence features through the forward propagation path and potential association features through the backward propagation path to generate a bidirectional fusion feature matrix; performing a strided convolution operation on the bidirectional fusion feature matrix to extract multi-scale spatio-temporal association patterns, and adding the weighted interaction matrix and the strided convolution output through skip connections to generate an association intensity matrix; inputting the association intensity matrix into a feature distillation network to extract local association patterns through separable convolutional layers and compress redundant features through pooling layers to output a risk association intensity matrix; inputting the cross-fusion feature tensor into a feature encoder to expand the feature receptive field through dilated convolutional layers, and filtering the output result of the dilated convolutional layers based on the feature selection weights generated by the risk association intensity matrix to generate a risk entity feature matrix.
5. The method according to claim 4, wherein Performing an outer product operation on the time-aligned feature matrix and the behavior pattern vector sequence to generate an initial interaction matrix includes: dividing the time-aligned feature matrix into multiple feature sub-blocks according to the row and column distribution of the time-aligned feature matrix, where each feature sub-block corresponds to a behavior pattern vector in the same time window of the behavior pattern vector sequence; performing an outer product operation on each feature sub-block and the behavior pattern vector of the corresponding time window to generate a sub-block interaction matrix; performing an edge padding operation on the sub-block interaction matrix to make its dimension consistent with that of the time-aligned feature matrix, generating a padded sub-block interaction matrix; splicing the padded sub-block interaction matrices in the order of time windows to generate an interaction tensor with a complete time dimension; performing channel dimension compression on the interaction tensor, mapping the multi-dimensional channels to a single correlation intensity channel through a summation operation, generating an initial interaction matrix; performing a residual connection on the initial interaction matrix and the time-aligned feature matrix to retain the original time-aligned feature information, generating an enhanced initial interaction matrix.
6. The method according to claim 1, wherein Determining the node distribution topology of the knowledge graph according to the entity feature vectors in the risk entity feature matrix includes: performing a non-linear dimensionality reduction process on each entity feature vector in the risk entity feature matrix to generate a low-dimensional embedding vector, and mapping the low-dimensional embedding vector to the topological space coordinate system; dividing space grid cells based on the distribution density of the embedding vectors in the topological space coordinate system, and calculating the density gradient value of the entity feature vectors in each space grid cell; performing a clustering operation on the space grid cells according to the density gradient value to generate an initial node cluster set, where each node cluster in the initial node cluster set corresponds to a local high-density region in the topological space; extracting the boundary overlapping regions between adjacent node clusters, calculating the feature similarity of the entity feature vectors in the boundary overlapping regions, and fusing the node clusters with the feature similarity exceeding the inter-cluster merging threshold to generate an optimized node cluster distribution; calculating the spatial coordinates of the center points of each node cluster according to the optimized node cluster distribution, and performing weighted interpolation on the spatial coordinates and the risk level scores in the risk entity feature matrix to generate a node distribution topology; performing back-projection matching on the center points of the node clusters in the node distribution topology and the entity feature vectors in the risk entity feature matrix to establish a one-to-many mapping relationship from entities to nodes, generating a final node distribution topology structure.
7. The method according to claim 6, wherein Determining the entity relationship topology of the knowledge graph according to the association strength values in the risk association strength matrix includes: performing a symmetrization process on the risk association strength matrix, setting the diagonal elements of the matrix to zero and taking the maximum value of the upper and lower triangular matrices to generate an undirected association strength matrix; inputting the undirected association strength matrix into a graph structure generation network, extracting the distribution quantiles of the association strength values through an adaptive threshold segmentation algorithm to generate a dynamic association threshold interval; dividing the association strength values into a strong association interval, a weak association interval, and a noise interval according to the dynamic association threshold interval, and retaining the matrix elements within the strong association interval to generate an initial edge connection set; performing path connectivity verification on each edge in the initial edge connection set, and when there are multiple association paths between two entity nodes, calculating the cumulative association strength of each path and retaining the maximum value path; generating a sparse association matrix according to the retained edge connection set, performing a subgraph partitioning operation on the sparse association matrix, and detecting fully connected subgraphs and bridge edge structures; compressing the fully connected subgraphs into super nodes, calculating the association strength of the bridge edges between the super nodes to generate a hierarchical entity relationship topology; performing time window slicing on the hierarchical entity relationship topology according to the temporal attributes in the risk entity feature matrix to generate a dynamic entity relationship topology structure with temporal dependence; spatially aligning the dynamic entity relationship topology structure with the node distribution topology, and generating a weighted edge connection topology according to the mapping relationship from the entity to the node.
8. The method according to claim 1, wherein Generating a dynamic knowledge graph of the target scientific and technological project based on the node distribution topology and the entity relationship topology includes: extracting the cosine similarity of the entity feature vectors in the node distribution topology to generate a topology similarity matrix, and performing channel concatenation on the topology similarity matrix and the risk association strength matrix to form a multi-modal association tensor; inputting the multi-modal association tensor into a graph convolutional network, aggregating the risk features of adjacent nodes through a neighborhood message passing mechanism to generate a node enhanced feature matrix; calculating the directional derivative of each node's feature vector in the node enhanced feature matrix with respect to the associated edges in the entity relationship topology to generate an edge weight gradient matrix; performing element-wise multiplication on the edge weight gradient matrix and the risk association strength matrix to generate a dynamic edge weight distribution matrix, and performing quantile normalization processing on the dynamic edge weight distribution matrix to map it to a probability distribution space; based on the normalized dynamic edge weight distribution matrix, performing adaptive pruning on the edge connections in the entity relationship topology, and retaining the associated edges whose gradient cumulative amount exceeds the average similarity between nodes; performing spatio-temporal alignment on the node enhanced feature matrix and the pruned entity relationship topology, and adjusting the temporal dependence of the node features according to the timestamps of the risk association strength matrix; jointly encoding the spatio-temporally aligned node features and associated edges through a graph structure encoder to generate a graph embedding representation of the dynamic knowledge graph, where the node embedding vector includes a risk propagation potential coefficient and the edge embedding vector includes a risk attenuation factor.
9. The method according to claim 8, characterized in that Identifying potential risk propagation paths in a dynamic knowledge graph includes: extracting risk propagation potential energy coefficients from node attributes in the dynamic knowledge graph and extracting risk attenuation factors from edge attributes; taking nodes with risk propagation potential energy coefficients greater than the average value of the same type of risk in the node attributes as candidate risk source nodes, and generating the reverse propagation path priority according to the reverse order of the risk attenuation factors of the edge attributes; performing breadth-first traversal on the candidate risk source nodes along the reverse propagation path priority, and recording the node sequence passed during the traversal and the risk attenuation factors of the corresponding edges; performing conduction path integrity verification on each node sequence, and when the risk propagation potential energy coefficients of adjacent nodes in the sequence satisfy the exponential attenuation law, calculating the cumulative conduction probability of the conduction path; performing convolution operation on the cumulative conduction probability and the risk propagation potential energy coefficient of the node sequence to generate a path comprehensive risk assessment value; comparing the path comprehensive risk assessment value with the benchmark value of the same type of historical risk cases, and screening out abnormal conduction paths whose assessment values deviate from the standard deviation range of the benchmark value; performing two-way risk diffusion simulation on the abnormal conduction paths to capture risk amplification nodes in the forward propagation path and key inhibition nodes in the reverse blocking path; generating a multi-dimensional risk conduction map including path weights and intervention priorities according to the spatial distribution density of the risk amplification nodes and the key inhibition nodes.
10. A computer system, characterized in that, including: A memory storing a computer program; a processor for loading the computer program to implement the knowledge graph generation method for scientific and technological project risk control according to any one of claims 1-9.
Citation Information
Patent Citations
Global financial risk knowledge graph construction method based on artificial intelligence technology
CN110717816A
Enterprise internal operation management risk identification and extraction method and system based on deep learning
CN112463981A
Granary risk event intelligent early warning and processing method based on ensemble learning and time sequence diagram neural network
CN118569634A
Chemical device question and answer method and system fused with knowledge graph
CN119862280A
Cited By
Enterprise data risk processing method and system based on dynamic knowledge graph
CN120494538A
Enterprise data risk processing method and system based on dynamic knowledge graph
CN120494538B
Electric power engineering design project management system and method
CN120579947A
AI-based security risk analysis method
CN120634819A
Knowledge graph-based risk early warning method and system for medical insurance settlement data
CN120725806A