An industry chain data feature reconstruction method based on an IDPC algorithm

By reconstructing data features across multiple industry chains using the IDPC algorithm, the problems of large differences and scarcity of data features are solved, the expressive power of data and the accuracy of risk prediction are improved, and the safe and stable operation of multiple industry chains is supported.

CN116881702BActive Publication Date: 2025-11-28SOUTHEAST UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310900710.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-21
Publication Date
2025-11-28
Estimated Expiration
2043-07-21

AI Technical Summary

Technical Problem

In a multi-industry chain environment, there are problems of large differences and lack of data characteristics, which increases the complexity of data integration and analysis, and affects the accuracy and reliability of risk management.

Method used

A method for reconstructing industry chain data features based on the IDPC algorithm is adopted. Through behavioral clustering and category reduction, enterprise data is transformed into data with production behavior categories, reducing redundant information and improving the expressiveness and interpretability of the data.

Benefits of technology

It improves the accuracy of data analysis and risk prediction, provides a more comprehensive data foundation, and offers a reliable basis for decision-making and risk management.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116881702B_ABST
    Figure CN116881702B_ABST
Patent Text Reader

Abstract

The application provides an industry chain data feature reconstruction method based on an IDPC algorithm, aiming to strengthen the correlation between data features and industry chain risks and improve the utility of data features in industry chain risk prediction. The method mainly realizes feature reconstruction through two steps of behavior clustering and category reduction. Behavior clustering converts the collected enterprise data into data with production behavior categories, improves the correlation between data and risks from the perspective of explainability, more accurately describes and understands the production behavior of enterprises, and provides a reliable basis for risk analysis. Category reduction reduces the production behavior category data to a feature space with production behavior preferences from the perspective of data effectiveness, realizing data reduction. Data reduction is beneficial to discovering the differences and features between different production behaviors, further improving the utility of data in risk prediction. The reduction based on the feature space with production behavior preferences enables decision makers to more accurately assess and manage the risks of enterprises and between enterprises.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of information technology and relates to an industry chain data feature reconstruction method based on an IDPC algorithm. BACKGROUND

[0002] The current multi-industry chain environment is a complex network of multiple levels and multiple enterprises interacting with each other. Different enterprises and nodes play a key role in this environment, and they exchange and cooperate with each other. This environment covers multiple industries and fields, involving a wide range of product and service supply chains. The characteristics of this multi-industry chain environment are frequent data flow and information exchange, and enterprises depend on each other and work together through data sharing. In the process of such collaborative work, there are often problems of large data feature differentiation and feature scarcity. Large data feature differentiation means that there are significant differences in data collection, recording and processing methods between different enterprises and nodes. First, different industries and fields use different data collection methods, recording formats and standards, resulting in significant differences in the way data features are expressed. This differentiation makes data diverse in terms of dimension, measurement and feature representation, increasing the complexity of data integration and analysis. Second, there is a problem of data feature scarcity in the multi-industry chain environment. This is because enterprises often limit data sharing and openness in the interests of protecting data privacy and commercial interests. Each enterprise has its own data resources, but is subject to confidentiality and competitive advantage constraints, resulting in relatively small amounts of data available for sharing. This data feature scarcity limits the comprehensive understanding and in-depth analysis of the multi-industry chain, and restricts the accuracy and reliability of related decision-making and risk management. Therefore, solving the problems of large data feature differentiation and data feature scarcity is a key challenge in the multi-industry chain environment. By proposing innovative data integration and feature extraction methods, it is possible to promote the effective sharing and integration of data between multiple enterprises and nodes, and thus obtain more comprehensive, accurate and reliable data features, providing a more reliable basis for decision-making and risk management.

[0003] To address the above challenges, the present invention proposes an industry chain data feature reconstruction method based on the IDPC algorithm, aiming to solve the problem of data differentiation and feature scarcity in multiple industry chains. Data feature reconstruction is a method of transforming, extracting, and reorganizing data features to improve the expressiveness and information value of data. This method processes the original data to make it more suitable for analysis and application in specific tasks or fields. By reorganizing and optimizing data features, the data becomes more distinctive, interpretable, and predictive. The method proposed in this invention can fully mine and reconstruct the features of multiple industry chain data from multiple levels and dimensions through behavior clustering and category reduction. Through behavior clustering, we can convert the collected enterprise data into data with production behavior categories, thereby enriching the expressiveness and interpretability of the data. This conversion process can better capture the characteristics and differences of different production behaviors in multiple industry chains, providing a more comprehensive data basis for risk analysis. At the same time, category reduction allows us to reduce the redundancy of data and extract more representative and important features, further improving the accuracy of data analysis and risk prediction.

[0004] The industry chain data feature reconstruction method based on the IDPC algorithm of the present invention has important application value in risk prediction in multiple industry chains. Through this method, we can fully utilize the potential information of multiple industry chain data, improve the utility and value of data, and more accurately assess and manage the risks in multiple industry chains. The application of this technology is expected to provide more comprehensive and accurate decision support for multiple industry chain risk management, promoting the safe and stable operation of the industry chain. SUMMARY

[0005] Technical problem: In view of the problems existing in the prior art, the present invention proposes an industry chain data feature reconstruction method based on the IDPC algorithm to solve the challenge of task dynamics in multiple networked industry chains. The current data features are scarce and cannot be directly applied to industry chain risk prediction, but we realize that the production behavior of enterprises is closely related to the safety of the industry chain, and the production behavior of enterprises directly affects the risk of the industry chain. Therefore, we propose an industry chain data feature reconstruction method based on the IDPC algorithm, which maps the original data to a feature space with production behavior preferences to solve this problem. This method mainly includes two key steps: clustering and reduction.

[0006] Technical Solution: The lack of data features in multiple industrial chains directly impacts risk prediction. Due to factors such as data confidentiality and commercial interests, the data features available for risk analysis are relatively insufficient. This problem is particularly prominent in multiple industrial chains, where complexity and overlap make data collection and analysis more difficult, thus reducing the effectiveness of data in risk prediction models. This invention proposes a method for reconstructing industrial chain data features based on the IDPC algorithm. Through behavioral clustering and category reduction, it mines and reconstructs the features of multiple industrial chain data from multiple levels and dimensions. Behavioral clustering transforms the collected enterprise data into data with production behavior categories, enhancing the data's expressiveness and interpretability. Category reduction reduces the production behavior category data to a feature space with production behavior preferences, reducing redundant information and improving the accuracy of data analysis and risk prediction. The technical solution of this industrial chain data feature reconstruction method based on the IDPC algorithm is as follows:

[0007] 1) Clustering Algorithm: The traditional Density Peaks Clustering (DPC) algorithm requires manual selection of cluster centers. Therefore, this section improves the DPC algorithm and proposes an Improved Density Peaks Clustering (IDPC) algorithm, which can adaptively select cluster centers. The IDPC algorithm can achieve the best clustering effect in industrial chain data and improve the effectiveness of risk prediction.

[0008] First, given a dataset, we need to calculate the Euclidean distance d between any two data points in the dataset. ij And calculate d ij The maximum distance d in the middle max And obtain the distance threshold d c The specific calculation formula is d c =0.2*d max By calculating the Euclidean distance between any two data points in the dataset, a distance matrix can be obtained. This distance matrix can be used as input to behavioral clustering algorithms to measure the similarity or difference between data points, thereby performing cluster analysis and grouping. Based on the distance matrix, clustering algorithms can group data points with similar characteristics into the same category or group, thus achieving the clustering operation on the data. Then, the local density ρ of all data points is calculated. Local density measures the spatial density of each data point; the higher the local density, the greater the probability of it becoming a cluster center. Afterwards, the relative distance δ needs to be calculated. Data point v i relative distance δ i This refers to the minimum distance between the point and other data points with higher local density, or if data point v i If δ is the data point with the highest local density in the dataset, then... iis the maximum distance to all other data points. The relative distance is used to measure the difference between different data categories by using the difference between two production behavior data, so as to find the center point of the category. The data point with larger local density and larger distance is selected as the center point of each transaction behavior category. In order to avoid the interference of the relative distance of a certain data point being large and the local density being small, an auxiliary parameter γ is set here. After normalization, γ i is obtained. i Then γ i is arranged in descending order and recorded as γ i ', which determines the clustering center point through γ i '.

[0009] 2) Reduction algorithm: After clustering, each production data of the enterprise will appear in a certain category, and each category represents a production preference. We reduce the data of these categories. According to the convergence of the data distribution of the same production behavior and the great difference between the data distribution of different production behaviors, we take the category as a new feature, map the original data to a feature space with production behavior preference by counting the number of data of each production behavior. First, use the IDPC algorithm to cluster the data, and divide similar data points into the same cluster, thereby forming a behavior label. Second, data aggregation and analysis are performed on each cluster, and the number of occurrences of each label is recorded to understand the frequency distribution of different behavior labels. Then, the number of behavior labels is assigned to the enterprise, and the statistical data is used to replace the original feature data to generate new feature data, which has the characteristics of reflecting production behavior preference. This method not only extracts and refines the production behavior preference in the original data, but also significantly improves the efficiency and accuracy of the industry chain risk prediction model, thereby providing more reliable risk assessment and prediction results for decision makers.

[0010] Advantages:

[0011] (1) The ability to fully mine and reconstruct data features. Through behavior clustering and category reduction technology, this method can fully mine and reconstruct the features of multiple industry chain data. Behavior clustering converts enterprise data into data with production behavior categories, improving the expressiveness and interpretability of data. This conversion process can better capture the characteristics and differences of different production behaviors in the multiple industry chain, providing a more comprehensive data basis for risk analysis. At the same time, the category reduction technology can reduce the production behavior category data to a feature space with production behavior preference, extracting more representative and important features.

[0012] (2) Improve data representation ability and interpretability. Data transformation is one of the key steps in the method of the present application. Through behavior clustering and category reduction technology, the original data can be reorganized and optimized. This data transformation process can improve the representation ability and interpretability of the data. The transformed data can better reflect the characteristics and behavior patterns in the multi-industrial chain, making the data more distinguishable, interpretable and predictive. This helps decision makers to understand and analyze the data, so as to more accurately assess and manage the risks in the multi-industrial chain.

[0013] (3) Reduce redundant information and improve data analysis and risk prediction accuracy. Feature reduction is another key step in the method of the present application, which can reduce redundant information in the data and extract more representative and important features. By accurately selecting and extracting features, the data dimension and complexity can be reduced, making the data more suitable for analysis and risk prediction. Such feature reduction process can improve the efficiency and accuracy of data analysis, making the related decision and risk management more reliable. BRIEF DESCRIPTION OF DRAWINGS

[0014] Figure 1 An algorithm implementation flowchart of an industrial chain data feature reconstruction method based on the IDPC algorithm. DETAILED DESCRIPTION

[0015] The present application will be further described in detail below in combination with the drawings and specific embodiments:

[0016] Due to the lack of existing data features, they cannot be directly used for industrial chain risk prediction. The present application is based on the fact that the production behavior of enterprises is closely related to the operation safety of the industrial chain, and the production behavior of enterprises is directly linked to the risk of the industrial chain. Therefore, a method of reconstructing the features of the industrial chain data based on the IDPC algorithm is proposed, which maps the original data to a feature space with production behavior preference. The algorithm mainly includes two parts: clustering and reduction, which will be described in detail below.

[0017] 1) Clustering algorithm: The traditional Density Peaks Clustering (DPC) algorithm needs to manually select the cluster center, so this section improves the DPC algorithm and proposes an Improved Density Peaks Clustering algorithm (IDPC) that can adaptively select the cluster center. The IDPC algorithm can achieve the best clustering effect in the industrial chain data and improve the effectiveness of risk prediction. The algorithm flow of adaptively selecting the cluster center in IDPC is as follows:

[0018] (1) Calculate the Euclidean distance d between any two data in the data set ij and calculate the maximum distance d ij in d maxAnd obtain the distance threshold d c By calculating the Euclidean distance between any two data points in a dataset, a distance matrix can be obtained. This distance matrix can be used as input to behavioral clustering algorithms to measure the similarity or difference between data points, thereby performing cluster analysis and grouping. Based on the distance matrix, clustering algorithms can group data points with similar characteristics into the same category or group, thus achieving the clustering operation on the data. The formula for calculating the distance threshold is shown below:

[0019] d c =0.2*d max

[0020] (2) Calculate the local density ρ of all data points. By calculating the local density, we can determine whether the data density in the region where a data point is located is high, thus revealing the clustering pattern and group structure of the data. High-density regions usually represent clusters of data points, corresponding to specific categories or clusters. In addition, calculating the local density is helpful in identifying data points with high local density. Density peaks indicate that there is high data density around them, and they often serve as the center of clusters or important data points. Accurately identifying density peaks helps to divide the clusters of data points, further revealing the internal structure and similarities of the data. The calculation formula is as follows:

[0021]

[0022] (3) Calculate the relative distance δ. Data point v i relative distance δ i This refers to the minimum distance between the point and other data points with higher local density, or if data point v i If δ is the data point with the highest local density in the dataset, then... i This refers to the maximum distance from all other data points. The calculation formula is as follows:

[0023]

[0024] Relative distance uses the difference between two production behavior data to measure the difference between different data categories, thereby finding the center point of the category.

[0025] (4) Select data points with high local density and large distances as the center points for each transaction behavior category. To avoid interference caused by a large relative distance and low local density of a certain data point, set an auxiliary parameter γ according to the following formula:

[0026] γ i =ρ i *δ i

[0027] (5) Normalize γ to obtain γ'. i Then γ' i Arranged in descending order, denoted as γ' i The normalization calculation formula is shown below:

[0028]

[0029] (6) via γ' i To determine the cluster centers, we define the first-order difference of the auxiliary parameters according to the following formula:

[0030]

[0031] The second difference of the auxiliary parameter is defined by the following formula:

[0032]

[0033] Based on the above definition, we find that cluster centers must meet the following conditions:

[0034]

[0035] non-central point and central point γ i The value jumps, that is, the value of γ at the center point... i The value is significantly greater than that of non-center points, and the γ value of the center point is also significantly greater. i The value changes are very chaotic, with γ' at non-central points. i The value changes approximately linearly. Let d be the last cluster center. i The first non-center point is d. i ' +1 Then the cluster center point d i 'To non-center point d i ' +1 This is the last place where the first-order difference changes significantly, i.e., α < -1; the subsequent first-order differences remain almost unchanged.

[0036] That is, α > -1. This is reflected in the piecewise linear graph of the second-order difference, v i This is the last place where the change is significant, i.e., β > 1. The subsequent second-order difference remains almost unchanged, i.e., β < 1.

[0037] 2) Reduction Algorithm: After clustering, each piece of production data from a company will appear in a certain category, and each category represents a production preference. We reduce the data in these categories. Since the data distribution of the same production behavior tends to converge, while the data distribution of different production behaviors differs greatly, we use the category as a new feature and map the original data to a feature space with production behavior preferences by counting the number of data points in each category. The specific steps of the data reduction method after clustering of industry chain data are as follows:

[0038] (1) Cluster division: First, according to the clustering results of the IDPC algorithm, the data to be clustered is divided, assuming that the data is divided into K clusters, where each cluster represents a specific behavior label. This process is based on the calculation of the Euclidean distance between data points and the determination of the core points of each cluster based on local density; this division process is based on the comprehensive consideration of the Euclidean distance calculation and local density between data points; specifically, the similarity and distance relationship between data points is measured by calculating the Euclidean distance between them, and then the core points of each cluster are identified according to the local density; through this division process, data points can be organized into different clusters, so that each cluster can represent a specific behavior label, thereby laying the foundation for subsequent data reduction and feature reconstruction;

[0039] (2) Data aggregation and label statistics: Second, the data in each cluster is aggregated and counted. This includes collecting all data samples in the cluster and recording the number of occurrences of each behavior label in order. By counting the frequency of each label, we can understand the importance and data distribution of each behavior label; this step covers collecting all data samples in the cluster and recording the number of occurrences of each behavior label in order; by counting the frequency of each label, we can better understand the importance and data distribution of each behavior label; this statistical analysis helps to reveal the relative importance of different behavior labels and the imbalance of data distribution; by recording the number of occurrences of each behavior label in detail, we can obtain the basis for further analysis and interpretation of data samples, providing valuable reference for subsequent data reduction and feature reconstruction;

[0040] (3) Label results are given to the enterprise: Next, the statistical results of the behavior labels are sequentially assigned to the enterprise. This means that each behavior label and its corresponding number of occurrences are associated with the enterprise and recorded and organized in a certain order. In this way, the enterprise can obtain behavior label information related to it, providing a basis for subsequent feature reconstruction and data analysis;

[0041] (4) Feature data replacement and reconstruction: Finally, by replacing the original feature data of the enterprise, new feature data with production behavior preferences is generated. This means that the original feature data of the enterprise is replaced by new feature data generated based on the statistical results of the behavior labels. The new feature data can better reflect the production behavior preferences of the enterprise, providing more accurate and valuable information for further data analysis and risk prediction.

[0042] Through the above four steps, the industry chain data specification method based on the IDPC algorithm can divide and aggregate data according to behavior tags, providing more accurate and useful feature data for enterprises, thereby realizing data specification and reconstruction. The application of this method can improve the accuracy and reliability of data analysis and risk prediction models, helping enterprises better understand and manage risks in the industry chain.

Claims

1. An industry chain data feature reconstruction method based on an IDPC algorithm, characterized in that, The method comprises two steps of clustering and reduction. Firstly, the DPC algorithm is improved to propose an IDPC algorithm for adaptive selection of clustering centers. After clustering, each production data of the enterprise will appear in a certain category, and each category represents a production preference. Then, the data of these categories are reduced. According to the convergence of the data distribution of the same production behavior and the great difference between the data distribution of different production behaviors, the original data are mapped into a feature space with production behavior preferences by taking the category as a new feature and by counting the number of data of each production behavior. The algorithm for selecting the cluster center is as follows: for a given data set, the Euclidean distance d ij between any two data in the data set is calculated ij , the maximum distance d max in d ij is calculated, and a distance threshold d c is obtained, and the specific calculation formula is d c =0.2*d max ; a distance matrix is obtained by calculating the Euclidean distance between any two data in the data set; the distance matrix is used as the input of the behavior clustering algorithm to measure the similarity or difference between the data points, so as to perform clustering analysis and group division; Based on the distance matrix, the clustering algorithm classifies data points with similar features into the same category or group, thereby realizing the clustering operation of the data. Then, the local density ρ of all data points is calculated. By calculating the local density, it is found whether the data density of the region where the data point is located is high, thereby revealing the clustering mode and group structure of the data. The high-density region represents the clustering of the data points, corresponding to a specific category or cluster. In addition, by calculating the local density, data points with high local density can be identified. The density peak point represents a high data density around it, serving as the center or important data point of clustering. Accurate identification of the density peak point helps to divide the clustering clusters of the data points and further reveals the internal structure and similarity of the data. The calculation formula is as follows: The relative distance δ is calculated. Data point v i relative distance δ i This refers to the minimum distance between the point and other data points with higher local density, or if data point v i If δ is the data point with the highest local density in the dataset, then... i This refers to the maximum distance from all other data points; the calculation formula is shown below: Relative distance is used to measure the difference between two production behavior data points, thereby finding the center point of the classification. Based on previous calculations, data points with high local density and large distances are selected as the center points of each transaction behavior category. To avoid interference from a data point with a large relative distance but low local density, the formula γ is used. i =ρ i *δ i Set an auxiliary parameter γ; normalize γ to obtain γ'. i Then γ' i Arranged in descending order, denoted as γ' i The normalization calculation formula is shown below: by γ" i determining cluster center points, defining a first order difference of the auxiliary parameter according to the following equation: The second-order difference of the auxiliary parameter is defined according to the following formula: According to the above definition, the clustering center needs to satisfy the following conditions: gamma of non-center points and center points i There is a jump in the value of gamma of center points i which is significantly larger than that of non-center points, and the value of gamma of center points i varies very chaotically, and the value of gamma of non-center points i changes nearly linearly; let the last cluster center point be d i , the first non-center point be d i+1 , then the cluster center point d i to the non-center point d i+1 is the last place where the first-order difference changes significantly, i.e., alpha < -1, and the subsequent first-order difference is nearly constant, i.e., alpha > -1; this is reflected in the broken line graph of the second-order difference, v i is the last place where the change is significant, i.e., beta > 1, and the subsequent second-order difference is nearly constant, i.e., beta < 1; The specific steps of the data reduction method after the clustering of the industrial chain data include clustering cluster division, data summarization and label statistics, label result assignment to enterprises and feature data replacement and reconstruction steps. According to the statistical results, the statistical results of the behavior labels are sequentially assigned to the enterprises. Each behavior label and its corresponding occurrence number are associated with the enterprise and recorded and organized in a certain order. In this way, the enterprise can obtain the behavior label information related to it, providing a basis for subsequent feature reconstruction and data analysis. By replacing the original feature data of the enterprise, new feature data with production behavior preferences are generated. The original feature data of the enterprise is replaced by new feature data generated based on the behavior label statistical results. The new feature data can better reflect the production behavior preferences of the enterprise and provide more accurate and valuable information for further data analysis and risk prediction.

2. The industry chain data feature reconstruction method based on the IDPC algorithm according to claim 1, characterized in that: According to the clustering results of the IDPC algorithm, the data to be clustered is divided; the data is divided into K clustering clusters, each of which represents a specific behavior label. This division process is based on the comprehensive consideration of the Euclidean distance calculation and local density between data points. Specifically, the similarity and distance relationship between data points are measured by calculating the Euclidean distance between them, and then the core points of each clustering cluster are identified according to the local density. Through this division process, data points can be organized into different clustering clusters, so that each clustering cluster can represent a specific behavior label, thereby laying a foundation for subsequent data reduction and feature reconstruction.

3. The industry chain data feature reconstruction method based on the IDPC algorithm according to claim 2, characterized in that: Based on the division result, the data in each cluster is summarized and counted; This step covers collecting all data samples in the cluster and recording the number of occurrences of each behavior label in order; By counting the frequency of each label, the importance and data distribution of each behavior label can be better understood; this statistical analysis helps to reveal the relative importance of different behavior labels and the imbalance of data distribution; By recording the number of occurrences of each behavior label in detail, the basis for further analysis and interpretation of data samples can be obtained, providing valuable reference for subsequent data reduction and feature reconstruction.

Citation Information

Patent Citations

  • Intelligent enterprise risk identification method based on financial big data

    CN115293641A

  • KR1025149930000B1