A method and system for enterprise digital management based on data mining
By building a dynamic network diagram within the enterprise, identifying key nodes and high-risk nodes, the problem of difficult to capture dynamic changes in the implicit dependencies in the existing technology is solved, and flexible and efficient optimization of resource allocation is achieved.
Patent Information
- Application Number
- CN202510742682.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-05
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-06-05
AI Technical Summary
The existing technology cannot accurately identify and analyze the dynamic changes in the implicit dependencies within the enterprise, resulting in inaccurate positioning of high-risk nodes and key propagation paths, affecting resource optimization.
By obtaining information flow data and resource allocation data between departments, building an initial network diagram, performing dependency relationship analysis, combining clustering analysis, path optimization and nonlinear feature extraction, key nodes and high-risk nodes are identified, and key locations for resource bottlenecks and conflicts occur.
Real-time updates and dynamic adjustments of internal dependencies of enterprises are realized, accurately identifying high-risk nodes and key propagation paths, optimizing resource allocation, and improving the company's ability to adapt to complex environments.
Smart Images

Figure CN120258481B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of enterprise digital management, and in particular to a method and system for enterprise digital management based on data mining. Background Art
[0002] Currently, traditional approaches to digital enterprise management rely on static organizational charts, business process management (BPM), or rule-based resource allocation schemes to describe interdepartmental collaboration. However, these approaches fail to accurately reflect the dynamic flow of internal resources and struggle to adapt to the rapid changes in the enterprise environment.
[0003] One existing technique uses graph theory analysis or social network analysis (SNA) to model internal enterprise dependencies. For example, a network diagram is constructed to describe resource allocation patterns between departments, and centrality metrics are used to analyze key nodes. However, this technique is primarily based on static network models, making it difficult to capture the dynamic changes in internal enterprise dependencies, resulting in low real-time and adaptable analysis results.
[0004] Existing technologies are unable to accurately identify and analyze the dynamic changes of implicit dependencies within an enterprise, resulting in inaccurate positioning of high-risk nodes and key transmission paths, affecting the enterprise's resource optimization. Summary of the Invention
[0005] The present invention provides a method and system for digital enterprise management based on data mining, which can identify the dynamic changes of implicit enterprise dependencies, locate high-risk nodes and key transmission paths, and optimize resource allocation.
[0006] In a first aspect, in order to solve the above technical problems, the present invention provides a method for digital enterprise management based on data mining, comprising:
[0007] Obtain data on information flow and resource allocation between departments;
[0008] Perform dependency relationship analysis based on the information flow data and the resource allocation data to obtain an initial network diagram of inter-departmental dependency relationships;
[0009] Based on the initial network diagram, cluster analysis is performed to analyze connection strength and path length, and the network is dynamically reconstructed and community mining and path optimization are integrated to identify key nodes, thereby obtaining a dynamic network diagram of information flow between departments;
[0010] Extracting interactive nonlinear features of the dynamic network graph to obtain predictive analysis data for determining node influence;
[0011] Identify nodes in the network that affect resource allocation and information transmission based on the dynamic network diagram and the prediction analysis data, and obtain key nodes and high-risk nodes;
[0012] Identify key communication paths in the network based on the key nodes and the high-risk nodes, and obtain key communication paths that affect resource allocation and information transmission;
[0013] Based on the key nodes, the high-risk nodes and the key propagation paths, key locations where resource bottlenecks and resource conflicts occur are predicted, and resource optimization suggestions are generated.
[0014] Preferably, the dependency relationship analysis is performed based on the information flow data and the resource allocation data to obtain an initial network diagram of the inter-department dependency relationship, including:
[0015] Calculating support and confidence between departments based on the information flow data and the resource allocation data;
[0016] When the support and the confidence are respectively greater than a preset support threshold and a preset confidence threshold, a network structure diagram of the dependency relationship between departments is established based on a graph theory method to obtain an initial network diagram of the dependency relationship between departments.
[0017] Preferably, the method of performing cluster analysis on the connection strength and path length based on the initial network diagram, dynamically reconstructing the network and integrating community mining and path optimization to identify key nodes to obtain a dynamic network diagram of information flow between departments includes:
[0018] Extracting connection strength data and path length data from the initial network graph;
[0019] Perform resource allocation pattern analysis based on the connection strength data and the path length data to obtain the main transmission paths of inter-departmental dependency relationships;
[0020] Perform dependency trend analysis on the main transmission paths to obtain dependency trends among departments;
[0021] When the dependency trend is greater than a preset trend range, the network structure of the initial network diagram is updated based on a graph theory method to obtain an updated network diagram of the dependency relationship between departments;
[0022] Identify key nodes for resource allocation based on the updated network graph to obtain priority nodes that require priority resource allocation;
[0023] Performing path weight analysis on the priority nodes to obtain weight factors reflecting the importance of information transmission between departments;
[0024] According to the weight factors, the dependency relationship between departments is updated to obtain a dynamic network diagram of information flow between departments.
[0025] Preferably, the extracting of interactive nonlinear features of the dynamic network graph to obtain predictive analysis data for determining node influence includes:
[0026] Extracting node connection relationships, weight data, and time series data from the dynamic network graph;
[0027] Calculating the influence index of each node based on the node connection relationship and the weight data;
[0028] When the influence index is greater than a preset influence threshold, the node corresponding to the influence index is determined as a main node with an amplification effect;
[0029] When the influence index is less than a preset influence threshold, the node corresponding to the influence index is determined as a secondary node with a weakening effect;
[0030] Based on the time series data, the evolution trend analysis of the main nodes and the secondary nodes is performed to obtain predictive analysis data of the evolution trend of the node influence.
[0031] Preferably, identifying nodes in the network that affect resource allocation and information transmission based on the dynamic network diagram and the predictive analysis data to obtain key nodes and high-risk nodes includes:
[0032] Extracting node connection data, shortest path data, network topology change trends, clustering coefficients, and community affiliation data from the dynamic network graph;
[0033] Based on the prediction and analysis data, combined with the node connection data, the shortest path data and the network topology change trend, a dynamic weight operation of the centrality index is modified to obtain a centrality index for evaluating the importance of the node in information dissemination;
[0034] Based on the prediction analysis data, combined with the clustering coefficient and the community belonging data, the dynamic weight of the vulnerability index is modified to obtain a vulnerability index for evaluating the degree of dependence of the node in information dissemination;
[0035] According to the centrality index, constructing each node feature vector to obtain a centrality feature vector;
[0036] The K-means clustering algorithm is used to perform importance ranking and screening on the centrality feature vectors to obtain preliminary key nodes of high importance, medium importance and low importance;
[0037] Ranking the importance of the preliminary key nodes to obtain key nodes with high centrality;
[0038] Analyze the connection characteristics and load capacity indicators of each node according to the vulnerability indicators to obtain a vulnerability assessment value;
[0039] When the vulnerability assessment value is less than a preset vulnerability threshold, determining the node corresponding to the vulnerability assessment value as a potential high-risk node;
[0040] Using a support vector machine algorithm to predict node failure of the potential high-risk nodes to obtain the node failure probability;
[0041] A node set whose node failure probability is greater than a preset probability threshold is extracted, and all nodes in the node set are determined as high-risk nodes.
[0042] Preferably, identifying key propagation paths in the network based on the key nodes and the high-risk nodes to obtain the key propagation paths that affect resource allocation and information transmission includes:
[0043] Constructing a risk network diagram using graph theory methods based on the key nodes and the high-risk nodes;
[0044] Calculating path parameters of the risk network graph to obtain propagation data and stability data of nodes on the path;
[0045] Evaluate node importance and path stability based on the propagation data and stability data to obtain a path evaluation result;
[0046] The path evaluation results are grouped according to path propagation priority, resource influence, and decision chain weight to obtain key propagation paths that affect resource allocation and information transmission.
[0047] Preferably, the predicting of key locations where resource bottlenecks and resource conflicts occur based on the key nodes, the high-risk nodes, and the key propagation paths, and generating resource optimization suggestions, includes:
[0048] Analyze resource allocation status according to the key nodes, the high-risk nodes, and the key propagation paths to obtain resource usage of key nodes, resource pressure of high-risk nodes, and traffic load of key propagation paths;
[0049] Predicting key locations where resource bottlenecks and resource conflicts occur based on the resource usage, resource pressure, and traffic load, combined with preset historical resource scheduling data;
[0050] Adjust resource allocation strategies for the key locations and generate resource optimization suggestions.
[0051] In a second aspect, the present invention provides an enterprise digital management system based on data mining, comprising:
[0052] Data acquisition module, used to obtain information flow data and resource allocation data between departments;
[0053] A dependency analysis module, configured to perform dependency relationship analysis based on the information flow data and the resource allocation data to obtain an initial network diagram of inter-departmental dependency relationships;
[0054] A path optimization module is used to perform cluster analysis on the connection strength and path length based on the initial network diagram, dynamically reconstruct the network, and integrate community mining and path optimization to identify key nodes, thereby obtaining a dynamic network diagram of information flow between departments;
[0055] A prediction analysis module is used to extract interactive nonlinear features of the dynamic network graph to obtain prediction analysis data for determining node influence;
[0056] A node identification module is used to identify nodes in the network that affect resource allocation and information transmission based on the dynamic network diagram and the prediction analysis data, and obtain key nodes and high-risk nodes;
[0057] A critical path module is used to identify the critical propagation paths in the network based on the key nodes and the high-risk nodes, and obtain the critical propagation paths that affect resource allocation and information transmission;
[0058] The resource optimization module is used to predict the key locations where resource bottlenecks and resource conflicts occur based on the key nodes, the high-risk nodes and the key propagation paths, and generate resource optimization suggestions.
[0059] In a third aspect, the present invention also provides an electronic device comprising a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, it implements any one of the above-mentioned data mining-based enterprise digital management methods.
[0060] In a fourth aspect, the present invention also provides a computer-readable storage medium, which includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute any one of the above-mentioned data mining-based enterprise digital management methods.
[0061] Compared with the prior art, the present invention has the following beneficial effects:
[0062] (1) The present invention provides a data mining-based enterprise digital management method. By acquiring information flow data and resource allocation data between enterprise departments, an initial network diagram is constructed. Combined with time series analysis and path optimization algorithms, the enterprise dependency network is dynamically adjusted. Existing technologies mainly rely on static analysis methods, which are difficult to accurately reflect the dynamic evolution of internal enterprise dependencies, resulting in delayed resource allocation decisions. The present invention can update the network structure in real time and identify the changing trends of information flow and resource flow within the enterprise.
[0063] (2) The present invention uses a nonlinear regression analysis method, combined with centrality indicators and vulnerability indicators, to identify high-risk nodes and key transmission paths. Existing technologies are based on linear analysis and have difficulty capturing the nonlinear changes in the influence of certain nodes in different business scenarios. The present invention calculates the node influence index and combines it with the support vector machine algorithm to predict the probability of node failure, accurately identifying nodes that have a greater impact on enterprise stability. At the same time, based on the path analysis algorithm, it discovers key transmission paths and reveals the core links within the enterprise that lead to business interruptions or resource conflicts.
[0064] (3) Based on path analysis and machine learning algorithms, the present invention predicts resource bottlenecks and resource conflicts and generates optimization recommendations. Compared to existing resource management methods that rely on fixed rules, the present invention dynamically calculates resource usage, traffic load, and critical path dependencies, and adjusts resource allocation strategies in combination with optimization algorithms, making resource scheduling more flexible and efficient. At the same time, the optimized dynamic network diagram can be integrated with the enterprise resource management system to achieve real-time monitoring of inter-departmental dependencies and resource flow status, improve the enterprise's adaptability in complex business environments, and optimize resource allocation plans. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 This is a flow chart of the enterprise digital management method based on data mining provided by the first embodiment of the present invention;
[0066] Figure 2 This is a structural diagram of an enterprise digital management system based on data mining provided by the second embodiment of the present invention. DETAILED DESCRIPTION
[0067] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0068] Reference Figure 1The first embodiment of the present invention provides a method for digital management of an enterprise based on data mining, comprising the following steps:
[0069] S11, obtain information flow data and resource allocation data between departments;
[0070] S12, performing dependency relationship analysis based on the information flow data and the resource allocation data to obtain an initial network diagram of inter-departmental dependency relationships;
[0071] S13, performing cluster analysis on the connection strength and path length based on the initial network diagram, dynamically reconstructing the network and integrating community mining and path optimization to identify key nodes, thereby obtaining a dynamic network diagram of information flow between departments;
[0072] S14, extracting interactive nonlinear features of the dynamic network graph to obtain prediction analysis data for determining node influence;
[0073] S15, identifying nodes in the network that affect resource allocation and information transmission based on the dynamic network diagram and the prediction analysis data, and obtaining key nodes and high-risk nodes;
[0074] S16, identifying key communication paths in the network based on the key nodes and the high-risk nodes, and obtaining key communication paths that affect resource allocation and information transmission;
[0075] S17: Based on the key nodes, the high-risk nodes, and the key propagation paths, predict key locations where resource bottlenecks and resource conflicts occur, and generate resource optimization suggestions.
[0076] In step S11, it is necessary to obtain information flow data and resource allocation data between departments, including:
[0077] In one specific embodiment, information flow data primarily originates from an enterprise's internal office systems, communication systems, business management systems, and data exchange platforms. For example, the document management system (DMS), enterprise email system (email), and instant messaging tools (such as WeChat for Business and DingTalk) within the office system can provide data such as file sharing, email exchanges, and communication records. Communication system data includes conference call records, video conference logs, internal enterprise forums, and work logs, reflecting the frequency of communication and the content of information exchange between different departments. Furthermore, an enterprise's business management systems (such as ERP, CRM, and SCM) can record cross-departmental business collaboration relationships, such as order flow, customer demand feedback, and production scheduling. Data exchange platforms (such as enterprise data sharing platforms and API logs) can provide information on the frequency of data calls and data sharing between different systems, further revealing information dependencies between departments.
[0078] In a specific embodiment, resource allocation data mainly involves aspects such as human resources, capital flow, equipment usage, and production resource allocation within the enterprise. Among them, human resource data can be obtained through the enterprise's human resource management system (HRM), including personnel distribution in various departments, job transfers, and working time allocation for cross-departmental collaboration. The financial management system (FMS) provides capital flow information, such as budget allocation, project costs, and cross-departmental capital transactions. The asset management system (AMS) is used to record the use of enterprise equipment, office resources, and production tools, and track the sharing and allocation of equipment between different departments. The production and supply chain management system (SCM) reflects inventory flow, production material scheduling, and logistics distribution, ensuring that resource allocation data comprehensively covers all aspects of enterprise operations.
[0079] Specifically, in order to ensure the integrity and timeliness of the data, the present invention adopts a variety of data collection methods, including system log analysis, database query, API data call and manual input review. System log analysis is mainly aimed at the data of office systems and communication systems, and automatically extracts email logs, file sharing records and meeting communication data. For example, by parsing the mail server log, cross-departmental email exchanges can be obtained, and the content relevance can be analyzed in combination with the email subject and keywords. The database query method is suitable for business management systems, and through SQL queries, internal enterprise resource flow data, such as purchase order approval processes and fund payment records, can be directly extracted from ERP, CRM and other databases. The API data call method is suitable for real-time data collection from different business systems, such as calling the interface between ERP and SCM systems to obtain inventory change information and material allocation status. In addition, for some unstructured data, such as inter-departmental resource coordination meeting records, data integrity can be ensured through manual review and re-recording.
[0080] It should be noted that since the collected data comes from various sources and in different formats, the present invention uses data cleaning, normalization and format conversion methods for pre-processing. Data cleaning mainly removes duplicate, invalid and abnormal data, such as eliminating automatic reply emails to avoid interfering with information flow analysis. Data normalization ensures that data from different sources have a unified measurement standard. For example, the working hour data recorded in different systems is converted into standard hour units to facilitate cross-departmental comparison. Format conversion is used to convert unstructured data (such as meeting minutes, text emails) into structured data table storage. For example, natural language processing (NLP) technology is used to extract keywords from email content and associate them with corresponding business processes.
[0081] Taking a manufacturing company as an example, the dependency between its R&D department and production department is mainly reflected in the interaction of technical documents, the sharing of experimental equipment, and the arrangement of testers. The present invention can extract the design documents uploaded by the R&D department through the DMS system, analyze their flow paths between different departments, and identify the dissemination pattern of technical knowledge. At the same time, the use log of experimental equipment is obtained through the asset management system (AMS), and the sharing of equipment between the R&D and production departments is tracked to analyze the cross-departmental flow of resources. The human resource management system (HRM) provides work scheduling data for testers, analyzes the distribution of testers' working hours in different projects, and determines the dependency of cross-departmental human resources. In the data preprocessing stage, all data are normalized according to a unified time dimension to facilitate subsequent time series analysis and trend prediction.
[0082] In step S12, dependency analysis needs to be performed based on the information flow data and the resource allocation data to obtain an initial network diagram of inter-departmental dependencies, including:
[0083] Calculating support and confidence between departments based on the information flow data and the resource allocation data;
[0084] When the support and the confidence are respectively greater than a preset support threshold and a preset confidence threshold, a network structure diagram of the dependency relationship between departments is established based on a graph theory method to obtain an initial network diagram of the dependency relationship between departments.
[0085] First, in order to ensure the accuracy of dependency analysis, the present invention performs data fusion on information flow data and resource allocation data. The main purpose of data fusion is to integrate multiple data sources, make up for the shortcomings of a single data source, and improve the integrity and consistency of the data. Data fusion adopts the methods of time alignment, entity matching and similarity analysis. For example, for the data of the office system, business management system and human resources system, they are first aligned according to the timestamp to ensure that the records of different systems can be analyzed according to the same time window. Secondly, the relevant data is matched based on the same project number, order number or equipment number to ensure that the information in the same business process can be integrated together. In addition, for unstructured data in emails, documents or business logs, natural language processing technology is used to calculate text similarity to identify different data entries belonging to the same business process.
[0086] After the data fusion is completed, the present invention uses an association rule mining algorithm to calculate the support and confidence between departments.
[0087] After completing the data fusion, the present invention uses an association rule mining algorithm to analyze the department collaboration pattern. Specifically, the present invention uses a frequent item set mining algorithm to extract stable department collaboration relationships from a large amount of business data. The method first constructs a transaction data set, in which each transaction represents a complete business interaction, and all departments in the transaction constitute an item set. For example, in a software development company, a complete product development process involves the product department, R&D department, testing department, and operations department, and the business process can be expressed as {product, R&D, testing, operations}. In a manufacturing company, a complete production task involves R&D, production, quality management, and supply chain, and the transaction can be expressed as {R&D, production, quality management, supply chain}. By constructing a large amount of business transaction data, the present invention can analyze the collaboration patterns of various departments in different business scenarios through an association rule mining algorithm.
[0088] Specifically, the present invention first performs frequent item set mining, that is, statistics the frequency of collaboration between departments within the enterprise and screens out department combinations with stable dependencies. The main steps of the algorithm are as follows:
[0089] The frequency of occurrence of each individual department is counted, the support is calculated, and the departments with support values above the threshold are selected to form frequent item sets.
[0090] Based on frequent itemsets, candidate department pairs are generated, and the number of their co-occurrences is calculated to screen out department pairs whose support is higher than the support threshold.
[0091] Further expand to combinations of multiple departments and calculate more complex dependencies until no new high-frequency collaborative department combinations can be found.
[0092] For example, a transaction data set of an enterprise includes the following business records: {R&D, production, quality management}, {production, supply chain, procurement}, {R&D, production, supply chain}, and {production, supply chain}.
[0093] After scanning the data set, calculate the support of each department pair, that is, the frequency of two departments appearing together in the same transaction. For example:
[0094] {R&D, Production} appear together 3 times, with a total of 4 transactions, and support = 3 / 4 = 0.75;
[0095] {production, supply chain} appear together 3 times, with a total of 4 transactions, and support = 3 / 4 = 0.75;
[0096] {R&D, Supply Chain} appear once together, with a total of 4 transactions, and support = 1 / 4 = 0.25;
[0097] If the support threshold set by the enterprise is 0.5, {R&D, Supply Chain} is screened out, while {R&D, Production} and {Production, Supply Chain} are retained.
[0098] Specifically, after obtaining the frequent item sets, the present invention calculates the confidence to measure the strength of the dependency relationship between departments. The confidence is calculated as follows:
[0099] Confidence = (number of co-occurrences of Department A and Department B) / number of occurrences of Department A
[0100] For example, the support from production to supply chain is 0.75, and production appears 4 times in total, so the confidence is 0.75; the support from R&D to production is 0.75, and R&D appears 3 times in total, so the confidence is 1.0.
[0101] Specifically, the present invention sets a confidence threshold, for example 0.6, that is, if department A occurs, there is at least a 60% probability that department B will also occur, then it is determined that A has a stable dependence on B. For example, the confidence from production to supply chain is 0.75, which meets the threshold, while the confidence from R&D to supply chain is 0.25, which is lower than 0.6, and is therefore screened out. When the calculated support and confidence are both higher than the set threshold, the present invention uses graph theory methods to establish the initial dependency network structure diagram of the enterprise. In this network, the nodes represent the various departments in the enterprise; the connecting lines represent the dependency relationship between departments, and the direction is from the dependent party to the dependent party; the weight of the connecting line is calculated comprehensively by the support and confidence, and the calculation method is as follows:
[0102] Dependency weight = 0.5 × support + 0.5 × confidence
[0103] For example, if the support of production to supply chain is 0.75 and the confidence is 0.75, then the weight of this dependency is 0.75.
[0104] Ultimately, through data fusion and association rule mining, the present invention can accurately construct the initial dependency network diagram within the enterprise, clearly display the information flow and resource allocation relationship between departments, and provide reliable data support for subsequent dynamic analysis, key node identification and resource optimization.
[0105] It should be noted that during implementation at a manufacturing enterprise, due to the stability of production processes and the long-standing existence of cross-departmental collaboration, the company sought to identify only dependencies with strong business relevance to optimize production scheduling and supply chain management. Therefore, when setting the support threshold, the company first compiled historical interaction data across all departments and found that, on average, each pair of departments interacted approximately 3% of the total interactions. Given the high stability of business flows in the manufacturing industry, the company decided to adopt a higher support threshold of 0.05, meaning only pairs of departments with interactions exceeding 5% were considered to have a dependency relationship. For example, in the past year, the R&D and Production departments collaborated 700 times, accounting for 7% of the total interactions and meeting the threshold. However, the R&D and Procurement departments interacted only 250 times, accounting for 2.5%, which was below the threshold. Therefore, the two were not considered to have a stable dependency relationship. When setting the confidence threshold, the manufacturing company focused on the closeness of collaboration within core business processes. Statistics revealed that the average collaboration confidence level between core departments was 0.6. Therefore, the company set a confidence threshold of 0.6, meaning that a department's behavior is considered stable only when it is at least 60% dependent on another department. For example, the collaboration confidence between the production and supply chain departments was 0.72, meeting the dependency screening criteria. However, the collaboration confidence between the R&D and procurement departments was only 0.4, below the set threshold, and therefore not included in the initial network diagram. Ultimately, the initial network diagram constructed by the manufacturing company included stable dependencies between R&D, production, and supply chain, providing data support for subsequent resource optimization and risk assessment.
[0106] For example, during implementation at an internet company, due to their flexible business models and dynamic interdepartmental collaboration, they sought to identify more potential dependencies to optimize product development and marketing strategies. Therefore, when setting the support threshold, the company analyzed all interdepartmental task allocation records from the past year and found that, on average, each pair of departments interacted for 2% of the total. Given the high frequency of collaboration but looser dependencies within internet companies, the company decided to use a lower support threshold of 0.02, requiring only department pairs with at least 2% of interactions to be included in the dependency network. For example, the product and R&D departments collaborated on 500 tasks, accounting for 10% and meeting the threshold. However, the marketing and R&D departments only interacted 80 times, accounting for 1.6%, falling below the threshold and therefore not included in the initial network diagram. Regarding the confidence threshold, internet companies have relatively flexible collaboration models, with the average collaboration confidence between core departments being approximately 0.5. Therefore, the company set a confidence threshold of 0.5, defining a stable dependency only when a department's actions are at least 50% dependent on another. For example, the collaboration confidence between the R&D department and the testing department is 0.7, which meets the dependency screening criteria, while the collaboration confidence between the marketing department and the R&D department is only 0.3, which is below the threshold and is therefore not included in the initial network diagram.
[0107] In step S13, it is necessary to perform cluster analysis on the connection strength and path length based on the initial network diagram, dynamically reconstruct the network, and integrate community mining and path optimization to identify key nodes to obtain a dynamic network diagram of information flow between departments, including:
[0108] Extracting connection strength data and path length data from the initial network graph;
[0109] Perform resource allocation pattern analysis based on the connection strength data and the path length data to obtain the main transmission paths of inter-departmental dependency relationships;
[0110] Perform dependency trend analysis on the main transmission paths to obtain dependency trends among departments;
[0111] When the dependency trend is greater than a preset trend range, the network structure of the initial network diagram is updated based on a graph theory method to obtain an updated network diagram of the dependency relationship between departments;
[0112] Identify key nodes for resource allocation based on the updated network graph to obtain priority nodes that require priority resource allocation;
[0113] Performing path weight analysis on the priority nodes to obtain weight factors reflecting the importance of information transmission between departments;
[0114] According to the weight factors, the dependency relationship between departments is updated to obtain a dynamic network diagram of information flow between departments.
[0115] First, we extract the connection strength data and path length data from the initial network graph. The connection strength represents the frequency of information exchange or the degree of resource sharing between different departments in the enterprise, such as the number of email exchanges, document sharing, and task handovers between two departments. The connection strength can be expressed as a weight matrix ,in Representative Department To the department The interaction frequency between them is calculated as follows:
[0116]
[0117] For example, in a manufacturing company, the R&D department sends 20 technical reports to the production department weekly, while the total number of reports between all departments within the company is 200. Therefore, the connection strength between the R&D and production departments is 0.1. Path length data measures the shortest number of steps required for information to propagate from one department to another, that is, the minimum number of hops required for information to travel within the enterprise network. This invention uses the Dijkstra shortest path algorithm or the Floyd-Warshall algorithm to calculate the shortest paths between departments to analyze the optimal transmission paths for information flow within the enterprise. For example, in an enterprise network, the marketing department (M) needs to transmit customer feedback to the R&D department (R). However, the marketing department cannot communicate directly with R&D and must instead go through the product department (P). The path can be represented as M→P→R, with a path length of 2, indicating that the information transfer requires two steps. If the marketing department can also reach the R&D department via the operations department (O), then the path M→O→R exists. The system compares the lengths of the two paths and selects the shortest one. For example, in the Dijkstra algorithm, if the edge weight of M→P is 1, the edge weight of P→R is 1, and the edge weight of M→O is 2, and the edge weight of O→R is 1, then the shortest path is still M→P→R, with a path length of 2, while the length of M→O→R is 3. Therefore, the marketing department should prioritize transmitting customer feedback to the R&D department through the product department to ensure efficient information flow. For more complex multi-department collaboration situations, the present invention uses the Floyd-Warshall algorithm to calculate the shortest path matrix between all departments, ensuring optimized information flow within the enterprise network and improving business collaboration efficiency.
[0118] Specifically, after obtaining connection strength and path length, the present invention uses a clustering algorithm to analyze resource allocation patterns on the data to identify key transfer paths between departments. Specifically, a K-Means clustering algorithm is used to group similar departments together to identify groups of departments with close business interactions. For example, in a certain software company, the R&D, Testing, and Operations departments have high connection strengths and can be classified into the same business group. Marketing, Sales, and Customer Support departments can also be classified into another group due to their frequent interactions. This clustering analysis helps identify critical paths for information dissemination within the enterprise and ensures that resource allocation prioritizes high-frequency business needs.
[0119] The present invention further conducts dependency trend analysis on the identified main transmission paths based on the time series analysis algorithm to predict the changing trend of inter-departmental dependency relationships. The time series analysis uses a sliding window method to calculate the department interaction change rate over a period of time. :
[0120]
[0121] in, Indicates the current time The connection strength, Indicates the strength of the connection in the previous time period. For example, in a manufacturing company, if the number of technical interaction reports between the R&D department and the production department increases from 20 to 40 in a certain quarter, the dependency trend of this path is 1.0.
[0122] Specifically, if the dependency trend exceeds a preset trend range (for example, the dependency growth threshold is set to 0.5), it indicates that the dependency level of the path is increasing and the network structure needs to be updated.
[0123] During the network structure update process, the present invention dynamically adjusts the initial network diagram based on graph theory methods. When the connection strength of a department increases beyond a threshold, or the path length shortens, the network structure will be automatically adjusted. For example, when the interaction frequency between the sales department and the product department increases by 200%, it means that the product improvement cycle is shortened, and the weight of the path needs to be increased, reflecting the direct impact of market demand on R&D. After the network is updated, the present invention uses a community discovery algorithm to analyze the updated network structure to identify key nodes that require priority resource allocation. The community discovery algorithm uses a modularity optimization method to group departments and calculate the degree of belonging of departments in different communities. For example:
[0124]
[0125] in, represents the edge ratio within the community, Represents the proportion of node degrees in the community. If a department has a high modularity, it indicates that it plays a key role in the network. For example, in a manufacturing company, the supply chain management department has dependencies with multiple departments and a high modularity, so it should be the priority node for resource allocation.
[0126] Specifically, after determining the priority nodes, the present invention employs a path optimization algorithm to analyze path weights and calculate the importance of information transfer. Path optimization utilizes the Analytic Hierarchy Process (AHP) to assign weights to different paths to ensure that key business paths are prioritized. For example, within an enterprise, the product development path can be divided into market feedback → product planning → R&D design → production and manufacturing. Market feedback and product planning have a greater impact, so resources should be allocated first to make the process more efficient. Finally, the present invention combines the optimized weight factors and uses graph theory to update the enterprise network structure, generating a final dynamic network diagram of information flow. This ensures that the network structure can reflect information flow trends within the enterprise in real time, improving the scientific nature and adaptability of resource management.
[0127] In step S14, it is necessary to extract interactive nonlinear features of the dynamic network graph to obtain predictive analysis data for determining node influence, including:
[0128] Extracting node connection relationships, weight data, and time series data from the dynamic network graph;
[0129] Calculating the influence index of each node based on the node connection relationship and the weight data;
[0130] When the influence index is greater than a preset influence threshold, the node corresponding to the influence index is determined as a main node with an amplification effect;
[0131] When the influence index is less than a preset influence threshold, the node corresponding to the influence index is determined as a secondary node with a weakening effect;
[0132] Based on the time series data, the evolution trend analysis of the main nodes and the secondary nodes is performed to obtain predictive analysis data of the evolution trend of the node influence.
[0133] First, it is necessary to extract node connectivity, weight data, and time series data from the dynamic network graph. Node connectivity describes the information exchange between departments within an enterprise and can be represented by establishing an adjacency matrix. For example, in a manufacturing company, the R&D department and the production department share technical documents and collaborate on processes, so there is a direct information connection between the two. Weight data is used to measure the frequency and intensity of interactions between departments. For example, if the marketing department provides 100 pieces of customer feedback to the R&D department each month, and the total number of interactions with other departments is 500, the interaction weight of this path accounts for 20%, indicating that the marketing department has a greater influence on the R&D department. In addition, time series data records the changes in the intensity of information flow over time. For example, the number of order approvals from the procurement department and the supply chain management department shows an increasing or decreasing trend in different months, which can be used to analyze the dynamic evolution of resource allocation.
[0134] After data extraction, nonlinear regression analysis was used to calculate each department's influence index to measure its central role in the enterprise information network. The influence index is calculated based on a combination of factors, including the number of directly connected departments (degree centrality), the role it plays in information flow (betweenness centrality), and the intensity of interaction (weight factor). For example, if a department has direct connections with multiple core business departments and serves as a key bridge in information transfer, its influence index will be high. Conversely, if a department has many connections but low interaction frequency, or its information flow has little impact on other departments, its influence index will be relatively low.
[0135] Specifically, degree centrality Indicates the number of other nodes directly connected to a node, calculated as follows:
[0136]
[0137] in, To connect a value; is the target node for which centrality is to be calculated; To traverse the network, remove the node All other nodes except
[0138] For example, if a department is directly connected to 7 other departments, then , indicating that the department has a strong direct influence.
[0139] Betweenness centrality Indicates the role of the node as a transit bridge in the network, that is, whether the internal information dissemination of the enterprise depends on this node. The calculation method is:
[0140]
[0141] in, It is a slave node arrive The total number of shortest paths, These paths pass through nodes For example, if a purchasing department has 15 out of 20 purchase approval processes that require it to go through the department, .
[0142] Weight impact factor : Indicates the interaction strength of the node, which is calculated as follows:
[0143]
[0144] in, is the weight value of the connection between node i and node j, quantifying the intensity or frequency of the interaction between the two; is the target node for which the weighted influence factor is to be calculated; To traverse the network and nodes All adjacent nodes are connected.
[0145] For example, if the total interaction weight of a node is 0.6, it means that it has a greater impact on the information flow. Finally, the node influence index can be calculated by the nonlinear regression formula:
[0146]
[0147] in, is degree centrality; is the betweenness centrality; is the weight influencing factor; It is the time series growth rate, which represents the change in historical trends and reflects the dynamic development potential (such as the quarterly growth rate of the business scale of a department).
[0148] in, is the regression weight, which can be obtained through historical data training. Specifically, first collect the data of degree centrality, betweenness centrality, weight factor and time series growth rate of each department in the enterprise over a period of time, and mark the actual business impact of each node. Then, use the multivariate regression model to calculate the influence index. As the dependent variable, each centrality index is used as the independent variable, and the least squares method (OLS) or gradient descent method is used to train the regression model to determine the optimal weight coefficient. For example, in a manufacturing company, by analyzing the data of the past year, it was found that degree centrality contributed the most to the influence index. After regression weight training, This indicates that the company's information flow relies primarily on the number of direct connections, followed by the transit function of information transmission, while historical trends have a relatively small impact. After training, the weight parameters are applied to new data to dynamically calculate the influence index of each department, and use this information to optimize the company's resource allocation and organizational structure.
[0149] In order to further distinguish between key nodes and secondary nodes, the present invention sets an influence threshold to ensure accurate identification of departments that play a key role in information flow and resource allocation in the enterprise network. The influence index is obtained from historical data statistics and classified using the mean ± standard deviation method. Specifically, the threshold of the main node is set ,in is the mean influence of all sectors, is the standard deviation, Take 1 or 1.5 and adjust it according to the actual business. For example, in a manufacturing company, the mean influence is 0.5 and the standard deviation is 0.15. If k=1, then Departments with an influence index higher than 0.65 (such as R&D centers and core production departments) are marked as primary nodes, indicating that they assume core decision-making or resource allocation functions in corporate operations. ,like , departments with a score below 0.35 (such as administrative logistics, human resources, and customer service centers) are marked as secondary nodes. These departments have little impact on the overall information flow and have a relatively low priority in resource optimization. and Departments in between (such as quality management and procurement departments) are classified as ordinary nodes, undertaking general business and are not information flow hubs.
[0150] Furthermore, this step combines time series analysis to predict the evolution of each department's influence, enabling the early identification of potential key or weakening nodes, thereby optimizing the company's resource allocation strategy. First, the system collects information flow data from each department over a period of time, including core indicators such as the frequency of business interactions, task allocation, and the number of approval processes. This data is then used to construct a time series dataset for analyzing the dynamic changes in each department's influence. To ensure the accuracy of the analysis, methods such as moving averages, exponential smoothing, or long-short-term memory (LSTM) neural networks are used to model the time series data and extract growth or decline trends in the influence index. Based on this trend analysis, departments that will become key nodes in the future can be identified, allowing the company's management and resource allocation strategies to be adjusted in advance.
[0151] For example, in a manufacturing company, the customer feedback data from the marketing department has shown an increasing trend over the past six months, increasing from 120 to 300, indicating that the impact of market feedback on corporate decision-making is increasing. By calculating the month-on-month growth rate, such as:
[0152]
[0153] The formula is used to calculate the time of a department The month-on-month growth rate of the influence index during the period is a measure of the change in the department's influence between two consecutive time periods (such as days, weeks, or months). Indicates time The influence growth rate at the moment, that is, the degree of growth or decline in the influence of the department in the current time period; Indicates time The department's influence index at any given moment reflects the department's importance in the company's information flow, decision-making participation, and resource allocation; Indicates time The influence index of the department at the moment, that is, the influence value of the previous time period.
[0154] For example, if a department's growth rate remains above 15% for multiple consecutive periods (e.g., the marketing department's growth rates are 12.5%, 18.5%, and 20%), it can be assumed that the department's influence on the company's operations will continue to increase. Companies can proactively adjust resource allocation based on this information, such as adding market data analysts, optimizing R&D and marketing collaboration processes, or adjusting production lines to match the rapid changes in market demand, thereby improving overall responsiveness and resource utilization.
[0155] Conversely, if a department's influence index continues to decline, such as if the production department of a product line is being marginalized due to declining market demand, companies can use the same testing method. For example, the order volume of a production department has continued to decline over the past six months, from 500 to 350, with a month-on-month decline rate of -6.25% and -10.26%, respectively. When this value remains below a set lower threshold (such as -5%) for a long period of time, it indicates that the department's business is declining. In this case, the company can consider adjusting the department's resource allocation, such as optimizing production line layout, reducing manpower and equipment investment, or even consolidating or transforming the business unit to reduce resource waste and improve overall operational efficiency.
[0156] In order to further improve the stability of the prediction, this paper uses exponential smoothing to reduce the noise of short-term fluctuations and predict the future influence index. The exponential smoothing model is calculated as follows:
[0157]
[0158] in, Indicates time The department's influence index at any given moment reflects the department's importance in the company's information flow, decision-making participation, and resource allocation; is the forecast value of the previous period; Forecast value for the future. is a smoothing factor (set to 0.8-0.9) used to balance the impact of historical data. For example, if the influence index of a production department in the current cycle is 350, and the index in the previous cycle was 390, the predicted influence index for the next cycle is: = 356. If the forecast values for multiple future cycles continue to show a downward trend, companies can consider making business adjustments in advance, such as reducing related job positions, reallocating production tasks, or identifying new business growth points to ensure optimal utilization of overall resources.
[0159] In step S15, it is necessary to identify nodes in the network that affect resource allocation and information transmission based on the dynamic network diagram and the prediction analysis data, and obtain key nodes and high-risk nodes, including:
[0160] Extracting node connection data, shortest path data, network topology change trends, clustering coefficients, and community affiliation data from the dynamic network graph;
[0161] Based on the prediction and analysis data, combined with the node connection data, the shortest path data and the network topology change trend, a dynamic weight operation of the centrality index is modified to obtain a centrality index for evaluating the importance of the node in information dissemination;
[0162] Based on the prediction analysis data, combined with the clustering coefficient and the community belonging data, the dynamic weight of the vulnerability index is modified to obtain a vulnerability index for evaluating the degree of dependence of the node in information dissemination;
[0163] According to the centrality index, constructing each node feature vector to obtain a centrality feature vector;
[0164] The K-means clustering algorithm is used to perform importance ranking and screening on the centrality feature vectors to obtain preliminary key nodes of high importance, medium importance and low importance;
[0165] Ranking the importance of the preliminary key nodes to obtain key nodes with high centrality;
[0166] Analyze the connection characteristics and load capacity indicators of each node according to the vulnerability indicators to obtain a vulnerability assessment value;
[0167] When the vulnerability assessment value is less than a preset vulnerability threshold, determining the node corresponding to the vulnerability assessment value as a potential high-risk node;
[0168] Using a support vector machine algorithm to predict node failure of the potential high-risk nodes to obtain the node failure probability;
[0169] A node set whose node failure probability is greater than a preset probability threshold is extracted, and all nodes in the node set are determined as high-risk nodes.
[0170] First, node connectivity data, shortest path data, network topology change trends, clustering coefficient, and community affiliation data are extracted from the dynamic network graph. This data is used to construct basic information about the network structure. Node connectivity data reflects the direct information exchange between departments, shortest path data is used to calculate the optimal path for information flow, network topology change trends monitor the dynamic adjustments of the network structure, the clustering coefficient measures the closeness of a department with surrounding departments, and community affiliation data is used to analyze whether a department belongs to a stable organizational structure. The acquisition of this data relies on the dynamic updating of the network model to ensure that the analyzed dependencies accurately reflect the real-time state of the enterprise organization.
[0171] Then, based on the predictive analysis data and combined with the above-mentioned extracted data, the centrality indicators are corrected and the weights of the centrality indicators are dynamically adjusted to evaluate the importance of each node in information dissemination. Specifically, the centrality indicators include degree centrality, betweenness centrality and closeness centrality, where degree centrality measures the number of direct connections of a node, betweenness centrality measures the transit role of the node in the shortest path, and closeness centrality is used to evaluate the accessibility of the node in the entire network. During the correction process, the present invention adopts a dynamic weight adjustment strategy to adjust the weights of different centrality indicators according to the predictive analysis data. For example, during the period of business adjustment of an enterprise, the propagation pattern of the information flow will change, and the influence of certain key departments will increase. At this time, the weight of the betweenness centrality can be increased to more accurately identify the nodes that play a hub role in the information transmission process.
[0172] At the same time, to identify high-risk nodes, it is necessary to calculate vulnerability indicators to measure the degree of dependence of each node on information flow and resource allocation. Vulnerability indicators are primarily corrected in conjunction with the clustering coefficient and community affiliation data to ensure that risk analysis reflects actual business dependencies. Specifically, a node with a low clustering coefficient means that its connections within the organizational structure are relatively isolated. Once an anomaly occurs in this node, it will be more likely to cause information interruption or resource allocation failure. At the same time, a node with a low degree of community affiliation indicates that the department has weaker connections within the organization and is relatively less stable. Therefore, during the correction process, the system adjusts the weights of the clustering coefficient and community affiliation data to optimize the calculation of vulnerability indicators.
[0173] After calculating the centrality and vulnerability indicators, the next step is to construct a centrality eigenvector for subsequent key node identification. This eigenvector contains each node's degree centrality, betweenness centrality, closeness centrality, and their modified weights, ensuring that the selection of key nodes in the network accurately reflects their role in resource allocation and information dissemination. Subsequently, a K-means clustering algorithm is used to classify and analyze these eigenvectors, categorizing all nodes into three levels: high importance, medium importance, and low importance, thereby forming a preliminary set of key nodes. This process effectively distinguishes between core business departments and general support departments, ensuring that the company prioritizes resource allocation for key departments.
[0174] After initially identifying key nodes, the system further ranks them by importance, identifying those with high centrality. Specifically, the system calculates the influence index of key nodes and ranks them from high to low. For example, in a given enterprise network, the R&D center, marketing department, and core production departments would be identified as high-centrality nodes, while functional departments like administrative support, finance, and auditing would fall into the lower centrality range. This process allows enterprises to identify the core departments that truly influence resource flow and information transmission and adjust resource allocation strategies.
[0175] To assess potential high-risk nodes, the present invention further calculates a vulnerability assessment value, analyzing it in conjunction with the node's connectivity characteristics and load capacity. When a node's vulnerability assessment value falls below a preset vulnerability threshold, indicating a high degree of reliance on information flow and resource allocation, coupled with poor stability, the node is marked as potentially high-risk. For example, in a supply chain management system, if a supplier node has few connections in the network (low degree centrality) and a low degree of belonging to the supply chain community, this indicates a lack of alternative supply channels. Any anomalies would impact the entire production process, and therefore the node should be marked as high-risk.
[0176] After identifying potential high-risk nodes, the present invention uses a support vector machine (SVM) algorithm to predict node failures and calculate the failure probability of each potential high-risk node. By analyzing the dynamic evolution of high-risk nodes in historical data, the SVM predicts the failure probability of a specific node within a future time window. For example, in a financial system, if abnormal fluctuations in the cash flow records of a business department indicate the risk of a future capital chain disruption, the system can train a SVM model based on historical transaction data to issue risk warnings for that department. When the failure probability of a node exceeds a preset probability threshold (e.g., 70%), the node is ultimately labeled as high-risk, allowing the enterprise to take proactive measures such as optimizing information transmission paths, adjusting resource allocation, or establishing alternative plans to mitigate potential losses.
[0177] The following describes step S15 of the present invention using a specific application scenario:
[0178] In the supply chain management of a certain intelligent manufacturing enterprise, the information flow and resource allocation relationship between different departments are complex, especially in the links of raw material supply, production, warehousing, sales, etc. The collaborative efficiency of each department directly affects the operational stability of the enterprise. However, due to the dynamic and nonlinear dependencies of the supply chain, it is difficult for enterprises to accurately identify key nodes and high-risk nodes, resulting in irrational resource allocation and even the risk of supply chain rupture. In order to optimize supply chain management, the present invention uses dynamic network graph analysis and prediction models to identify key nodes and high-risk nodes in the enterprise supply chain, optimize resource allocation strategies, and improve operational stability and adaptability.
[0179] First, supply chain network data is extracted from the company's supply chain management system (SCM), manufacturing execution system (MES), and enterprise resource planning system (ERP). This data includes information flow and resource allocation across multiple dimensions. For example, during the production process, the raw material procurement department needs to maintain constant information exchange with the production department to ensure timely delivery of raw materials. The production department also shares finished product inventory data with the warehousing department to facilitate production planning and inventory management. To quantify these dependencies, the system extracts node connectivity data, shortest path data, network topology change trends, clustering coefficients, and community affiliation data. Among them, node connection data is used to record direct interactions between departments, such as the cooperative network between the procurement department and multiple suppliers; shortest path data is used to calculate the optimal path for information flow and logistics allocation, such as the transportation route of raw materials from suppliers to production lines; network topology change trends are used to monitor structural adjustments in the supply chain, such as the addition of new suppliers or changes in production lines; the clustering coefficient measures the closeness of a department in the supply chain. For example, if the warehousing department is closely connected with multiple production lines, it plays a greater role in resource allocation; community affiliation data is used to divide business modules in the supply chain, such as classifying raw material procurement, production and quality inspection departments as "production communities" and sales, logistics and customer service as "sales communities" to identify the collaborative relationships between various links in the supply chain.
[0180] After acquiring data, the system calculates centrality metrics to identify key nodes in the supply chain. Degree centrality measures the number of direct connections a department has within the supply chain. For example, if a procurement department has established long-term partnerships with 10 suppliers, its degree centrality is high, indicating that the department is crucial to the raw materials supply chain. Betweenness centrality measures the transit role of a node in the supply chain network. For example, if a warehousing department handles multiple functions, such as finished product allocation and raw materials storage, its betweenness centrality is high, indicating that it plays a key role in information and resource transfer. Closeness centrality assesses a department's accessibility within the entire supply chain. For example, a high closeness centrality for a production department indicates that it has direct interactions with multiple business modules, indicating that the department is crucial to the operation of the entire supply chain. By calculating these metrics and combining them with predictive analysis data, the system can dynamically adjust the weighting of each metric to ensure the accurate identification of key nodes.
[0181] After identifying key nodes, the system further calculates vulnerability indicators to identify high-risk nodes in the supply chain. First, the system assesses the stability of a node by combining its clustering coefficient and community affiliation data. A low clustering coefficient for a supplier indicates that the supplier's collaboration with the enterprise is relatively independent, making any supply chain issues more impactful. For example, if Enterprise A relies solely on Supplier B for critical raw materials, a supply chain disruption with Supplier B would significantly reduce Enterprise A's production capacity. Therefore, Supplier B's vulnerability assessment is high, and it should be marked as a high-risk node. Furthermore, a low community affiliation for a node indicates that the department's collaboration within the supply chain network is weak. For example, a logistics company may only maintain a temporary partnership with the enterprise and lack a stable transportation network. If the logistics company ceases service, the enterprise's delivery chain would be disrupted, thus posing a high vulnerability risk. Similarly, if a production line node has few connections but produces critical products, its supply chain dependency is high and its risk is relatively high.
[0182] After identifying potential high-risk nodes, the system uses a support vector machine (SVM) algorithm to predict node failures and calculate the failure probability for each potential high-risk node. The SVM analyzes the dynamic evolution of high-risk nodes in historical data to predict the failure probability of a specific node within a future time window. For example, if a supplier's historical transaction data indicates significant fluctuations in delivery cycles, low inventory turnover, and restricted cash flow over the past three months, its supply stability is poor, with a failure probability exceeding 70%. In this case, the system automatically marks it as a high-risk node and issues an alert to enterprise managers, recommending adjustments to supply chain strategies, such as identifying alternative suppliers, increasing inventory reserves, or optimizing procurement cycles, to reduce supply chain risks.
[0183] In step S16, it is necessary to identify the key propagation paths in the network based on the key nodes and the high-risk nodes, and obtain the key propagation paths that affect resource allocation and information transmission, including:
[0184] Constructing a risk network diagram using graph theory methods based on the key nodes and the high-risk nodes;
[0185] Calculating path parameters of the risk network graph to obtain propagation data and stability data of nodes on the path;
[0186] Evaluate node importance and path stability based on the propagation data and stability data to obtain a path evaluation result;
[0187] The path evaluation results are grouped according to path propagation priority, resource influence, and decision chain weight to obtain key propagation paths that affect resource allocation and information transmission.
[0188] First, the present invention uses graph theory to construct a risk network diagram, which represents the information flow paths and potential risks within an enterprise. In this risk network diagram, key nodes are central to information flow and resource allocation, such as the R&D department, marketing department, and core suppliers. These nodes exhibit high degree centrality, betweenness centrality, or nearness centrality, and play a pivotal role in organizational decision-making and operational management. Meanwhile, high-risk nodes are those where insufficient resource carrying capacity, high dependency, or a high probability of failure hinder information flow and resource supply, such as a single supplier, key equipment maintenance nodes, or departments facing significant financial pressure. To quantify these dependencies, the system constructs an adjacency matrix A to represent the information and resource flow relationships between departments. The values in the matrix represent the intensity of interactions between departments. For example, in a manufacturing enterprise, the R&D department has more information interaction with the production department, resulting in a higher edge weight, while the finance department has less direct interaction with the production department, resulting in a lower edge weight. This risk network diagram allows for a clear visualization of the paths of information and resource flows and further analysis of their propagation characteristics.
[0189] After constructing the risk network graph, the present invention uses a shortest path algorithm (such as the Dijkstra algorithm or the Floyd-Warshall algorithm) to calculate path propagation and stability data within the network. Path propagation measures the efficiency of information or resource transmission along that path, while path stability measures whether the path is susceptible to high-risk nodes. Specifically, the system calculates all information transmission paths and calculates path propagation using the following formula:
[0190]
[0191] in, It represents the betweenness centrality of each node on the path. The higher the value, the more important the node is in information transmission. The larger the value, the greater the influence of the path on resource allocation and decision-making. At the same time, the influence of high-risk nodes is combined to calculate the stability weight of the path:
[0192]
[0193] in, Indicates the failure risk of each node on the path. The higher the value, the lower the stability of the path. Therefore, path stability The higher it is, the more stable the path is.
[0194] After calculating the path propagation and stability, the present invention uses a weighted algorithm to comprehensively evaluate the importance of the path. Calculated by the following formula:
[0195]
[0196] in, , The weight coefficient should be adjusted according to the management needs and risk preferences of the enterprise to ensure that the evaluation of key communication paths is more in line with the actual business scenario. When the enterprise pays more attention to the efficiency of information transmission, the weight of communication can be appropriately increased. , so that the path with smooth information flow gets higher priority in the optimization process; conversely, when the enterprise pays more attention to system stability and risk prevention, the stability weight can be increased. , giving more weight to paths with higher stability and less impact from high-risk nodes. In addition, companies can optimize weight settings through historical data analysis and backtesting. For example, calculate whether paths with higher propagation but lower stability often experience information lag or resource allocation problems over a period of time. If so, it is necessary to increase the weight appropriately. If some high-spreading paths can effectively support resource allocation in the past decision-making process, it can improve The value of .
[0197] After determining the importance of a path, the present invention further uses clustering algorithms (such as K-means or hierarchical clustering) to categorize the path evaluation results to better understand and optimize the management of an enterprise's information and resource flows. Key communication paths are primarily categorized as follows: High-priority paths, which are efficient and stable, such as R&D → Production → Logistics. These paths are crucial to the flow of enterprise resources and their smooth operation must be prioritized; High-risk paths, which are efficient but less stable, such as key suppliers → Procurement → Production. Problems in these paths could hinder enterprise operations and therefore require focused monitoring and contingency plans; Redundant paths, which are inefficient but highly stable, such as routine management processes like Administration → Finance → Auditing. These paths can be used for general management optimization without requiring excessive attention.
[0198] In step S17, it is necessary to predict the key locations where resource bottlenecks and resource conflicts occur based on the key nodes, the high-risk nodes, and the key propagation paths, and generate resource optimization suggestions, including:
[0199] Analyze resource allocation status according to the key nodes, the high-risk nodes, and the key propagation paths to obtain resource usage of key nodes, resource pressure of high-risk nodes, and traffic load of key propagation paths;
[0200] Predicting key locations where resource bottlenecks and resource conflicts occur based on the resource usage, resource pressure, and traffic load, combined with preset historical resource scheduling data;
[0201] Adjust resource allocation strategies for the key locations and generate resource optimization suggestions.
[0202] First, the system extracts information about key nodes, high-risk nodes, and critical communication paths from the previous step and analyzes their resource allocation status. Resource usage at key nodes includes the node's current resource occupancy rate, consumption rate, and remaining available resources. For example, in a manufacturing enterprise, the production department, as a key node, can measure its resource usage by indicators such as the utilization rate of production equipment, the rate of consumption of inventory raw materials, and staffing load. Resource pressure at high-risk nodes involves the stability of the node in resource scheduling, such as when a single supplier, a single logistics channel, or a key server becomes a system bottleneck. This component analyzes the resource redundancy (current inventory / demand), load ratio (current workload / maximum capacity), and historical resource fluctuations of key computing nodes to determine the node's stability. Traffic load at critical communication paths measures the flow of resources or information along different paths, including metrics such as resource transmission rate, average response time, and peak traffic. For example, in a logistics transportation network, the transportation capacity from supplier to warehouse and from warehouse to production line can be assessed by hourly freight volume and historical delays.
[0203] Specifically, after completing the resource status analysis, this step combines historical resource scheduling data with machine learning algorithms to predict resource bottlenecks and resource conflicts. First, regression analysis is used to predict resource demand trends at key nodes. Specifically, time series models (such as ARIMA and LSTM) are used to analyze resource consumption data over a period of time to predict future resource demand growth trends. For example, if sales data indicates that market demand for a certain product is expected to grow by 30% over the next month, the system can predict that the production department's material demand will also increase during that period and adjust procurement plans in advance. Second, a resource shortage risk model is trained using classification algorithms (such as random forests and support vector machines (SVMs). By analyzing factors such as supply shortages and equipment failures in historical data, it predicts which high-risk nodes will become bottlenecks due to resource shortages. For example, if a key supplier experienced delivery delays in four of the past 12 months, the system can predict a 30%-40% risk of supply instability and recommend adding backup suppliers. Furthermore, traffic analysis models (such as Bayesian networks) are used to predict traffic changes along key transmission paths to proactively identify emerging congestion issues. For example, in an enterprise's information system, if it is predicted that server access volume will surge by 50% during the annual audit period, the system can warn that the computing resource demand for this path is insufficient and recommend expanding server capacity in advance.
[0204] Specifically, based on predictions of resource bottlenecks and conflict risks, the present invention employs optimization algorithms to adjust resource allocation at key locations to reduce systemic risk and improve resource utilization. Optimization methods include dynamic resource scheduling, resource redundancy optimization, and intelligent path adjustment. In dynamic resource scheduling, the system reallocates available resources across different departments based on resource demand forecasts. For example, when a production line faces the risk of shutdown due to a shortage of raw materials, the system can recommend allocating raw materials from other production lines with higher inventory levels to ensure continued production. Regarding resource redundancy optimization, if the system detects that the resource carrying capacity of a high-risk node is approaching its upper limit, such as when a key supplier's supply capacity is saturated, the system can recommend adding backup resources, such as adding new suppliers or increasing safety stock levels, to mitigate supply chain risk. In intelligent path adjustment, the system can optimize resource allocation for critical transmission paths where resource flow is blocked. For example, in a data transmission network, if the system predicts that a server's load will exceed 80% during peak hours, it can preemptively transfer some computing tasks to backup servers to improve overall network stability.
[0205] In a specific example, suppose a multinational manufacturing company wants to optimize resource scheduling within its global supply chain. The system first analyzes key nodes (production workshops, suppliers), high-risk nodes (single raw material supplier), and critical transmission paths (supplier → warehouse → production line). It discovers that a major supplier's order fulfillment capacity fluctuates, while inventory levels on the production line are low. Based on historical data analysis, the system predicts a 10%-15% risk of delivery delays for this supplier over the next three weeks, potentially leading to production bottlenecks. To address this, the system generates the following optimization recommendations: Increase backup suppliers to reduce the company's reliance on a single supplier to ensure a stable supply of raw materials; adjust inventory management strategies, increase inventory safety thresholds, and increase reserves of key materials to mitigate the impact of supply chain fluctuations; optimize production plans by adjusting production schedules to prioritize products that don't rely on this supplier's raw materials to avoid downtime due to stockouts; and optimize logistics routes. If peak-hour traffic congestion on a particular route impacts supply chain efficiency, the system recommends rerouting or increasing shipments to ensure timely delivery of raw materials.
[0206] In summary, the present invention provides a method and system for enterprise digital management based on data mining. The present invention can identify the dynamic changes of the enterprise's implicit dependencies, locate high-risk nodes and key transmission paths, and optimize resource allocation.
[0207] Reference Figure 2 The second embodiment of the present invention provides an enterprise digital management system based on data mining, including:
[0208] Data acquisition module, used to obtain information flow data and resource allocation data between departments;
[0209] A dependency analysis module, configured to perform dependency relationship analysis based on the information flow data and the resource allocation data to obtain an initial network diagram of inter-departmental dependency relationships;
[0210] A path optimization module is used to perform cluster analysis on the connection strength and path length based on the initial network diagram, dynamically reconstruct the network, and integrate community mining and path optimization to identify key nodes, thereby obtaining a dynamic network diagram of information flow between departments;
[0211] A prediction analysis module is used to extract interactive nonlinear features of the dynamic network graph to obtain prediction analysis data for determining node influence;
[0212] A node identification module is used to identify nodes in the network that affect resource allocation and information transmission based on the dynamic network diagram and the prediction analysis data, and obtain key nodes and high-risk nodes;
[0213] A critical path module is used to identify the critical propagation paths in the network based on the key nodes and the high-risk nodes, and obtain the critical propagation paths that affect resource allocation and information transmission;
[0214] The resource optimization module is used to predict the key locations where resource bottlenecks and resource conflicts occur based on the key nodes, the high-risk nodes and the key propagation paths, and generate resource optimization suggestions.
[0215] It should be noted that the enterprise digital management system based on data mining provided in an embodiment of the present invention is used to execute all the process steps of the enterprise digital management method based on data mining in the above embodiment. The working principles and beneficial effects of the two correspond one to one, so they will not be repeated here.
[0216] An embodiment of the present invention further provides an electronic device. The electronic device includes: a processor, a memory, and a computer program stored in the memory and executable on the processor, such as a prediction analysis program. When the processor executes the computer program, the steps in the above-mentioned embodiments of the enterprise digital management method based on data mining are implemented, such as Figure 1 Alternatively, when the processor executes the computer program, the functions of the modules / units in the above-mentioned device embodiments are realized, such as the node identification module.
[0217] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory and executed by the processor to implement the present invention. The one or more modules / units may be a series of computer program instruction segments capable of implementing specific functions, and the instruction segments are used to describe the execution process of the computer program in the electronic device.
[0218] The electronic device may be a computing device such as a desktop computer, notebook, PDA, or smart tablet. The electronic device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that the aforementioned components are merely examples of electronic devices and do not constitute a limitation of the electronic device. The electronic device may include more or fewer components than those described above, or a combination of certain components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, and the like.
[0219] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor. The processor is the control center of the electronic device and connects various parts of the entire electronic device using various interfaces and lines.
[0220] The memory can be used to store the computer programs and / or modules. The processor implements the various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory and accessing the data stored in the memory. The memory may primarily include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function or an image playback function); the data storage area may store data generated based on the use of the mobile phone (such as audio data, a phone book, etc.). Furthermore, the memory may include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.
[0221] If the module / unit integrated into the electronic device is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the process steps in the above-mentioned method embodiments by using a computer program to instruct the relevant hardware. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments. The computer program includes computer program code, which can be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium can be appropriately increased or decreased based on the requirements of legislation and patent practice in a jurisdiction. For example, in some jurisdictions, based on legislation and patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.
[0222] It should be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. A person of ordinary skill in the art can understand and implement the present invention without inventive effort.
[0223] The specific embodiments described above further illustrate the objectives, technical solutions, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. In particular, it should be noted that any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included within the scope of protection of the present invention for those skilled in the art.
Claims
1. A data mining-based enterprise digital management method, characterized by include: Obtain data on information flow and resource allocation between departments; Based on the information flow data and the resource allocation data, dependency analysis is performed to obtain an initial network diagram of inter-departmental dependencies, including: Performing data fusion on information flow data and resource allocation data, wherein the data fusion includes time alignment, entity matching, and similarity analysis; Based on the fused data, an association rule mining algorithm is used to calculate the support and confidence between departments, including: constructing a transaction data set, where each transaction represents a complete business interaction, and all departments in the transaction constitute an itemset; performing frequent item set mining to screen out department combinations with stable dependencies, wherein the frequent item set mining includes counting the occurrence frequency of each individual department, calculating the support, and screening out departments with a support threshold above the support threshold to form a frequent item set; based on the frequent item set, generating candidate department pairs, and calculating the number of their co-occurrences, screening out department pairs with a support higher than the support threshold; after obtaining the frequent item set, calculating the confidence of each candidate department pair, and screening out department pairs with a confidence higher than a preset threshold; further expanding to combinations of multiple departments, calculating more complex dependencies, until no new high-frequency collaborative department combinations can be found; When the support and the confidence are respectively greater than a preset support threshold and a preset confidence threshold, a dependency network structure diagram between departments is established based on a graph theory method to obtain an initial network diagram of the dependency relationship between departments; Based on the initial network diagram, cluster analysis is performed to analyze connection strength and path length, and the network is dynamically reconstructed and community mining and path optimization are integrated to identify key nodes, thereby obtaining a dynamic network diagram of information flow between departments; Extracting interactive nonlinear features of the dynamic network graph to obtain predictive analysis data for determining node influence; Identify nodes in the network that affect resource allocation and information transmission based on the dynamic network diagram and the prediction analysis data, and obtain key nodes and high-risk nodes; Identify key communication paths in the network based on the key nodes and the high-risk nodes, and obtain key communication paths that affect resource allocation and information transmission; Based on the key nodes, the high-risk nodes and the key propagation paths, key locations where resource bottlenecks and resource conflicts occur are predicted, and resource optimization suggestions are generated.
2. The enterprise digital management method based on data mining according to claim 1 is characterized in that: According to the initial network diagram, cluster analysis is performed on connection strength and path length, and the network is dynamically reconstructed and community mining and path optimization are integrated to identify key nodes to obtain a dynamic network diagram of information flow between departments, including: Extracting connection strength data and path length data from the initial network graph; Perform resource allocation pattern analysis based on the connection strength data and the path length data to obtain the main transmission paths of inter-departmental dependency relationships; Perform dependency trend analysis on the main transmission paths to obtain dependency trends among departments; When the dependency trend is greater than a preset trend range, the network structure of the initial network diagram is updated based on a graph theory method to obtain an updated network diagram of the dependency relationship between departments; Identify key nodes for resource allocation based on the updated network graph to obtain priority nodes that require priority resource allocation; Performing path weight analysis on the priority nodes to obtain weight factors reflecting the importance of information transmission between departments; According to the weight factors, the dependency relationship between departments is updated to obtain a dynamic network diagram of information flow between departments.
3. The enterprise digital management method based on data mining according to claim 1 is characterized in that: The interactive nonlinear feature extraction of the dynamic network graph to obtain predictive analysis data for determining node influence includes: Extracting node connection relationships, weight data, and time series data from the dynamic network graph; Calculating the influence index of each node based on the node connection relationship and the weight data; When the influence index is greater than a preset influence threshold, the node corresponding to the influence index is determined as a main node with an amplification effect; When the influence index is less than a preset influence threshold, the node corresponding to the influence index is determined as a secondary node with a weakening effect; Based on the time series data, the evolution trend analysis of the main nodes and the secondary nodes is performed to obtain predictive analysis data of the evolution trend of the node influence.
4. The enterprise digital management method based on data mining according to claim 1 is characterized in that: The step of identifying nodes in the network that affect resource allocation and information transmission based on the dynamic network diagram and the prediction analysis data, and obtaining key nodes and high-risk nodes, includes: Extracting node connection data, shortest path data, network topology change trends, clustering coefficients, and community affiliation data from the dynamic network graph; Based on the prediction and analysis data, combined with the node connection data, the shortest path data and the network topology change trend, a dynamic weight operation of the centrality index is modified to obtain a centrality index for evaluating the importance of the node in information dissemination; Based on the prediction analysis data, combined with the clustering coefficient and the community belonging data, the dynamic weight of the vulnerability index is modified to obtain a vulnerability index for evaluating the degree of dependence of the node in information dissemination; According to the centrality index, constructing each node feature vector to obtain a centrality feature vector; The K-means clustering algorithm is used to perform importance ranking and screening on the centrality feature vectors to obtain preliminary key nodes of high importance, medium importance and low importance; Ranking the importance of the preliminary key nodes to obtain key nodes with high centrality; Analyze the connection characteristics and load capacity indicators of each node according to the vulnerability indicators to obtain a vulnerability assessment value; When the vulnerability assessment value is less than a preset vulnerability threshold, determining the node corresponding to the vulnerability assessment value as a potential high-risk node; Using a support vector machine algorithm to predict node failure of the potential high-risk nodes to obtain the node failure probability; A node set whose node failure probability is greater than a preset probability threshold is extracted, and all nodes in the node set are determined as high-risk nodes.
5. The enterprise digital management method based on data mining according to claim 1 is characterized in that: The identifying of key propagation paths in the network based on the key nodes and the high-risk nodes to obtain the key propagation paths that affect resource allocation and information transmission includes: Constructing a risk network diagram using graph theory methods based on the key nodes and the high-risk nodes; Calculating path parameters of the risk network graph to obtain propagation data and stability data of nodes on the path; Evaluate node importance and path stability based on the propagation data and stability data to obtain a path evaluation result; The path evaluation results are grouped according to path propagation priority, resource influence, and decision chain weight to obtain key propagation paths that affect resource allocation and information transmission.
6. The enterprise digital management method based on data mining according to claim 1 is characterized in that: The method of predicting key locations where resource bottlenecks and resource conflicts occur based on the key nodes, the high-risk nodes, and the key propagation paths, and generating resource optimization suggestions, includes: Analyze resource allocation status according to the key nodes, the high-risk nodes, and the key propagation paths to obtain resource usage of key nodes, resource pressure of high-risk nodes, and traffic load of key propagation paths; Predicting key locations where resource bottlenecks and resource conflicts occur based on the resource usage, resource pressure, and traffic load, combined with preset historical resource scheduling data; Adjust resource allocation strategies for the key locations and generate resource optimization suggestions.
7. An enterprise digital management system based on data mining, characterized in that: A method for implementing a digital enterprise management method based on data mining as claimed in any one of claims 1 to 6, comprising: Data acquisition module, used to obtain information flow data and resource allocation data between departments; A dependency analysis module, configured to perform dependency relationship analysis based on the information flow data and the resource allocation data to obtain an initial network diagram of inter-departmental dependency relationships; A path optimization module is used to perform cluster analysis on the connection strength and path length based on the initial network diagram, dynamically reconstruct the network, and integrate community mining and path optimization to identify key nodes, thereby obtaining a dynamic network diagram of information flow between departments; A prediction analysis module is used to extract interactive nonlinear features of the dynamic network graph to obtain prediction analysis data for determining node influence; A node identification module is used to identify nodes in the network that affect resource allocation and information transmission based on the dynamic network diagram and the prediction analysis data, and obtain key nodes and high-risk nodes; A critical path module is used to identify the critical propagation paths in the network based on the key nodes and the high-risk nodes, and obtain the critical propagation paths that affect resource allocation and information transmission; The resource optimization module is used to predict the key locations where resource bottlenecks and resource conflicts occur based on the key nodes, the high-risk nodes and the key propagation paths, and generate resource optimization suggestions.
8. An electronic device, characterized in that: The method comprises a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the enterprise digital management method based on data mining as described in any one of claims 1 to 6 is implemented.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the enterprise digital management method based on data mining according to any one of claims 1 to 6.
Citation Information
Patent Citations
Enterprise digital intelligent operation method and system based on data mining
CN115222301A
Cloud computing task tracking processing method and system
CN118656200A
Electric power project risk prediction method and system based on data mining
CN119130112A
Operation and maintenance management system and method based on artificial intelligence
CN119676055A