Multi-source heterogeneous operation and maintenance data analysis method and system based on industry subject portraits
By constructing an industry knowledge graph and generating industry entity profiles, the problem of integrating multi-source heterogeneous data has been solved, enabling accurate anomaly detection, root cause localization, and risk warning, improving the depth and timeliness of analysis, and supporting intelligent industry situation awareness and decision-making.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANDONG RONGWEI INFORMATION TECH CO LTD
- Filing Date
- 2025-12-25
- Publication Date
- 2026-04-17
AI Technical Summary
Existing technologies struggle to effectively integrate multi-source heterogeneous data, making it impossible to achieve accurate anomaly detection, rapid root cause localization, and proactive risk warning. The depth and timeliness of analytical conclusions are insufficient, and reliance on human experience leads to inefficiency.
By constructing an industry knowledge graph, we can generate profiles of industry stakeholders, perform anomaly detection, root cause localization, and predictive early warning. We can also use graph feature indicators and profile tags for multi-dimensional analysis, and combine machine learning models for root cause localization and risk warning.
It achieves deep integration and knowledge-based processing of multi-source heterogeneous data, improving the accuracy and automation of anomaly detection, root cause localization, and risk prediction, and providing intelligent industry situational awareness and decision support.
Smart Images

Figure CN121880784A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of data processing technology, specifically relating to a method and system for analyzing multi-source heterogeneous operation and maintenance data based on industry subject profiles. Background Technology
[0002] In the current wave of digital transformation, the operations of many industries, such as finance, energy, and manufacturing, are becoming increasingly complex, generating massive amounts of multi-source, heterogeneous data. This data is scattered across various monitoring systems, business platforms, and public channels, forming serious data silos and making it difficult to grasp the industry landscape from a holistic perspective. Traditional analytical methods are often limited to single data sources or post-event statistics, lacking the ability to mine deep relationships between data and effectively build an industry knowledge system. Furthermore, due to the failure to combine micro-level entity behavior with macro-level industry characteristics, existing technologies struggle to achieve accurate anomaly detection, rapid root cause localization, and proactive risk warnings. The depth and timeliness of analytical conclusions are insufficient, relying heavily on human experience and resulting in low efficiency. Therefore, the industry urgently needs an integrated analytical method that can break down data silos, integrate knowledge and profiles, and support intelligent analysis. Summary of the Invention
[0003] To address the aforementioned shortcomings of existing technologies, this invention provides a method and system for analyzing multi-source heterogeneous operation and maintenance data based on industry entity profiles, in order to solve the above-mentioned technical problems.
[0004] In a first aspect, the present invention provides a method for analyzing multi-source heterogeneous operation and maintenance data based on industry entity profiles, including: Multi-source heterogeneous data from multiple heterogeneous data sources is collected from industry entities, and the multi-source heterogeneous data is cleaned and standardized to form standard data. Based on the standard data, an industry knowledge graph containing industry subject nodes, knowledge element nodes, and their relationships is constructed through entity recognition and relationship extraction. Based on the part of the standard data that is related to the target industry entity, and by integrating the graph feature indicators of the entity extracted from the industry knowledge graph, a quantitative score is calculated for the industry entity in the multi-dimensional profile dimension. The quantitative score is then mapped to a profile label according to a preset threshold to generate an industry entity profile. Based on the industry knowledge graph and the industry entity profile, at least one of the following analyses is performed on the industry entity: anomaly detection, root cause localization, predictive early warning, or multi-dimensional dynamic analysis. Based on the analysis results, conclusions characterizing the health status and risk situation of the industry are generated.
[0005] In one optional implementation, multi-source heterogeneous data from multiple heterogeneous data sources is collected from industry entities, including: By configuring agents, API interfaces, and log collection tools, multi-source heterogeneous data can be collected in real time or offline from the operation and maintenance systems of any industry, such as finance, energy, or manufacturing. The multi-source heterogeneous data includes publicly available data from industry entities and internal platform behavioral data; the publicly available data includes at least one of business registration information, patent information, financial reports, or public opinion data; the internal platform behavioral data includes at least one of search records, transaction behavior, or content publishing data.
[0006] In one optional implementation, the multi-source heterogeneous data is cleaned and standardized, including: The collected operation and maintenance data is cleaned to remove noise, missing values and outliers, and the data with different formats is converted into a standardized format based on a predefined unified data model. By using host IP, device ID, or transaction serial number as a unified identifier, standardized data from different data sources are associated to form an information view with complete context. The associated data is stored in a unified time-series database.
[0007] In an optional implementation, based on the standard data, an industry knowledge graph is constructed through entity recognition and relationship extraction, comprising industry subject nodes, knowledge element nodes, and the relationships between them, including: Natural language processing technology is used to identify entities and extract relationships between entities from unstructured or semi-structured text in the standard data. Based on entities and relationships, an industry knowledge graph is constructed with "enterprise", "product" and "technology" as nodes and "dependency" and "competition" as associations. Based on the dependence of industry entities in the knowledge graph and based on preset multi-level thresholds, industry entities are divided into different levels; wherein, the level includes at least a first-level entity with high dependence and a second-level entity with medium dependence. The industry knowledge graph is stored in a graph database and supports traversal and querying using a graph query language.
[0008] In an optional implementation, based on the portion of the standard data relevant to the target industry entity, and integrating the graph feature indicators of the entity extracted from the industry knowledge graph, a quantitative score is calculated for the industry entity on a multi-dimensional profile dimension. The quantitative score is then mapped to profile tags according to a preset threshold to generate an industry entity profile, including: For each preset profile dimension, corresponding multi-source raw indicators are collected from the standard data and the industry knowledge graph; wherein, the profile dimension includes technical strength, market influence and innovation activity. The multi-source raw indicators include statistical indicators obtained from the standard data and graph feature indicators extracted from the industry knowledge graph, wherein the graph feature indicators include node centrality or PageRank value. For direct quantitative indicators, their values are obtained directly; for complex quantitative indicators, they are calculated using a preset rule model to obtain quantitative values; wherein, the complex quantitative indicators include at least a patent quality score, which is calculated by combining the basic score of a single patent, the weight of the number of citations, and the weight of the International Patent Classification number. Use Min-Max standardization, Z-Score standardization, or quantile standardization to uniformly scale the values of various indicators with different dimensions to the same comparable scale. Weights are assigned to each standardized indicator under the same profile dimension, and the weights are determined by at least one of the following methods: expert scoring, entropy weighting, or AHP (analytic hierarchy process). By weighted summing of the quantitative values of each indicator, the quantitative score of the industry entity in each profile dimension is obtained. The quantitative score is compared with a preset threshold to generate a corresponding profile label; wherein the preset threshold is set by quantile method, absolute standard method or cluster analysis method; the profile label includes a single-dimensional level label and an overall qualitative label generated by combining multi-dimensional labels based on decision tree rules.
[0009] In an optional implementation, the quantified score is compared with a preset threshold to generate a corresponding profile label, including: Single-dimensional level labeling: The quantitative score of a single profile dimension is compared with the level threshold preset for that dimension to generate a qualitative label describing the level of that dimension; wherein, the level threshold is set by the quantile method, including the threshold range for classifying industry entities into "leaders", "leaders", "followers" or "potential stocks". The overall qualitative label generated based on the combination of decision tree rules: Multiple single-dimensional level labels are taken as input, and logical judgments are performed through predefined decision tree rules to output a comprehensive overall qualitative label; wherein, the decision tree rules include at least: When an industry entity is rated as a "leader" in both the "technological strength" and "market influence" dimensions, it is given the overall qualitative label of "industry leader". When an industry entity is rated as a "leader" in the "innovation activity" dimension, but is rated as a "potential stock" in the "market influence" dimension, it is given the overall qualitative label of "hidden champion".
[0010] In one optional implementation, based on the industry knowledge graph and the industry entity profile, at least one of the following analyses is performed on the industry entity: anomaly detection, root cause localization, predictive early warning, or multi-dimensional dynamic analysis, including: Based on the tags in the industry entity profile, the industry entities are divided into different analysis groups; Establish dynamic behavioral baselines based on historical core indicator data for each analysis group; The current indicator data of the industry entity is calculated in real time by the stream processing engine and compared with the dynamic behavior baseline of the analysis group. When the calculated deviation exceeds the threshold determined based on the historical volatility of the group, an anomaly alarm is triggered. In response to an anomaly alert for a target entity, a subgraph containing the target entity and its associated nodes is extracted from the industry knowledge graph using a graph query language. Examine the status indicators of each associated node in the subgraph near the abnormal time window, and filter out the associated nodes that have abnormal indicators at the same time. By combining the profile tags, graph feature attributes, and indicator changes of the target subject and its associated nodes, the contribution of each potential root cause is calculated through a machine learning model, and a root cause localization report is output.
[0011] Secondly, this invention provides a multi-source heterogeneous operation and maintenance data analysis system based on industry entity profiles, including: The data acquisition module is used to collect multi-source heterogeneous data from multiple heterogeneous data sources from industry entities, and to clean and standardize the multi-source heterogeneous data to form standard data. The graph construction module is used to construct an industry knowledge graph based on the standard data, through entity recognition and relationship extraction, which includes industry subject nodes, knowledge element nodes and the relationships between them. The profile generation module is used to calculate a quantitative score for the industry entity in the multi-dimensional profile dimension based on the part of the standard data that is related to the target industry entity and to integrate the graph feature indicators of the entity extracted from the industry knowledge graph. The quantitative score is then mapped to profile tags according to a preset threshold to generate an industry entity profile. The subject analysis module is used to perform at least one of the following analyses on the industry subject based on the industry knowledge graph and the industry subject profile: anomaly detection, root cause localization, predictive early warning, or multi-dimensional dynamic analysis. The conclusion generation module is used to generate conclusions characterizing the health status and risk situation of the industry based on the obtained analysis results.
[0012] Thirdly, a device is provided, comprising: The memory is used to store multi-source heterogeneous operation and maintenance data analysis programs based on industry entity profiles; The processor is used to implement the steps of the multi-source heterogeneous operation and maintenance data analysis method based on industry entity profiles as provided in the first aspect when executing the multi-source heterogeneous operation and maintenance data analysis program based on industry entity profiles.
[0013] Fourthly, a computer-readable medium is provided, on which a multi-source heterogeneous operation and maintenance data analysis program based on industry entity profiles is stored. When the multi-source heterogeneous operation and maintenance data analysis program based on industry entity profiles is executed by a processor, it implements the steps of the multi-source heterogeneous operation and maintenance data analysis method based on industry entity profiles provided in the first aspect.
[0014] The beneficial effects of this invention lie in the fact that the multi-source heterogeneous operation and maintenance data analysis method and system based on industry entity profiles provided by this invention completely breaks down the barriers between multi-source heterogeneous data by constructing a unified data model and industry knowledge graph, realizing deep data integration and knowledge-based processing. By generating dynamically updated industry entity profiles, the system can establish accurate baselines for individual and group behavior, thereby achieving a shift from passive response to proactive early warning. Its core beneficial effect is that by combining the association analysis capabilities of knowledge graphs with the group segmentation capabilities of profile tags, the accuracy, automation, and interpretability of anomaly detection, root cause localization, and risk prediction are significantly improved, ultimately providing intelligent and visual support for industry situational awareness and decision-making, and significantly enhancing the efficiency and reliability of operation and maintenance management. Attached Figure Description
[0015] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0016] Figure 1 This is a schematic flowchart of a method according to an embodiment of the present invention.
[0017] Figure 2 This is a schematic block diagram of a system according to an embodiment of the present invention.
[0018] Figure 3 This is a schematic diagram of the structure of a device provided in an embodiment of the present invention. Detailed Implementation
[0019] To enable those skilled in the art to better understand the technical solutions of this invention, the technical solutions of the embodiments of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this invention, and not all embodiments. Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this invention.
[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains. The terminology used herein in the description of the invention is for the purpose of describing particular embodiments only and is not intended to be limiting of the invention.
[0021] The multi-source heterogeneous operation and maintenance data analysis method based on industry entity profiles provided in this embodiment of the invention is executed by computer equipment, and correspondingly, the multi-source heterogeneous operation and maintenance data analysis system based on industry entity profiles runs on computer equipment.
[0022] Figure 1 This is a schematic flowchart illustrating a method according to an embodiment of the present invention. Wherein, Figure 1 The executing entity can be a multi-source heterogeneous operation and maintenance data analysis system based on industry entity profiles. Depending on different needs, the order of the steps in this flowchart can be changed, and some can be omitted.
[0023] like Figure 1 As shown, the method includes: S1. Collect multi-source heterogeneous data from multiple heterogeneous data sources of industry entities, and clean and standardize the multi-source heterogeneous data to form standard data; S2. Based on the standard data, an industry knowledge graph containing industry subject nodes, knowledge element nodes and their relationships is constructed through entity recognition and relationship extraction. S3. Based on the part of the standard data that is related to the target industry entity, and by integrating the graph feature indicators of the entity extracted from the industry knowledge graph, calculate the quantitative score of the industry entity in the multi-dimensional profile dimension, and map the quantitative score to the profile label according to the preset threshold to generate an industry entity profile. S4. Based on the industry knowledge graph and the industry entity profile, perform at least one of the following analyses on the industry entity: anomaly detection, root cause localization, predictive early warning, or multi-dimensional dynamic analysis; S5. Based on the obtained analysis results, generate conclusions characterizing the health status and risk situation of the industry.
[0024] In one embodiment of the present invention, based on step S1, the following will provide a possible embodiment and describe its specific implementation in a non-limiting manner.
[0025] Data Acquisition Tool Configuration and Deployment: The system configures and deploys corresponding acquisition components based on the data source type and access method. For servers and network devices that need to actively pull data, we deploy lightweight agents, such as Telegraf or Prometheus Exporter, on the target nodes to collect performance metrics periodically. For external systems or public data sources that provide standardized interfaces, we configure API interfaces for periodic polling or receive pushed data via webhooks. For example, we use Python's requests library to call the Tianyancha API to obtain enterprise business information, or call the State Intellectual Property Office interface to obtain patent data. For streaming log data generated by applications and systems, we use log collection tools, such as Filebeat or Fluentd, for real-time tracking and collection.
[0026] Real-time and Offline Data Acquisition from Multiple Industry Data Sources: This system is applicable to multiple industries, including finance, energy, and manufacturing. In the financial sector, we collect real-time data such as trading volume, response latency, and error codes from core trading systems, network management platforms, and application logs. In the energy sector, we collect real-time data on power generation equipment status, grid load, and environmental sensor data through agents and interfaces deployed in the SCADA system. The system supports both real-time streaming and offline batch acquisition modes. Real-time data is accessed via Apache Kafka message queues for immediate analysis; offline data (such as historical financial reports and archived logs) can be imported in batches via SFTP or object storage services (such as Amazon S3).
[0027] Classification and Examples of Multi-Source Heterogeneous Data: The collected data is clearly divided into two main categories: Publicly available data: This data is obtained from external data sources through the aforementioned API interfaces. Specific implementations may include: registered capital and business scope of enterprises obtained from business information query platforms; patent application numbers and authorization announcement dates obtained from patent databases; annual financial reports of listed companies obtained from financial information sources; and news and social media sentiment analysis data about specific enterprises obtained from public opinion monitoring systems.
[0028] Internal platform behavioral data: Captured from internal business systems through log collection tools and agents, specifically including: user search keywords and click records within the platform; transaction success / failure status, transaction amount and channel generated during the transaction process; and articles, comments and solution content published by users in the community or knowledge base.
[0029] The process of cleaning and standardizing the multi-source heterogeneous data includes: cleaning the collected operation and maintenance data to remove noise, missing values, and outliers; converting the data of different formats into a standardized format based on a predefined unified data model; associating the standardized data from different data sources using host IP, device ID, or transaction serial number as a unified identifier to form an information view with a complete context; and storing the associated data in a unified time-series database.
[0030] In one embodiment of the present invention, based on step S2, the following will provide a possible embodiment and describe its specific implementation in a non-limiting manner.
[0031] S201. Natural Language Processing Implementation of Entity and Relation Extraction The system employs a deep learning-based natural language processing pipeline to extract knowledge from unstructured / semi-structured text. In practice: Named entity recognition is performed using pre-trained models (such as BERT-BiLSTM-CRF) to accurately identify entities such as "enterprise" (e.g., "Huawei Technologies Co., Ltd."), "product" (e.g., "Kirin chip"), and "technology" (e.g., "5G communication") from news, financial reports, and patent documents.
[0032] A relation extraction method based on dependency parsing is adopted, combined with predefined relation patterns, to extract "dependency" relations (such as "Company A depends on Company B's chip supply") and "competition" relations (such as "Company C and Company D compete in the smartphone market") from the text.
[0033] For semi-structured data (such as shareholder information of listed companies), a special parser is developed to directly extract relationships such as "holding" and "participation".
[0034] S202. Dependency Calculation and Hierarchical Partitioning Based on Graph Algorithms After the map is constructed, the system executes the following analysis process: The dependency index of each industry entity node is calculated using a graph algorithm, specifically including: PageRank algorithm: Evaluates the global influence of a node in the entire graph; Degree centrality: Calculates the in-degree and out-degree of a node to measure its dependence on others; Betweenness centrality: Identifying key nodes that act as bridges in a network.
[0035] Based on the above indicators, a comprehensive dependency score is constructed, and a hierarchical division is performed using a multi-level threshold based on quantiles: First-level entities: Nodes with comprehensive dependency scores in the top 15% are marked as core dependency nodes; Second-level entities: Nodes with comprehensive dependency scores between 15% and 50% are marked as important dependency nodes; The remaining nodes can be further subdivided or uniformly classified according to business requirements.
[0036] S203. Optimization of Graph Database Storage and Query The constructed industry knowledge graph is stored in the Neo4j graph database. The specific implementation details include: Design an optimized graph data model, store entities as nodes, relationships as edges, and establish indexes for different types of nodes and relationships; Use the Cypher query language to achieve efficient graph traversal, for example: MATCH (company:Enterprise)-[r:Dependency|Competition]->(other) WHERE company.name = "Target Enterprise" RETURN company, r, other; To achieve high-performance queries, establish indexes for common query paths and configure appropriate caching policies; Develop a graph service layer based on Spring Boot to provide RESTful APIs for the front end and other systems to call and query.
[0037] In an embodiment of the present invention, based on step S3, a possible embodiment will be given below to non-restrictively elaborate on its specific implementation scheme.
[0038] S301. For each preset portrait dimension, collect corresponding multi-source raw metrics from the standard data and the industry knowledge graph; wherein, the portrait dimensions include technical strength, market influence, and innovation activity; the multi-source raw metrics include statistical metrics obtained from the standard data and graph feature metrics extracted from the industry knowledge graph, and the graph feature metrics include node centrality or PageRank value.
[0039] In the data collection stage, the system constructs a multi-source data foundation for the three core dimensions of technical strength, market influence, and innovation activity. Structured data such as "total number of patent applications" and "number of R & D personnel" are obtained from the enterprise information database through SQL queries, and relevant metrics are extracted from the patent database and financial reports by calling API interfaces. In addition, the system also uses the Cypher query language to extract graph feature metrics such as node centrality and PageRank value from the Neo4j graph database. These graph metrics can effectively reflect the structural position and influence of an enterprise in the industry network, providing multi-dimensional data support for subsequent analysis.
[0040] S302. For direct quantitative indicators, their values are obtained directly; for complex quantitative indicators, they are calculated using a preset rule model to obtain quantitative values; wherein, the complex quantitative indicators include at least a patent quality score, which is calculated by combining the basic score of a single patent, the weight of the number of citations, and the weight of the International Patent Classification number.
[0041] Entering the indicator quantification stage, the system adopts differentiated processing schemes for different types of indicators. For directly quantifiable indicators such as "operating revenue" and "number of patents," their raw values are used directly; for complex quantifiable indicators such as "patent quality score," a specially designed calculation model is used. This model first assigns a base score based on patent type, with invention patents receiving 10 points, utility models 5 points, and designs 2 points; then, it introduces a citation count weight, using a logarithmic function (log). 10 (C+1) is smoothed; finally, it is multiplied by the technical field weight coefficient based on IPC classification, and the quality score of each patent is obtained by weighted calculation.
[0042] S303. Use Min-Max standardization, Z-Score standardization, or quantile standardization to uniformly scale the values of various indicators with different dimensions to the same comparable scale.
[0043] In the data standardization process, the system automatically selects the most suitable standardization method based on the data distribution characteristics. For relatively uniformly distributed data, Min-Max standardization is used to linearly map it to the [0,1] interval; for data with outliers, Z-Score standardization is used to transform it based on the mean and standard deviation; for data that requires robust processing, quantile standardization is used to map it according to industry ranking position to ensure that each indicator is comparable.
[0044] S304. Assign weights to each standardized indicator under the same profile dimension, wherein the weights are determined by at least one of the following methods: expert scoring, entropy weighting, or AHP (Analytic Hierarchy Process).
[0045] The weight allocation process is implemented using the Analytic Hierarchy Process (AHP). The system constructs a hierarchical model containing an objective layer, a criterion layer, and a scheme layer. A judgment matrix is generated through pairwise comparisons by experts. The eigenvectors are solved using a numerical calculation library to determine the weights of each indicator. A rigorous consistency check is performed to ensure the rationality and scientific nature of the weight allocation.
[0046] S305. By weighted summing of the quantitative values of each indicator, the quantitative score of the industry entity in each profile dimension is obtained.
[0047] After completing the above preparations, the system calculates the scores for each dimension using the weighted summation formula P_i = Σ(W_j * I_j_normalized), where the weights W_j come from the AHP analysis results, and I_j_normalized are the standardized indicator values. This calculation process can comprehensively reflect the overall level of the enterprise across various dimensions.
[0048] S306. The quantitative score is compared with a preset threshold to generate a corresponding profile label; wherein the preset threshold is set by quantile method, absolute standard method or cluster analysis method; the profile label includes a single-dimensional level label and an overall qualitative label generated by combining multi-dimensional labels based on decision tree rules.
[0049] The generation of profile tags is achieved through a hierarchical tag system, which includes two levels: single-dimensional level tags and overall qualitative tags.
[0050] In the single-dimensional rank label generation stage, the system employs a dynamic threshold setting method based on quantiles. Specifically, the system first uses Python's pandas library to statistically analyze the quantitative scores of all entities within the industry in a specific dimension (such as technical strength). The 90th, 70th, and 30th quantile values are calculated using NumPy's percentile function as the critical points for rank division. Entities ranking in the top 10% are labeled "Leaders," those ranking between 11% and 30% are labeled "Advanced Players," those ranking between 31% and 70% are labeled "Followers," and the bottom 30% are labeled "Potential Stars." This process is automatically executed through a pre-defined threshold mapping function, which compares each entity's dimensional score with the calculated threshold range and outputs the corresponding rank label.
[0051] In the overall qualitative label generation stage, the system builds a decision tree reasoning mechanism based on the Drools rule engine. The rule engine loads a predefined set of business rules, and automatically triggers rule evaluation after the system completes the generation of all single-dimensional level labels. In specific implementation, the system is configured with multiple business rules, including: when the system detects that the entity's technical strength dimension label is "leader" and its market influence dimension label is also "leader," the rule engine executes the corresponding action, assigning the entity the overall qualitative label of "industry leader"; when the entity's innovation activity dimension label is "leader" and its market influence dimension label is "potential stock," the system automatically generates the overall qualitative label of "hidden champion."
[0052] To ensure the timeliness of the tagging system, a dynamic update mechanism has been established. A monthly recalculation of quantile thresholds is performed via a scheduled task configured through Apache Airflow, ensuring that the grading reflects the latest industry trends. Simultaneously, the business rules in the rules engine support online hot updates, allowing the tag generation logic to be adjusted promptly according to changes in business needs, guaranteeing the accuracy and usability of the profile tags.
[0053] In one embodiment of the present invention, based on step S4, the following will provide a possible embodiment and describe its specific implementation in a non-limiting manner.
[0054] Anomaly detection and root cause analysis are achieved through a complete technical loop. The system first uses Spark SQL window functions to intelligently group all entities based on an industry entity profile tagging system, classifying entities with the same or similar profile tags into the same analysis group. For example, the system executes the query `SELECT company_id FROM company_profile WHERE tech_capability = 'Leader' AND innovation_level = 'High'` to categorize all technology-leading and innovative companies into the "Technology Innovation Pioneer Group," and establishes an independent analysis unit for each group.
[0055] During the dynamic behavior baseline establishment phase, the system constructs adaptive benchmarks for the core indicators of each analysis group. The STL time series decomposition algorithm is used to perform in-depth analysis of the group's historical data, decomposing the indicator series into three components: trend, seasonality, and residuals. Based on the decomposition results, the system uses the Holt-Winters exponential smoothing model to predict the expected baseline value for the next 24 hours, while dynamically calculating the threshold range based on the standard deviation of historical residuals. For the more volatile startup group, the system automatically sets a wider threshold boundary (±3σ), while a stricter threshold standard (±1.5σ) is used for the more stable utility group to ensure the accuracy of anomaly detection.
[0056] Real-time monitoring is implemented using the Apache Flink stream processing engine. The system assigns an analysis group ID as a routing key to each data stream and calculates the Z-score deviation between the current metric value and the dynamic baseline of its group in real time. When the deviation exceeds a threshold for three consecutive time points, the system immediately generates a tiered alarm event and pushes the alarm information to downstream processing modules via a Kafka message queue. During this process, the system adjusts the alarm priority based on the entity's profile weight to ensure that anomalies from key supply chain enterprises are handled with priority.
[0057] Upon triggering an anomaly alarm, the system automatically initiates the root cause localization process. First, the Cypher graph query language is used to extract the related subgraphs of the target entity from the Neo4j graph database. The query statement is like MATCH(target:Company{id: $alert_company_id})-[r:SUPPLIES_TO|COMPETES_WITH*..2]-(related) RETURN target,r, related, which retrieves all related nodes within the second-degree range of the target entity. The system then uses batch graph traversal technology to examine the status indicators of these related nodes within 30 minutes before and after the anomaly time window, quickly identifying related nodes that exhibit synchronously abnormal fluctuations.
[0058] During the feature engineering phase, the system constructs an analysis dataset containing multi-dimensional features, including: profile label encoding of the target subject and related nodes, graph structure features (PageRank value, betweenness centrality, etc.), and the change in metrics for each node. These features are input into a pre-trained XGBoost model, which is trained by analyzing historical anomaly cases and can accurately calculate the contribution of each feature to the current anomaly. Finally, the system generates a detailed root cause localization report, listing not only the potential root cause nodes with the highest contribution but also providing complete association path analysis and confidence assessment, providing a reliable basis for decision-making. The entire processing can be completed in seconds, ensuring rapid response and accurate diagnosis of anomalies.
[0059] In one embodiment of the present invention, based on step S5, a possible embodiment will be given below, and its specific implementation will be described in a non-limiting manner.
[0060] First, the various analysis results are aggregated in multiple dimensions: based on the real-time anomaly detection results, the anomaly occurrence rate of each sub-industry is calculated by sliding time window calculation, and stratified statistics are performed according to the profile group dimension; at the same time, the high-frequency root cause node information in the root cause location report is integrated to construct an industry-level risk propagation chain model.
[0061] In the health status assessment phase, the system employs a weighted comprehensive index algorithm to generate an industry health score. This algorithm uses the industry anomaly incidence rate as the basic indicator, combined with weights based on the severity of anomalies (graded according to the magnitude of deviation) and key node impact factors (determined based on PageRank values from a knowledge graph) for comprehensive calculation. The system establishes a health baseline for each industry, triggering an industry health warning when the comprehensive score falls below the baseline by 20%. In practice, the system uses Python's pandas library for data aggregation and NumPy for matrix operations, ultimately generating a health analysis report that includes trend comparisons and dimensional breakdowns.
[0062] The risk situation awareness module is implemented through a three-layer analysis architecture. Short-term risk identification is based on real-time anomaly density changes output by the stream processing engine, using an exponentially weighted moving average algorithm to monitor abrupt changes in abnormal indicators. Mid-term risk warning analyzes frequently occurring risk patterns in root cause localization results, utilizing association rule mining algorithms to discover potential systemic risks. Long-term risk prediction combines the output of the predictive warning module, using an LSTM neural network model to deduce industry risk evolution trends. The system pays particular attention to state changes of key nodes in the knowledge graph; when multiple core nodes simultaneously exhibit anomalies, the industry risk level is automatically upgraded.
[0063] At the visualization level, the system utilizes an industry status dashboard developed based on the Spring Boot framework. The front-end uses ECharts to achieve multi-dimensional data visualization, including: displaying the health status distribution of various sub-sectors through heatmaps, illustrating the transmission path of risks in the industry chain using Sankey diagrams, and showcasing historical changes in industry health using trend curves. The system also establishes a tiered push mechanism; for identified major risks, it automatically pushes early warning information through platforms such as WeChat Work and DingTalk, ensuring that relevant personnel can obtain key status information in a timely manner.
[0064] To ensure the timeliness of the conclusions, the system schedules the entire analysis process through the Airflow workflow engine, performing a health status assessment every hour and updating the risk situation analysis every 6 hours. All analysis results are persisted to Elasticsearch, allowing users to search by keyword and review historical trends, providing continuous data support for industry regulation and investment decisions.
[0065] In some embodiments, the multi-source heterogeneous operation and maintenance data analysis system based on industry entity profiles may include multiple functional modules composed of computer program segments. The computer programs of each program segment in the multi-source heterogeneous operation and maintenance data analysis system based on industry entity profiles may be stored in the memory of a computer device and executed by at least one processor to perform (see details). Figure 1 (Description) Functionality for analyzing multi-source heterogeneous operation and maintenance data based on industry entity profiles.
[0066] In this embodiment, the multi-source heterogeneous operation and maintenance data analysis system based on industry entity profiles can be divided into multiple functional modules according to its functions, such as... Figure 2 As shown. The module referred to in this invention is a series of computer program segments that can be executed by at least one processor and perform a fixed function, and is stored in memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0067] The data acquisition module is used to collect multi-source heterogeneous data from multiple heterogeneous data sources from industry entities, and to clean and standardize the multi-source heterogeneous data to form standard data. The graph construction module is used to construct an industry knowledge graph based on the standard data, through entity recognition and relationship extraction, which includes industry subject nodes, knowledge element nodes and the relationships between them. The profile generation module is used to calculate a quantitative score for the industry entity in the multi-dimensional profile dimension based on the part of the standard data that is related to the target industry entity and to integrate the graph feature indicators of the entity extracted from the industry knowledge graph. The quantitative score is then mapped to profile tags according to a preset threshold to generate an industry entity profile. The subject analysis module is used to perform at least one of the following analyses on the industry subject based on the industry knowledge graph and the industry subject profile: anomaly detection, root cause localization, predictive early warning, or multi-dimensional dynamic analysis. The conclusion generation module is used to generate conclusions characterizing the health status and risk situation of the industry based on the obtained analysis results.
[0068] Figure 3 The multi-source heterogeneous operation and maintenance data analysis method based on industry subject profiling provided in this application embodiment can be applied to devices. Those skilled in the art will understand that the device structure involved in the embodiments of this invention does not constitute a limitation on the device. A device may include more or fewer components than illustrated, or combine certain components, or have different component arrangements. In the embodiments of this invention, the device includes, but is not limited to, laptop computers, desktop computers, workbenches, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the embodiments of this application described and / or claimed herein.
[0069] The device 300 may include a processor 310, a memory 320, and a communication unit 330. These components communicate via one or more buses. Those skilled in the art will understand that the server structure shown in the figure does not constitute a limitation of the present invention. It may be a bus topology or a star topology, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0070] The memory 320 can be used to store execution instructions of the processor 310. The memory 320 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. When the execution instructions in the memory 320 are executed by the processor 310, the device 300 is able to perform some or all of the steps in the above method embodiments.
[0071] The processor 310 serves as the control center of the storage device, connecting various parts of the electronic device via various interfaces and lines. It executes software programs and / or modules stored in the memory 320, and calls data stored in the memory to perform various functions of the electronic device and / or process data. The processor can be composed of integrated circuits (ICs), such as a single packaged IC or multiple packaged ICs with the same or different functions connected together. For example, the processor 310 may consist only of a central processing unit (CPU). In this embodiment of the invention, the CPU may have a single processing core or include multiple processing cores.
[0072] The communication unit 330 is used to establish a communication channel, enabling the storage device to communicate with other devices. It can receive user data sent by other devices or send user data to other devices.
[0073] The present invention also provides a computer medium, wherein the computer medium may store a program, which, when executed, may include some or all of the steps provided in the embodiments of the present invention. The medium may be a magnetic disk, an optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0074] Those skilled in the art will clearly understand that the techniques in the embodiments of the present invention can be implemented using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions in the embodiments of the present invention, or the parts that contribute to the prior art, can be embodied in the form of a software product. This computer software product is stored in a medium such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, or any other medium capable of storing program code. It includes several instructions to cause a computer device (which may be a personal computer, a server, or a second device, network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention.
[0075] The same or similar parts between the various embodiments in this specification can be referred to mutually. In particular, the device embodiments are basically similar to the method embodiments, so the description is relatively simple, and the relevant parts can be referred to the description in the method embodiments.
[0076] In the embodiments provided by this invention, it should be understood that the disclosed systems and methods can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between systems or modules may be electrical, mechanical, or other forms.
[0077] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical modules; that is, they may be located in one place or distributed across multiple network modules. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.
[0078] In addition, the functional modules in the various embodiments of the present invention can be integrated into one processing module, or each module can exist physically separately, or two or more modules can be integrated into one module.
[0079] Although the present invention has been described in detail with reference to the accompanying drawings and preferred embodiments, the present invention is not limited thereto. Various equivalent modifications or substitutions can be made to the embodiments of the present invention by those skilled in the art without departing from the spirit and essence of the invention, and such modifications or substitutions should all be within the scope of the present invention. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should also be covered within the protection scope of the present invention.
Claims
1. A multi-source heterogeneous operation and maintenance data analysis method based on industry principal portrait, characterized in that, include: Multi-source heterogeneous data from multiple heterogeneous data sources is collected from industry entities, and the multi-source heterogeneous data is cleaned and standardized to form standard data. Based on the standard data, an industry knowledge graph containing industry subject nodes, knowledge element nodes, and their relationships is constructed through entity recognition and relationship extraction. Based on the part of the standard data that is related to the target industry entity, and by integrating the graph feature indicators of the entity extracted from the industry knowledge graph, a quantitative score is calculated for the industry entity in the multi-dimensional profile dimension. The quantitative score is then mapped to a profile label according to a preset threshold to generate an industry entity profile. Based on the industry knowledge graph and the industry entity profile, at least one of the following analyses is performed on the industry entity: anomaly detection, root cause localization, predictive early warning, or multi-dimensional dynamic analysis. Based on the analysis results, conclusions characterizing the health status and risk situation of the industry are generated.
2. The method of claim 1, wherein, Collect multi-source heterogeneous data from multiple heterogeneous data sources from industry stakeholders, including: By configuring agents, API interfaces, and log collection tools, multi-source heterogeneous data can be collected in real time or offline from the operation and maintenance systems of any industry, such as finance, energy, or manufacturing. The multi-source heterogeneous data includes publicly available data from industry entities and internal platform behavioral data; the publicly available data includes at least one of business registration information, patent information, financial reports, or public opinion data; the internal platform behavioral data includes at least one of search records, transaction behavior, or content publishing data.
3. The method of claim 1, wherein, The multi-source heterogeneous data is cleaned and standardized, including: The collected operation and maintenance data is cleaned to remove noise, missing values and outliers, and the data with different formats is converted into a standardized format based on a predefined unified data model. By using host IP, device ID, or transaction serial number as a unified identifier, standardized data from different data sources are associated to form an information view with complete context. The associated data is stored in a unified time-series database.
4. The method of claim 1, wherein, Based on the aforementioned standard data, an industry knowledge graph is constructed through entity recognition and relationship extraction, comprising industry entity nodes, knowledge element nodes, and the relationships between them, including: Natural language processing technology is used to identify entities and extract relationships between entities from unstructured or semi-structured text in the standard data. Based on the entities and relationships, an industry knowledge graph is constructed with "enterprise", "product" and "technology" as nodes and "dependence" and "competition" as associations. Based on the dependence of industry entities in the knowledge graph and based on preset multi-level thresholds, industry entities are divided into different levels; wherein, the level includes at least a first-level entity with high dependence and a second-level entity with medium dependence. The industry knowledge graph is stored in a graph database and supports traversal and querying using a graph query language.
5. The method of claim 1, wherein, Based on the portion of the standard data relevant to the target industry entity, and integrating the graph feature indicators of the entity extracted from the industry knowledge graph, a quantitative score is calculated for the industry entity in the multi-dimensional profile dimension. This quantitative score is then mapped to profile tags according to a preset threshold to generate an industry entity profile, including: For each preset profile dimension, corresponding multi-source raw indicators are collected from the standard data and the industry knowledge graph; wherein, the profile dimension includes technical strength, market influence and innovation activity. The multi-source raw indicators include statistical indicators obtained from the standard data and graph feature indicators extracted from the industry knowledge graph, wherein the graph feature indicators include node centrality or PageRank value. For direct quantitative indicators, their values are obtained directly; for complex quantitative indicators, they are calculated using a preset rule model to obtain quantitative values; wherein, the complex quantitative indicators include at least a patent quality score, which is calculated by combining the basic score of a single patent, the weight of the number of citations, and the weight of the International Patent Classification number. Use Min-Max standardization, Z-Score standardization, or quantile standardization to uniformly scale the values of various indicators with different dimensions to the same comparable scale. Weights are assigned to each standardized indicator under the same profile dimension, and the weights are determined by at least one of the following methods: expert scoring, entropy weighting, or AHP (analytic hierarchy process). By weighted summing of the quantitative values of each indicator, the quantitative score of the industry entity in each profile dimension is obtained. The quantitative score is compared with a preset threshold to generate a corresponding profile label; wherein the preset threshold is set by quantile method, absolute standard method or cluster analysis method; the profile label includes a single-dimensional level label and an overall qualitative label generated by combining multi-dimensional labels based on decision tree rules.
6. The method of claim 5, wherein, The quantified score is compared with a preset threshold to generate a corresponding profile label, including: Single-dimensional level labeling: The quantitative score of a single profile dimension is compared with the level threshold preset for that dimension to generate a qualitative label describing the level of that dimension; wherein, the level threshold is set by the quantile method, including the threshold range for classifying industry entities into "leaders", "leaders", "followers" or "potential stocks". The overall qualitative label generated based on the combination of decision tree rules: Multiple single-dimensional level labels are taken as input, and logical judgments are performed through predefined decision tree rules to output a comprehensive overall qualitative label; wherein, the decision tree rules include at least: When an industry entity is rated as a "leader" in both the "technological strength" and "market influence" dimensions, it is given the overall qualitative label of "industry leader". When an industry entity is rated as a "leader" in the "innovation activity" dimension, but is rated as a "potential stock" in the "market influence" dimension, it is given the overall qualitative label of "hidden champion".
7. The method of claim 1, wherein, Based on the industry knowledge graph and the industry entity profile, at least one of the following analyses is performed on the industry entities: anomaly detection, root cause localization, predictive early warning, or multi-dimensional dynamic analysis, including: Based on the tags in the industry entity profile, the industry entities are divided into different analysis groups; Establish dynamic behavioral baselines based on historical core indicator data for each analysis group; The current indicator data of the industry entity is calculated in real time by the stream processing engine and compared with the dynamic behavior baseline of the analysis group. When the calculated deviation exceeds the threshold determined based on the historical volatility of the group, an anomaly alarm is triggered. In response to an anomaly alert for a target entity, a subgraph containing the target entity and its associated nodes is extracted from the industry knowledge graph using a graph query language. Examine the status indicators of each associated node in the subgraph near the abnormal time window, and filter out the associated nodes that have abnormal indicators at the same time. By combining the profile tags, graph feature attributes, and indicator changes of the target subject and its associated nodes, the contribution of each potential root cause is calculated through a machine learning model, and a root cause localization report is output.
8. An industry subject portrait-based multi-source heterogeneous operation and maintenance data analysis system, characterized in that, include: The data acquisition module is used to collect multi-source heterogeneous data from multiple heterogeneous data sources from industry entities, and to clean and standardize the multi-source heterogeneous data to form standard data. The graph construction module is used to construct an industry knowledge graph based on the standard data, through entity recognition and relationship extraction, which includes industry subject nodes, knowledge element nodes and the relationships between them. The profile generation module is used to calculate a quantitative score for the industry entity in the multi-dimensional profile dimension based on the part of the standard data that is related to the target industry entity and to integrate the graph feature indicators of the entity extracted from the industry knowledge graph. The quantitative score is then mapped to profile tags according to a preset threshold to generate an industry entity profile. The subject analysis module is used to perform at least one of the following analyses on the industry subject based on the industry knowledge graph and the industry subject profile: anomaly detection, root cause localization, predictive early warning, or multi-dimensional dynamic analysis. The conclusion generation module is used to generate conclusions characterizing the health status and risk situation of the industry based on the obtained analysis results.
9. An industry subject portrait-based multi-source heterogeneous operation and maintenance data analysis + device, characterized in that, include: The memory is used to store multi-source heterogeneous operation and maintenance data analysis programs based on industry entity profiles; The processor is configured to implement the steps of the multi-source heterogeneous operation and maintenance data analysis method based on industry entity profiles as described in any one of claims 1-7 when executing the multi-source heterogeneous operation and maintenance data analysis program based on industry entity profiles.
10. A computer readable medium having stored thereon a computer program, characterized in that, The readable medium stores a multi-source heterogeneous operation and maintenance data analysis program based on industry entity profiles. When the multi-source heterogeneous operation and maintenance data analysis program based on industry entity profiles is executed by the processor, it implements the steps of the multi-source heterogeneous operation and maintenance data analysis method based on industry entity profiles as described in any one of claims 1-7.