Intelligent data production method and system based on graph neural network and adaptive learning

By introducing intelligent data production methods based on graph neural networks and adaptive learning in the data production process, the defects of multimodal data processing, real-time and dynamic adaptability, data quality and generalization, and knowledge reasoning capabilities are solved, and efficient, real-time and flexible data processing and analysis are achieved.

CN119988647AInactive Publication Date: 2025-05-13AACAT TECHNOLOGY LTD

Patent Information

Application Number
CN202510465539.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing data production process has defects in multimodal data processing, real-time and dynamic adaptability, data quality and generalization, and knowledge reasoning capabilities. It is difficult to effectively process heterogeneous data, meet the needs of real-time streaming data analysis, automated verification and optimization, and complex logical reasoning.

Method used

Using an intelligent data production method based on graph neural networks and adaptive learning, a variety of data types are collected in real time through distributed crawler clusters and multi-protocol adaptation engines, a large language model and graph neural network are used to model cross-modal entity association relationships, dynamically optimize detection thresholds, and a spatio-temporal graph convolutional network is used to predict data trends, and a dynamic knowledge graph and relationship graph are constructed to achieve data fusion and task priority balance.

Benefits of technology

It realizes unified processing of multimodal data, improves real-time and dynamic adaptability, reduces false positive rates, enhances data quality and generalization capabilities, and supports complex logical reasoning and multi-industry customized services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988647A_ABST
    Figure CN119988647A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent data production method and system based on a graph neural network and adaptive learning, and the method comprises the steps: collecting a text, a time sequence, an image and sensor data in real time, and converting unstructured data into a structured JSON format through a format analysis tool; extracting text entity features based on a large language model, modeling a cross-modal entity association relationship in combination with a graph neural network, and fusing multi-source heterogeneous data through a dynamic knowledge graph; the detection threshold is dynamically optimized by adopting a reinforcement learning algorithm, and the false alarm rate is reduced in combination with a double-track verification mechanism; and predicting a data trend by using the space-time diagram convolutional network, and outputting an interpretability analysis report through the generative large model. According to the invention, a multi-protocol adaptation engine and a distributed stream processing architecture are adopted, and a protocol analysis layer automatically identifies a plurality of industrial protocols, so that the manual adaptation time is reduced; and real-time synchronization of production line-level data is realized through edge node parallel acquisition and compression transmission.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing technology, and in particular, to an intelligent data production method and system based on graph neural network and adaptive learning. Background Art

[0002] With the rapid development of big data and artificial intelligence technologies, the traditional data production process has gradually exposed the following core defects: 1. Rigid multimodal data processing: Existing systems rely on manual rules and static models (such as traditional ETL tools), which makes it difficult to uniformly process heterogeneous data such as text, images, and tables. For example, the correlation analysis between sensor time series data and equipment description documents in industrial scenarios requires manual writing of rules, which is inefficient and has poor scalability.

[0003] 2. Insufficient real-time and dynamic adaptability: Traditional batch processing architecture cannot meet the needs of real-time streaming data analysis (such as financial high-frequency trading monitoring), and the model is difficult to adapt to data distribution drift (such as changes in user behavior leading to prediction failure).

[0004] 3. The contradiction between quality and generalization: Data noise (such as sensor errors) and sparsity (such as rare medical cases) lead to model overfitting or failure, and existing technologies lack automated verification and continuous optimization mechanisms.

[0005] 4. Weak knowledge reasoning capabilities: Relying on predefined rule bases (such as expert systems), it is unable to automatically build dynamic knowledge graphs or implement complex logical reasoning (such as supply chain risk transmission analysis).

[0006] Patent application document CN102981485A discloses a method and system for processing real-time OEE data information of production line equipment operation. By collecting real-time data information of the production line, a recording system with production line equipment operation time and OEE real-time data information is established. In its real-time data information system, production shifts, equipment, products and other production factor information are collected and stored in real time. However, the real-time data collection system of the patent only supports structured data and cannot process unstructured formats such as PDF / HTML.

[0007] Patent application document CN116828001A discloses a smart factory production efficiency optimization system and method based on big data analysis. The system includes a data acquisition module, a data storage module, a data analysis module, a wireless communication module, an information management module, a data application module, and an early warning monitoring module. The distributed production line network slicing algorithm provides differentiated networks for different core businesses, realizes data diversion through a shared edge algorithm, improves the network security level through an intrusion protection algorithm, and predicts the best planning scheme based on real-time changes in market conditions through a commercial optimization model. However, although the distributed network slicing algorithm of this patent improves the collection efficiency, it lacks the ability to deeply mine semantic relationships. Summary of the invention

[0008] In view of the defects in the prior art, the purpose of the present invention is to provide an intelligent data production method and system based on graph neural network and adaptive learning.

[0009] The intelligent data production method based on graph neural network and adaptive learning provided by the present invention includes: Step 1: Use distributed crawler clusters, multi-protocol adaptation engines, and edge computing nodes to collect text, time series, image, and sensor data in real time, and use format parsing tools to convert unstructured data into structured JSON format; Step 2: Extract text entity features based on the large language model, combine graph neural network to model cross-modal entity associations, and fuse multi-source heterogeneous data through dynamic knowledge graphs; Step 3: Use reinforcement learning algorithm to dynamically optimize the detection threshold and combine the dual-track verification mechanism to reduce the false alarm rate; Step 4: Use the spatiotemporal graph convolutional network to predict data trends, and output an interpretable analysis report through a generative large model. At the same time, a standardized RESTful API interface is generated to support customized services for multiple industries. Step 5: Construct a device-material-order relationship graph through a spatiotemporal graph convolutional network to capture implicit dependency chains; combine equipment health prediction, demand prediction, and risk level classification to balance task priorities through dynamic weight allocation; and trigger an alert when the risk probability exceeds the threshold.

[0010] Preferably, the step 1 comprises: Build a distributed crawler cluster based on the Scrapy framework, configure a dynamic IP proxy pool and User-Agent rotation mechanism, simulate real user behavior to avoid anti-crawling detection; use Apache Tika to parse HTML / PDF page content, extract structured text fields, and store them in a MongoDB shard cluster; Collect sensor data streams in real time through Modbus / TCP and OPC-UA protocols, aggregate them into batch data by time window, and implement real-time streaming data processing through Apache Flink; Call the OpenCV library to normalize the image resolution and convert it to PNG format, while using a lossless compression algorithm to reduce storage overhead; The distributed crawler cluster is optimized in the following ways: Use Kubernetes for containerized deployment, supporting elastic expansion and contraction; Integrate Zstandard compression algorithm to compress transmission data in real time, with a compression ratio of 3:1; Data storage adopts a hot-cold tiering strategy, with hot data stored in SSDs and cold data migrated to the HDFS cluster.

[0011] Preferably, step 2 comprises: Extract entity-relationship triples from text data and calculate semantic similarity as the initial edge weight of the graph neural network; The GAT architecture is used to optimize node embedding, and the node feature update formula is:

[0012] in, Representation Node In the The embedding vector of the layer, is the activation function, It means summing up the neighboring nodes u of node v. represents the weight matrix of the lth layer, Represents the embedding vector of node u at layer l; Add and delete entity relationships based on real-time data streams to dynamically update the knowledge graph; The deployment optimization of graph neural networks includes: using the METIS algorithm to split the industrial equipment topology map into 8 sub-graphs according to node degree centrality, distributing them to 4 NVIDIA A100 GPUs, and achieving cross-card communication through the NCCL library; applying FP16 mixed precision training and combining gradient accumulation to reduce the peak memory occupancy; and automatically adjusting the batch size according to the GPU memory utilization, with the range set to 16-128.

[0013] Preferably, the step 3 comprises: A lightweight MobileNetV3 model is used to implement initial screening of image defects, and YOLOv5 is integrated to achieve multi-target detection; The classification threshold is dynamically adjusted based on the deep deterministic policy gradient algorithm. The state space includes the historical false alarm rate, noise level and real-time throughput. The reward function is designed as:

[0014] Among them, λ is the balance coefficient, TP is the true positive, FP is the false positive, and TN is the true negative; A web interface is provided for experts to annotate disputed data, and the annotation results are fed back to the detection model for incremental learning.

[0015] Preferably, step 4 comprises: Use a bidirectional LSTM model to analyze historical data, output future trend baselines, and integrate the Prophet model for seasonal correction; Integrate industry news, policy texts and social media public opinion data to generate analytical reports including visual charts; The prediction results are automatically compiled into RESTful API based on the Swagger framework, and support OAuth2.0 authentication and dynamic data format conversion.

[0016] The intelligent data production system based on graph neural network and adaptive learning provided by the present invention includes: Module M1: collects text, time series, image and sensor data in real time through distributed crawler clusters, multi-protocol adaptation engines and edge computing nodes, and converts unstructured data into structured JSON format using format parsing tools; Module M2: Extract text entity features based on a large language model, combine graph neural networks to model cross-modal entity associations, and fuse multi-source heterogeneous data through dynamic knowledge graphs; Module M3: Uses reinforcement learning algorithm to dynamically optimize detection thresholds and combines dual-track verification mechanism to reduce false alarm rate; Module M4: Use spatiotemporal graph convolutional networks to predict data trends, and output interpretable analysis reports through generative large models. At the same time, it generates standardized RESTful API interfaces to support customized services for multiple industries. Module M5: Construct the equipment-material-order relationship graph through the spatiotemporal graph convolutional network to capture the implicit dependency chain; combine equipment health prediction, demand prediction and risk level classification, balance task priority through dynamic weight allocation; trigger an early warning when the risk probability exceeds the threshold.

[0017] Preferably, the module M1 comprises: Build a distributed crawler cluster based on the Scrapy framework, configure a dynamic IP proxy pool and User-Agent rotation mechanism, simulate real user behavior to avoid anti-crawling detection; use Apache Tika to parse HTML / PDF page content, extract structured text fields, and store them in a MongoDB shard cluster; Collect sensor data streams in real time through Modbus / TCP and OPC-UA protocols, aggregate them into batch data by time window, and implement real-time streaming data processing through Apache Flink; Call the OpenCV library to normalize the image resolution and convert it to PNG format, while using a lossless compression algorithm to reduce storage overhead; The distributed crawler cluster is optimized in the following ways: Use Kubernetes for containerized deployment, supporting elastic expansion and contraction; Integrate Zstandard compression algorithm to compress transmission data in real time, with a compression ratio of 3:1; Data storage adopts a hot-cold tiering strategy, with hot data stored in SSDs and cold data migrated to the HDFS cluster.

[0018] Preferably, the module M2 comprises: Extract entity-relationship triples from text data and calculate semantic similarity as the initial edge weight of the graph neural network; The GAT architecture is used to optimize node embedding, and the node feature update formula is:

[0019] in, Representation Node In the The embedding vector of the layer, is the activation function, It means summing up the neighboring nodes u of node v. represents the weight matrix of the lth layer, Represents the embedding vector of node u at layer l; Add and delete entity relationships based on real-time data streams to dynamically update the knowledge graph; The deployment optimization of graph neural networks includes: using the METIS algorithm to split the industrial equipment topology map into 8 sub-graphs according to node degree centrality, distributing them to 4 NVIDIA A100 GPUs, and achieving cross-card communication through the NCCL library; applying FP16 mixed precision training and combining gradient accumulation to reduce the peak memory occupancy; and automatically adjusting the batch size according to the GPU memory utilization, with the range set to 16-128.

[0020] Preferably, the module M3 comprises: A lightweight MobileNetV3 model is used to implement initial screening of image defects, and YOLOv5 is integrated to achieve multi-target detection; The classification threshold is dynamically adjusted based on the deep deterministic policy gradient algorithm. The state space includes the historical false alarm rate, noise level and real-time throughput. The reward function is designed as:

[0021] Among them, λ is the balance coefficient, TP is the true positive, FP is the false positive, and TN is the true negative; A web interface is provided for experts to annotate disputed data, and the annotation results are fed back to the detection model for incremental learning.

[0022] Preferably, the module M4 comprises: Use a bidirectional LSTM model to analyze historical data, output future trend baselines, and integrate the Prophet model for seasonal correction; Integrate industry news, policy texts and social media public opinion data to generate analytical reports including visual charts; The prediction results are automatically compiled into RESTful API based on the Swagger framework, and support OAuth2.0 authentication and dynamic data format conversion.

[0023] Compared with the prior art, the present invention has the following beneficial effects: (1) Using a multi-protocol adaptation engine and distributed stream processing architecture (based on the Apache Flink framework), the protocol parsing layer automatically identifies more than 20 industrial protocols such as Modbus and OPC-UA, reducing manual adaptation time by 80%; through parallel collection and compression transmission of edge nodes (data throughput reaches 1.2TB / h), real-time synchronization of production line-level data is achieved; (2) Construct a BERT-GRU hybrid model optimized by cross-modal knowledge graph + prompt engineering. The knowledge graph dynamically integrates multi-source data such as equipment parameters (numeric type), maintenance records (text type), and vibration spectrum (time series type), and the entity relationship coverage is increased by 60%. The prompt engineering guides the large model to focus on key features in the field (such as "fault code: XJ203 → associated bearing temperature threshold > 85°C"), reducing semantic ambiguity errors by 47%; (3) A dual-track verification mechanism (CNN visual inspection + reinforcement learning dynamic threshold) is adopted. The primary detection layer uses lightweight MobileNetV3 to achieve 95% defect screening (FPS = 120); the reinforcement learning layer analyzes false alarm patterns (such as reflection misjudgment) in real time and dynamically adjusts the classification threshold, reducing the pass rate from 12% to 3%; (4) Using the spatiotemporal graph convolutional network (ST-GCN) + multi-task transfer learning framework, we constructed a three-dimensional relationship diagram of equipment, materials, and orders to capture implicit dependencies (such as the failure of machine tool A affecting five downstream assembly lines); and improved generalization capabilities through multi-task joint training such as equipment health prediction (MAE < 0.08) and demand prediction (RMSE reduced by 32%). BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Other features, objects and advantages of the present invention will become more apparent from the detailed description of non-limiting embodiments made with reference to the following drawings: Figure 1 This is the system architecture diagram; Figure 2 Flowchart for GNN semantic relationship mining; Figure 3 Schematic diagram of iterative optimization for adaptive learning mechanism; Figure 4 A case study on trend forecasting in the financial industry. DETAILED DESCRIPTION

[0025] The present invention is described in detail below in conjunction with specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but are not intended to limit the present invention in any form. It should be noted that, for those of ordinary skill in the art, several changes and improvements can also be made without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0026] Example 1 The present invention provides an intelligent data production system based on graph neural network and adaptive learning, comprising: 1. Basic collection layer: Supports fully automatic (Scrapy framework), semi-automatic (API call) and full-network crawling; Multi-format parsing module: Integrates Apache Tika to achieve unified conversion of PDF / HTML and other files into JSON format, increasing parsing speed by 40%.

[0027] 2. Intelligent data processing layer: Semantic understanding: Use a large language model (GPT-4) to extract text features and combine it with graph neural networks to model entity relationships; Industry knowledge fine-tuning: Through knowledge distillation loss function:

[0028] in, represents the knowledge distillation loss, x is the data point, D is the dataset, represents the Kullback-Leibler divergence, represents the probability distribution of the teacher model's output for input x, Represents the probability distribution of the student model output for input x; Inject medical ICD-10 terminology database; Relationship mining: GNN node embedding formula:

[0029] in, Representation Node In the The embedding vector of the layer, is the activation function, It means summing up the neighboring nodes u of node v. represents the weight matrix of the lth layer, Represents the embedding vector of node u at layer l; The accuracy of entity association calculation reaches 92%; Computational optimization: Using model parallel technology, GNN is deployed on a GPU cluster (NVIDIA A100), increasing the training speed by 3 times.

[0030] 3. Data quality detection layer: Adaptive learning mechanism: Dynamic threshold adjustment based on the DQN algorithm, the formula is:

[0031] in, represents the parameter vector at time step t+1, is the learning rate, is the objective function J in terms of parameters The gradient at .

[0032] Dual-track verification: Automatic detection (model prediction) and manual annotation platform are collaboratively optimized to reduce the false alarm rate by 15%.

[0033] 4. Data application layer: Trend prediction: Time series model (LSTM) and GPT-4 work together to generate industry analysis reports; Customized service: Provides RESTful API interface with response delay less than 200ms.

[0034] like Figure 1 , is a system architecture diagram, showing the interaction process of the basic collection layer, intelligent data processing layer, data quality detection layer and data application layer.

[0035] Key components: crawler scheduler, Apache Tika parsing engine, GNN relationship mining module, DQN optimizer.

[0036] Technical points: 1. The basic collection layer realizes multi-source heterogeneous data collection and format conversion through fully automatic crawlers and Apache Tika parsing engine; 2. The intelligent data processing layer integrates the large language model (GPT-4) and graph neural network (GNN) for semantic understanding and relationship mining; 3. The data quality detection layer dynamically optimizes the detection standards through an adaptive learning mechanism and collaborates with the manual annotation platform to improve accuracy; 4. The data application layer provides trend forecasting and API interfaces to support industry customized services.

[0037] like Figure 2 , is the GNN semantic relationship mining flowchart; Input: text entities (such as "drug A" and "side effect B"); Process: node feature embedding → edge weight calculation → output association score.

[0038] Technical points: 1. Input text entities (such as "drug A" and "side effect B") generate vector representations through node feature embedding; 2. The edge weight calculation module models entity associations based on cosine similarity or attention mechanism; 3. The graph neural network model iteratively optimizes the relevance score through multi-layer graph convolution.

[0039] like Figure 3 , which is a schematic diagram of iterative optimization of the adaptive learning mechanism; Quality inspection results → DQN reward calculation: ( ) →Threshold update.

[0040] in, is the reward value, is the balance detection accuracy, is the false alarm rate.

[0041] Technical points: 1. Dynamic threshold adjustment strategy based on reinforcement learning (DQN algorithm); 2. The reward function is optimized by balancing the detection accuracy (Precision) and the false positive rate (FalsePositiveRate); 3. The closed-loop feedback mechanism continuously improves quality inspection standards.

[0042] like Figure 4 , a case study on trend prediction in the financial industry; Input: Nasdaq index data for the past five years; Output: LSTM prediction curve and analysis report generated by GPT-4 (accuracy 89% vs 78% of ARIMA model).

[0043] Technical points: 1. LSTM model predicts future trends (such as the Nasdaq index); 2. GPT-4 generates an explainability analysis report based on the prediction results; 3. Output industry-customized decision-making recommendations (such as "increase holdings in the technology sector").

[0044] Basic collection layer Fully automatic crawling: 1. Use the Scrapy framework to build a distributed crawler, configure User-Agent rotation and IP proxy pool (number of requests per second ≤ 5, to avoid anti-crawling mechanism); 2. Data is stored in a MongoDB cluster, supporting horizontal expansion.

[0045] Format conversion: Parse PDF / HTML files through Apache Tika and output structured JSON data (fields include text content, metadata, source URL).

[0046] Intelligent data processing layer Knowledge fine-tuning steps: 1. Extract 100,000 terms from the medical ICD-10 database to build a knowledge base; 2. Use the PyTorch framework to implement knowledge distillation and fine-tune the GPT-4 model (learning rate 2e-5, batch size 32).

[0047] GNN Relationship Mining: Based on PyTorchGeometric implementation, the node feature dimension is 256 and the number of graph convolution layers is 3; Computing resources: NVIDIA A100 GPU × 4, training time 8 hours.

[0048] Data quality detection layer Dynamic Threshold Adjustment: State space: historical detection results (accuracy and false alarm rate of the past 100 data items); Action space: threshold adjustment range (±0.1~0.3); Reward function: .

[0049] Data application layer (financial case) Experimental setup: Data source: Yahoo Finance 2018-2023 Nasdaq 100 Index daily data (a total of 1,260 items); Comparison methods: ARIMA model, Prophet model; Evaluation indicators: root mean square error (RMSE), prediction direction accuracy.

[0050] result:

[0051] Example 2 The present invention also provides an intelligent data production system based on graph neural network and adaptive learning. The intelligent data production system based on graph neural network and adaptive learning can be realized by executing the process steps of the intelligent data production method based on graph neural network and adaptive learning. That is, those skilled in the art can understand the intelligent data production method based on graph neural network and adaptive learning as a preferred implementation of the intelligent data production system based on graph neural network and adaptive learning.

[0052] According to the present invention, the intelligent data production system based on graph neural network and adaptive learning includes: module M1: through distributed crawler clusters, multi-protocol adaptation engines and edge computing nodes, text, time series, image and sensor data are collected in real time, and the format parsing tool is used to convert unstructured data into structured JSON format; module M2: based on the large language model, text entity features are extracted, cross-modal entity association relationships are modeled in combination with graph neural network, and multi-source heterogeneous data are integrated through dynamic knowledge graph; module M3: reinforcement learning algorithm is used to dynamically optimize the detection threshold, and the false alarm rate is reduced in combination with the dual-track verification mechanism; module M4: spatiotemporal graph convolutional network is used to predict data trends, and an interpretable analysis report is output through a generative large model, and a standardized RESTful API interface is generated at the same time to support customized services for multiple industries; module M5: a device-material-order relationship diagram is constructed through a spatiotemporal graph convolutional network to capture implicit dependency chains; combined with equipment health prediction, demand prediction and risk level classification, task priorities are balanced through dynamic weight allocation; when the risk probability exceeds the threshold, an early warning is triggered.

[0053] The module M1 includes: building a distributed crawler cluster based on the Scrapy framework, configuring a dynamic IP proxy pool and a User-Agent rotation mechanism, simulating real user behavior to avoid anti-crawling detection; parsing HTML / PDF page content through Apache Tika, extracting structured text fields, and storing them in a MongoDB shard cluster; collecting sensor data streams in real time through Modbus / TCP and OPC-UA protocols, aggregating them into batch data according to time windows, and realizing real-time processing of streaming data through Apache Flink; calling the OpenCV library to normalize the image resolution and convert it into PNG format, while using a lossless compression algorithm to reduce storage overhead; the distributed crawler cluster is optimized in the following ways: using Kubernetes for containerized deployment to support elastic expansion and contraction; integrating the Zstandard compression algorithm to compress the transmitted data in real time with a compression ratio of 3:1; data storage adopts a hot and cold tiering strategy, with hot data stored in SSD and cold data migrated to the HDFS cluster.

[0054] The module M2 includes: extracting entity-relationship triples from text data and calculating semantic similarity as the initial edge weight of the graph neural network; using the GAT architecture to optimize node embedding, and the node feature update formula is: ,in, Representation Node In the The embedding vector of the layer, is the activation function, It means summing up the neighboring nodes u of node v. represents the weight matrix of the lth layer, Represents the embedding vector of node u in the lth layer; adds and deletes entity relationships according to real-time data streams, and updates the knowledge graph dynamically; the deployment optimization of the graph neural network includes: using the METIS algorithm to split the industrial equipment topology map into 8 subgraphs according to node degree centrality, and distribute them to 4 NVIDIA A100 GPUs, and realize cross-card communication through the NCCL library; applying FP16 mixed precision training, combined with gradient accumulation to reduce the peak memory usage; automatically adjusting the batch size according to the GPU memory utilization, and setting the range to 16-128.

[0055] The module M3 includes: using a lightweight MobileNetV3 model to implement initial screening of image defects, and integrating YOLOv5 to implement multi-target detection; dynamically adjusting the classification threshold based on a deep deterministic policy gradient algorithm, the state space includes historical false alarm rate, noise level and real-time throughput, and the reward function is designed as: , where λ is the balance coefficient, TP is the true positive, FP is the false positive, and TN is the true negative; a web interface is provided for experts to annotate disputed data, and the annotation results are fed back to the detection model for incremental learning.

[0056] The module M4 includes: using a bidirectional LSTM model to analyze historical data, outputting a future trend baseline, and integrating a Prophet model for seasonal correction; integrating industry news, policy texts, and social media public opinion data to generate an analysis report containing visual charts; automatically compiling prediction results into a RESTful API based on the Swagger framework, and supporting OAuth2.0 authentication and dynamic conversion of data formats.

[0057] Those skilled in the art know that, in addition to implementing the system, device and its various modules provided by the present invention in a purely computer-readable program code, it is entirely possible to implement the same program in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers and embedded microcontrollers by logically programming the method steps. Therefore, the system, device and its various modules provided by the present invention can be considered as a hardware component, and the modules included therein for implementing various programs can also be considered as structures within the hardware component; the modules for implementing various functions can also be considered as both software programs for implementing the method and structures within the hardware component.

[0058] The above describes the specific embodiments of the present invention. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. In the absence of conflict, the embodiments of the present application and the features in the embodiments can be combined with each other arbitrarily.

Claims

1. An intelligent data production method based on graph neural network and adaptive learning, characterized in that: include: Step 1: Use distributed crawler clusters, multi-protocol adaptation engines, and edge computing nodes to collect text, time series, image, and sensor data in real time, and use format parsing tools to convert unstructured data into structured JSON format; Step 2: Extract text entity features based on the large language model, combine graph neural network to model cross-modal entity associations, and fuse multi-source heterogeneous data through dynamic knowledge graphs; Step 3: Use reinforcement learning algorithm to dynamically optimize the detection threshold and combine the dual-track verification mechanism to reduce the false alarm rate; Step 4: Use the spatiotemporal graph convolutional network to predict data trends, and output an interpretable analysis report through a generative large model. At the same time, a standardized RESTful API interface is generated to support customized services for multiple industries. Step 5: Construct a device-material-order relationship graph through a spatiotemporal graph convolutional network to capture implicit dependency chains; combine equipment health prediction, demand prediction, and risk level classification to balance task priorities through dynamic weight allocation; and trigger an alert when the risk probability exceeds the threshold.

2. The intelligent data production method based on graph neural network and adaptive learning according to claim 1 is characterized in that: The step 1 comprises: Build a distributed crawler cluster based on the Scrapy framework, configure a dynamic IP proxy pool and User-Agent rotation mechanism, simulate real user behavior to avoid anti-crawling detection; use Apache Tika to parse HTML / PDF page content, extract structured text fields, and store them in a MongoDB shard cluster; Collect sensor data streams in real time through Modbus / TCP and OPC-UA protocols, aggregate them into batch data by time window, and implement real-time streaming data processing through Apache Flink; Call the OpenCV library to normalize the image resolution and convert it to PNG format, while using a lossless compression algorithm to reduce storage overhead; The distributed crawler cluster is optimized in the following ways: Use Kubernetes for containerized deployment, supporting elastic expansion and contraction; Integrate Zstandard compression algorithm to compress transmission data in real time, with a compression ratio of 3:1; Data storage adopts a hot-cold tiering strategy, with hot data stored in SSDs and cold data migrated to the HDFS cluster.

3. The intelligent data production method based on graph neural network and adaptive learning according to claim 1 is characterized in that: The step 2 comprises: Extract entity-relationship triples from text data and calculate semantic similarity as the initial edge weight of the graph neural network; The GAT architecture is used to optimize node embedding, and the node feature update formula is: in, Representation Node In the The embedding vector of the layer, is the activation function, It means summing up the neighboring nodes u of node v. represents the weight matrix of the lth layer, Represents the embedding vector of node u at layer l; Add and delete entity relationships based on real-time data streams to dynamically update the knowledge graph; The deployment optimization of graph neural networks includes: using the METIS algorithm to split the industrial equipment topology map into 8 sub-graphs according to node degree centrality, distributing them to 4 NVIDIA A100 GPUs, and achieving cross-card communication through the NCCL library; applying FP16 mixed precision training and combining gradient accumulation to reduce the peak memory occupancy; and automatically adjusting the batch size according to the GPU memory utilization, with the range set to 16-128.

4. The intelligent data production method based on graph neural network and adaptive learning according to claim 1 is characterized in that: The step 3 comprises: A lightweight MobileNetV3 model is used to implement initial screening of image defects, and YOLOv5 is integrated to achieve multi-target detection; The classification threshold is dynamically adjusted based on the deep deterministic policy gradient algorithm. The state space includes the historical false alarm rate, noise level and real-time throughput. The reward function is designed as: Among them, λ is the balance coefficient, TP is the true positive, FP is the false positive, and TN is the true negative; A web interface is provided for experts to annotate disputed data, and the annotation results are fed back to the detection model for incremental learning.

5. The intelligent data production method based on graph neural network and adaptive learning according to claim 1 is characterized in that: The step 4 comprises: Use a bidirectional LSTM model to analyze historical data, output future trend baselines, and integrate the Prophet model for seasonal correction; Integrate industry news, policy texts and social media public opinion data to generate analytical reports including visual charts; The prediction results are automatically compiled into RESTful API based on the Swagger framework, and support OAuth2.0 authentication and dynamic data format conversion.

6. An intelligent data production system based on graph neural network and adaptive learning, characterized in that: include: Module M1: collects text, time series, image and sensor data in real time through distributed crawler clusters, multi-protocol adaptation engines and edge computing nodes, and converts unstructured data into structured JSON format using format parsing tools; Module M2: Extract text entity features based on a large language model, combine graph neural networks to model cross-modal entity associations, and fuse multi-source heterogeneous data through dynamic knowledge graphs; Module M3: Uses reinforcement learning algorithm to dynamically optimize detection thresholds and combines dual-track verification mechanism to reduce false alarm rate; Module M4: Use spatiotemporal graph convolutional networks to predict data trends, and output interpretable analysis reports through generative large models. At the same time, it generates standardized RESTful API interfaces to support customized services for multiple industries. Module M5: Construct the equipment-material-order relationship graph through the spatiotemporal graph convolutional network to capture the implicit dependency chain; combine equipment health prediction, demand prediction and risk level classification, balance task priority through dynamic weight allocation; trigger an early warning when the risk probability exceeds the threshold.

7. The intelligent data production system based on graph neural network and adaptive learning according to claim 6, characterized in that: The module M1 comprises: Build a distributed crawler cluster based on the Scrapy framework, configure a dynamic IP proxy pool and User-Agent rotation mechanism, simulate real user behavior to avoid anti-crawling detection; use Apache Tika to parse HTML / PDF page content, extract structured text fields, and store them in a MongoDB shard cluster; Collect sensor data streams in real time through Modbus / TCP and OPC-UA protocols, aggregate them into batch data by time window, and implement real-time streaming data processing through Apache Flink; Call the OpenCV library to normalize the image resolution and convert it to PNG format, while using a lossless compression algorithm to reduce storage overhead; The distributed crawler cluster is optimized in the following ways: Use Kubernetes for containerized deployment, supporting elastic expansion and contraction; Integrate Zstandard compression algorithm to compress transmission data in real time, with a compression ratio of 3:1; Data storage adopts a hot-cold tiering strategy, with hot data stored in SSDs and cold data migrated to the HDFS cluster.

8. The intelligent data production system based on graph neural network and adaptive learning according to claim 6, characterized in that: The module M2 comprises: Extract entity-relationship triples from text data and calculate semantic similarity as the initial edge weight of the graph neural network; The GAT architecture is used to optimize node embedding, and the node feature update formula is: in, Representation Node In the The embedding vector of the layer, is the activation function, It means summing up the neighboring nodes u of node v. represents the weight matrix of the lth layer, Represents the embedding vector of node u at layer l; Add and delete entity relationships based on real-time data streams to dynamically update the knowledge graph; The deployment optimization of graph neural networks includes: using the METIS algorithm to split the industrial equipment topology map into 8 sub-graphs according to node degree centrality, distributing them to 4 NVIDIA A100 GPUs, and achieving cross-card communication through the NCCL library; applying FP16 mixed precision training and combining gradient accumulation to reduce the peak memory occupancy; and automatically adjusting the batch size according to the GPU memory utilization, with the range set to 16-128.

9. The intelligent data production system based on graph neural network and adaptive learning according to claim 6, characterized in that: The module M3 comprises: A lightweight MobileNetV3 model is used to implement initial screening of image defects, and YOLOv5 is integrated to achieve multi-target detection; The classification threshold is dynamically adjusted based on the deep deterministic policy gradient algorithm. The state space includes the historical false alarm rate, noise level and real-time throughput. The reward function is designed as: Among them, λ is the balance coefficient, TP is the true positive, FP is the false positive, and TN is the true negative; A web interface is provided for experts to annotate disputed data, and the annotation results are fed back to the detection model for incremental learning.

10. The intelligent data production system based on graph neural network and adaptive learning according to claim 6, characterized in that: The module M4 comprises: Use a bidirectional LSTM model to analyze historical data, output future trend baselines, and integrate the Prophet model for seasonal correction; Integrate industry news, policy texts and social media public opinion data to generate analytical reports including visual charts; The prediction results are automatically compiled into RESTful API based on the Swagger framework, and support OAuth2.0 authentication and dynamic data format conversion.

Citation Information

Patent Citations

  • Real-time OEE (Overall Equipment Effectiveness) data information processing method and system for equipment operation of production line

    CN102981485A

  • Smart factory production efficiency optimization system and method based on big data analysis

    CN116828001A

  • Network security knowledge graph generation method based on threat intelligence

    CN113282759A

  • An intelligent prediction system for social public opinion risk

    CN119760361A

Cited By

  • Method and device for generating enterprise health degree analysis report and storage medium

    CN120317231A

  • Marine oil and gas equipment data monitoring method and system based on enhanced graph learning

    CN120492825A

  • An ocean oil and gas equipment data monitoring method and system based on enhanced graph learning

    CN120492825B

  • Industrial chain knowledge graph construction method and system

    CN120525062A

  • Preplan digitization method and system based on large model technology

    CN120632112A