Multi-source data automatic acquisition method based on large language model

Through the multi-source data automation acquisition method based on large language models, the optimal acquisition strategy and real-time abnormal detection are dynamically generated, which solves the shortcomings of multi-source data acquisition technology in protocol adaptation and strategy optimization, and realizes efficient and reliable data acquisition and processing, and provides knowledge-enhanced data assets.

CN120450015AActive Publication Date: 2025-08-08SHENZHEN WEIPINZHIYUAN INFORMATION TECH CO LTD

Patent Information

Application Number
CN202510829167.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-20
Publication Date
2025-08-08
Estimated Expiration
2045-06-20

AI Technical Summary

Technical Problem

The existing multi-source data acquisition technology cannot perform protocol adaptation and adaptive strategy optimization, resulting in poor scalability when facing the rapid iteration of new IoT protocols or protocol versions, unable to meet dynamically changing business needs, and there are problems such as leakage of high-value data and redundant collection of low-value data.

Method used

Using a multi-source data automation acquisition method based on a large language model, by analyzing the feature distribution of historically collected data and the current data source status, dynamically generate the optimal acquisition strategy, schedule the acquisition tasks in real time, and perform real-time abnormality detection and adaptive compensation, build a domain knowledge graph, identify protocol features for analysis and semantic noise filtering, and realize the acquisition of protocol-independent standardized data.

Benefits of technology

It realizes efficient and reliable acquisition and processing of multi-source data in a dynamic environment, overcomes the shortcomings of protocol adaptation and policy optimization, ensures the continuity and accuracy of data acquisition, and provides knowledge-enhanced data assets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120450015A_ABST
    Figure CN120450015A_ABST
Patent Text Reader

Abstract

The multi-source data automatic collection method based on the large language model comprises the steps that feature distribution of historically collected data is analyzed, and an optimal collection strategy is dynamically generated in combination with the state of a current data source; scheduling the acquisition task in real time to obtain a dynamically scheduled data stream; performing real-time anomaly detection and self-adaptive compensation on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream; analyzing entity association and service logic of the self-fault-tolerant data stream, and predicting potential data demand points; automatically expanding a data acquisition range to obtain knowledge-enhanced data assets; identifying protocol features of the data transmission stream of the data assets, and analyzing the data by an analyzer based on protocol feature mapping to obtain protocol-independent standardized data; and performing semantic noise filtering and cross-modal cleaning to obtain semantic pure data. According to the invention, the defect that protocol adaptation and adaptive strategy optimization cannot be carried out in the existing multi-source data acquisition technology is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method for automatically collecting multi-source data based on a large language model. Background Art

[0002] With the deep integration of big data and artificial intelligence, multi-source data collection has become a critical component of digital transformation across various industries. Currently, fields such as the Industrial Internet of Things, smart cities, and financial risk management all face the need to collaboratively collect data from massive heterogeneous data sources (such as sensors, databases, APIs, and web pages). However, existing multi-source data collection technologies suffer from the following significant drawbacks: First, protocol adaptability is limited. Traditional data collection relies on pre-programmed protocol parsing rules (e.g., fixed parsing modules for HTTP and Modbus protocols). This makes it difficult to adapt to new IoT protocols (e.g., LoRaWAN and custom industrial protocols) or scenarios with rapidly evolving protocol versions. When unknown protocols are encountered, significant manpower is required for protocol reverse analysis and parser development, resulting in poor scalability of the collection system and an inability to meet dynamically changing business needs.

[0003] Second, static collection strategies. Most collection systems use a fixed collection pattern (e.g., collecting log data once an hour), and are unable to dynamically adjust strategies based on data value density, data source load, and other factors. This results in both missed high-value data and redundant collection of low-value data, leading to the dual problems of wasted computing resources and missing critical information.

[0004] In summary, the existing multi-source data collection technology has obvious deficiencies in protocol adaptation, policy optimization, etc., and cannot meet the data collection needs in complex business scenarios. Summary of the Invention

[0005] The main purpose of the present invention is to provide a multi-source data automatic collection method based on a large language model, aiming to overcome the defects of existing multi-source data collection technology that is unable to perform protocol adaptation and adaptive strategy optimization.

[0006] To achieve the above objectives, the present invention provides a method for automatically collecting multi-source data based on a large language model, comprising the following steps: Analyze the characteristic distribution of historically collected data and dynamically generate the optimal collection strategy based on the current state of the data source; Based on the optimal acquisition strategy, real-time scheduling of acquisition tasks is performed to obtain a dynamically scheduled data stream; real-time anomaly detection and adaptive compensation are performed on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream; Building a domain knowledge graph based on a large language model, analyzing the entity associations and business logic of the self-fault-tolerant data stream, and predicting potential data demand points; based on the predicted potential data demand points, automatically expanding the data collection scope to obtain knowledge-enhanced data assets; Identify the protocol features of the data transmission stream of the data asset, parse the data based on the parser of the protocol feature mapping to obtain protocol-independent standardized data; perform semantic noise filtering and cross-modal cleaning on the protocol-independent standardized data to obtain semantically pure data.

[0007] Furthermore, the protocol-independent standardized data is data that eliminates the communication protocol characteristics that the original data relies on and is uniformly converted into an independent data format.

[0008] Furthermore, the acquisition strategy includes sampling frequency, batch size and priority sorting.

[0009] Furthermore, we analyze the characteristic distribution of historically collected data and, combined with the current state of the data source, dynamically generate the optimal collection strategy, including: Slice the historically collected data and extract the data volume fluctuation characteristics, value density distribution, and field integrity characteristics of each slice to obtain the data feature vector; Monitor the connection success rate, response timeliness, and server load pressure of each data source, and conduct hierarchical evaluation to obtain a data source status matrix; A reinforcement learning model is constructed, and multi-objective reinforcement learning optimization processing is performed on the data feature vector and the data source state matrix to generate an optimal parameter vector, and an optimal acquisition strategy is obtained based on the optimal parameter vector.

[0010] Furthermore, a domain knowledge graph is constructed based on the large language model to analyze the entity associations and business logic of the self-fault-tolerant data flow and predict potential data demand points, including: Identifying entities in the self-fault-tolerant data stream, building semantic associations between entities, and generating a set of entity relationship triples; Build a domain knowledge graph based on a large language model, map the entity relationship triple set to the domain knowledge graph, calculate the similarity between entities, and complete the implicit relationship to obtain a dynamic knowledge graph; Analyze the association paths of entities in the dynamic knowledge graph, extract business rules, identify event sequence patterns through time series analysis, and generate a business logic rule base; Through the Transformer-based demand forecasting model, the business logic rule base and dynamic knowledge graph are analyzed to identify data missing points and business association gaps, and output potential data demand points.

[0011] Furthermore, real-time anomaly detection and adaptive compensation are performed on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream, including: Slicing the dynamically scheduled data stream, extracting numerical features, temporal features, and semantic features, and constructing a feature vector set; Performing anomaly detection based on the feature vector set to generate an abnormal event including an abnormality type and a confidence level; Map abnormal events to a pre-built fault knowledge graph, analyze the abnormal propagation path through a graph neural network, locate the root cause of the problem, and generate a diagnostic report; An optimal compensation strategy is automatically matched based on the diagnosis report, and parallel compensation is performed on the dynamically scheduled data stream according to the optimal compensation strategy to obtain a self-fault-tolerant data stream.

[0012] Furthermore, the protocol characteristics of the data transmission flow of the data asset are identified, and the data is parsed by a parser based on the protocol characteristic mapping to obtain protocol-independent standardized data, including: Using a finite state automaton to scan the data transmission stream of the data asset bit by bit, and identifying the protocol type and data structure characteristics according to preset protocol state transition rules; The parser corresponding to the protocol type and data structure characteristics is retrieved from the parser rule library, and the data is decapsulated and field mapped to obtain protocol-independent standardized data including data source labels, timestamps and business fields.

[0013] Furthermore, the protocol-independent standardized data is subjected to semantic noise filtering and cross-modal cleaning to obtain semantically pure data, including: Extracting features from the protocol-independent standardized data to obtain a neural morphological feature map, including spike timing features, synaptic weight features, and population coding features; performing noise filtering on the neuromorphic feature map to obtain an enhanced feature map; Performing entity parsing, event extraction, and causal reasoning on the enhanced feature graph to construct a cognitive graph; The cognitive graph is subjected to cross-modal semantic alignment and a meta-learning-driven cleaning process to obtain semantically pure data.

[0014] The present invention also provides a multi-source data automatic acquisition device based on a large language model, comprising: The analysis unit is used to analyze the characteristic distribution of historically collected data and dynamically generate the optimal collection strategy based on the current state of the data source; A compensation unit is configured to schedule acquisition tasks in real time based on the optimal acquisition strategy to obtain a dynamically scheduled data stream; perform real-time anomaly detection and adaptive compensation on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream; An expansion unit is used to build a domain knowledge graph based on a large language model, analyze the entity associations and business logic of the self-fault-tolerant data stream, and predict potential data demand points; based on the predicted potential data demand points, automatically expand the data collection scope to obtain knowledge-enhanced data assets; A cleaning unit is used to identify the protocol characteristics of the data transmission flow of the data asset, parse the data based on the parser of the protocol feature mapping to obtain protocol-independent standardized data; and perform semantic noise filtering and cross-modal cleaning on the protocol-independent standardized data to obtain semantically pure data.

[0015] The present invention also provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of any one of the above methods when executing the computer program.

[0016] The present invention also provides a computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the steps of any of the above methods are implemented.

[0017] The present invention provides a multi-source data automatic collection method based on a large language model, comprising: analyzing the characteristic distribution of historical collection data, combining the current state of the data source, and dynamically generating an optimal collection strategy; scheduling collection tasks in real time based on the optimal collection strategy to obtain a dynamically scheduled data stream; performing real-time anomaly detection and adaptive compensation on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream; constructing a domain knowledge graph based on the large language model, analyzing the entity association and business logic of the self-fault-tolerant data stream, and predicting potential data demand points; automatically expanding the data collection scope based on the predicted potential data demand points to obtain knowledge-enhanced data assets; identifying the protocol characteristics of the data transmission stream of the data asset, parsing the data based on a parser of protocol feature mapping to obtain protocol-independent standardized data; performing semantic noise filtering and cross-modal cleaning on the protocol-independent standardized data to obtain semantically pure data. In the present invention, by analyzing the characteristic distribution of historical collection data, combining the current state of the data source, dynamically generating an optimal collection strategy, and at the same time, predicting potential data demand points and automatically expanding the data collection scope; then parsing the data based on a parser of protocol feature mapping to obtain protocol-independent standardized data; overcoming the defects of existing multi-source data collection technologies that cannot perform protocol adaptation and adaptive strategy optimization. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 This is a schematic diagram of the steps of a multi-source data automatic collection method based on a large language model in one embodiment of the present invention; Figure 2This is a structural block diagram of a multi-source data automatic acquisition device based on a large language model in one embodiment of the present invention; Figure 3 It is a schematic block diagram of the structure of a computer device according to an embodiment of the present invention.

[0019] The implementation, functional features and advantages of the present invention will be further described with reference to the accompanying drawings in conjunction with embodiments. DETAILED DESCRIPTION

[0020] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0021] Reference Figure 1 In one embodiment of the present invention, a method for automatically collecting multi-source data based on a large language model is provided, comprising the following steps: Step S1: Analyze the characteristic distribution of historically collected data and dynamically generate the optimal collection strategy based on the current state of the data source; Step S2, scheduling the acquisition task in real time based on the optimal acquisition strategy to obtain a dynamically scheduled data stream; performing real-time anomaly detection and adaptive compensation on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream; Step S3: constructing a domain knowledge graph based on the large language model, analyzing the entity associations and business logic of the self-fault-tolerant data stream, and predicting potential data demand points; based on the predicted potential data demand points, automatically expanding the data collection scope to obtain knowledge-enhanced data assets; Step S4: Identify the protocol features of the data transmission stream of the data asset, parse the data based on the parser of the protocol feature mapping to obtain protocol-independent standardized data; perform semantic noise filtering and cross-modal cleaning on the protocol-independent standardized data to obtain semantically pure data.

[0022] In this embodiment, as described in step S1 above, during the multi-source data collection process, historically collected data contains key information such as data volume fluctuation patterns and data value distribution, while the current data source status (e.g., server load and network latency) reflects the real-time collection environment. This step first conducts in-depth analysis of the historically collected data, employing time series analysis to extract characteristics such as data volume fluctuations over different time periods (e.g., differences between weekdays and holidays) and the frequency distribution of high-value data. Simultaneously, a real-time monitoring module is used to obtain status parameters such as response time, connection success rate, and server resource utilization for each data source. Then, using the historical data characteristics and current data source status information as input, the system leverages the data analysis and decision-making capabilities of a large language model, combined with a reinforcement learning algorithm, to dynamically generate optimal collection strategies for different data sources with the optimization goals of maximizing collection efficiency, minimizing resource consumption, and ensuring data integrity. This includes parameter configurations such as sampling frequency (e.g., high-frequency collection for real-time data streams and low-frequency batch collection for historical logs) and collection priority (e.g., prioritizing alarm data).

[0023] As described in step S2 above, based on the optimal collection strategy generated in step S1, the real-time scheduling module precisely allocates collection tasks to the appropriate computing resource nodes (e.g., in-memory computing nodes for processing real-time data and distributed storage clusters for processing batch data) according to priority and time schedule. This ensures efficient execution of collection tasks, thereby forming a dynamically scheduled data stream. To ensure the stability and reliability of the data stream, a real-time anomaly detection mechanism is deployed. By building a Transformer-based time series analysis model and semantic recognition model, the dynamically scheduled data stream is monitored in real time. For numerical data, the mechanism analyzes its changing trends and fluctuation ranges to detect abnormal jumps, missing values, and other issues. For textual data, the mechanism identifies noise such as semantic contradictions and incorrect representations. Once an anomaly is detected, an adaptive compensation mechanism is immediately activated. Based on the anomaly type (e.g., data source failure, network interruption, data format error), appropriate compensation strategies are automatically implemented. These strategies include switching to an alternative data source, adjusting data request parameters, and generating data repair rules using a large language model. This eliminates the impact of the anomaly, resulting in a self-resilient data stream and ensuring the continuity and accuracy of data collection.

[0024] As described in step S3 above, leveraging the powerful knowledge understanding and graph-building capabilities of the large language model, the self-fault-tolerant data stream generated in step S2 is deeply analyzed. First, various entities (such as users, devices, and orders) are automatically extracted from the data stream. Semantic association analysis is used to identify relationships between entities (e.g., "user - ownership - device," "order - association - product"), thereby constructing a domain knowledge graph. Next, graph analysis algorithms are used to mine the knowledge graph, identifying potential patterns and business logic rules in the entity associations within the data stream (e.g., the correlation between the order volume of a certain product type and the user's region within a specific time period). Based on these discovered business logic and historical data patterns, combined with the predictive capabilities of the large language model, forward-looking predictions are made for potential future data needs (e.g., predicting the need for new user attribute collection based on user growth trends). Finally, based on the prediction results, the data collection scope expansion process is automatically triggered, including operations such as configuring new data source access, adjusting web crawler rules, and modifying API call parameters. This incorporates the new data into the collection scope, forming a knowledge-enhanced data asset that provides richer and more valuable data support for subsequent data analysis and decision-making.

[0025] As described in step S4 above, since data assets originate from a variety of different data sources, the protocols (such as HTTP, MQTT, Modbus) and data formats (JSON, XML, binary) used for their data transmission vary, which makes unified data processing and analysis difficult. This step first identifies the protocol features of the data transmission stream in the data asset. By deploying an adaptive protocol parsing engine, using deep packet inspection technology and machine learning classification algorithms, it extracts features such as protocol identifiers, message structures, and interaction timing from the data transmission stream to accurately determine the protocol type used by the data. Then, based on the pre-established protocol-parser mapping relationship, the corresponding parser is automatically called to decapsulate and convert the data format, remove protocol-specific information, and uniformly convert data from different protocols into a standardized format containing data source tags, timestamps, and business fields to achieve protocol independence. On this basis, semantic noise filtering and cross-modal cleaning are performed on the protocol-independent standardized data. For numerical data, statistical methods and anomaly detection algorithms are used to eliminate outliers and invalid data; for text data, the semantic understanding ability of large language models is used to identify and filter duplicate content, incorrect expressions, and semantically ambiguous information; in terms of cross-modal data processing, by establishing a unified semantic space, unstructured data such as images and audio are semantically aligned and fused with structured data, ultimately obtaining semantically pure and uniformly formatted data, laying a solid foundation for subsequent data mining, analysis, and application.

[0026] In this embodiment, by analyzing the characteristic distribution of historically collected data and combining it with the status of the current data source, the optimal collection strategy is dynamically generated. At the same time, the potential data demand points are predicted to automatically expand the data collection scope; then, the data is parsed by a parser based on protocol feature mapping to obtain protocol-independent standardized data; this overcomes the defect that existing multi-source data collection technology cannot perform protocol adaptation and adaptive strategy optimization.

[0027] In one embodiment, the protocol-independent standardized data is data that eliminates the communication protocol characteristics that the original data relies on and is uniformly converted into an independent data format.

[0028] In one embodiment, the acquisition strategy includes sampling frequency, batch size, and priority sorting.

[0029] In one embodiment, the characteristic distribution of historically collected data is analyzed and combined with the current state of the data source to dynamically generate an optimal collection strategy, including: Slice the historically collected data and extract the data volume fluctuation characteristics, value density distribution, and field integrity characteristics of each slice to obtain the data feature vector; Monitor the connection success rate, response timeliness, and server load pressure of each data source, and conduct hierarchical evaluation to obtain a data source status matrix; A reinforcement learning model is constructed, and multi-objective reinforcement learning optimization processing is performed on the data feature vector and the data source state matrix to generate an optimal parameter vector, and an optimal acquisition strategy is obtained based on the optimal parameter vector.

[0030] In this embodiment, to comprehensively analyze the underlying patterns inherent in historically collected data, we first employ the sliding window technique used in time series analysis to slice the historically collected data at fixed time intervals (e.g., hours or days), dividing the continuous data stream into multiple independent time segments. For each data slice, we perform deep feature extraction along three dimensions: First, we calculate data volume fluctuation characteristics. By counting the number of data records within the slice and combining statistical indicators such as mean, standard deviation, and coefficient of variation, we quantify the fluctuation and stability of data volume over different time periods and identify abnormal intervals of sudden growth or decline in data volume. Second, we assess the value density distribution. By using a large language model to perform semantic analysis on the data content within the slice, we annotate the proportion of high-value information (e.g., records containing key business indicators and abnormal alerts) within the data through keyword extraction, sentiment analysis, and importance scoring, and construct a curve showing the change in value density over time. Third, we detect field integrity characteristics. By traversing each data record within the slice, we count the number of missing fields in each business field (e.g., "user ID" and "transaction amount"), calculate the field missing rate, and identify unstable data areas with persistent field missing issues. Finally, the above three types of features are normalized and integrated to form a multi-dimensional data feature vector, providing a quantitative data basis for subsequent strategy generation.

[0031] To monitor the operational status of data sources in real time, a multi-dimensional data source monitoring module was deployed. First, the connection success rate of each data source is continuously monitored. This module periodically sends connection requests and calculates the ratio of successful connections to total requests to assess the stability of the data source's network connectivity. Second, the module measures the response time of each data source, recording the time interval between sending and receiving a response for each data request. Different levels of response are defined, including real-time (e.g., response time less than 100ms), near-real-time (100-500ms), and offline (over 500ms). Furthermore, the module collects load pressure indicators from the data source's server, including CPU utilization, memory usage, and disk I / O throughput, to determine whether server resources are overloaded. Based on this monitoring data, each data source is graded (e.g., high, medium, and low) based on pre-set thresholds (e.g., a connection success rate below 90% indicates instability, and a CPU utilization exceeding 80% indicates high load). Finally, the connection success rate level, response timeliness level and server load pressure level of each data source are structured and stored in a matrix form to construct a data source status matrix, which intuitively reflects the real-time health status and service capabilities of all data sources.

[0032] To dynamically optimize the acquisition strategy, a decision-making model based on deep reinforcement learning was constructed. This model takes data feature vectors and a data source state matrix as input, and optimizes the multi-objective objectives of maximizing acquisition efficiency, minimizing resource consumption, and optimizing data value. During model training, the data feature vectors and the data source state matrix are mapped into the reinforcement learning "state" space, while different acquisition strategy parameters (such as sampling frequency, batch size, and acquisition priority) are defined as the "action" space. By designing a well-designed reward function (e.g., positively rewarding the complete acquisition of high-value data and negatively penalizing excessive resource consumption), the agent is guided to explore different action combinations in the state space. Using classic reinforcement learning algorithms such as actor-criticism, the agent continuously interacts with the environment, learning the optimal decision-making strategy from historical data and gradually adjusting the strategy parameters to maximize long-term cumulative rewards. After multiple rounds of iterative training, when the model converges, it outputs an optimal parameter vector containing parameters such as the optimal sampling frequency, batch size, and data source priority. Finally, the optimal parameter vector is converted into executable collection strategy instructions, and collection resources are dynamically allocated to different data sources. While ensuring data integrity and value, efficient use of computing resources is achieved, and the optimal collection strategy that adapts to data changes and data source status is generated.

[0033] In one embodiment, a domain knowledge graph is constructed based on a large language model, the entity associations and business logic of the self-fault-tolerant data stream are analyzed, and potential data demand points are predicted, including: Identifying entities in the self-fault-tolerant data stream, building semantic associations between entities, and generating a set of entity relationship triples; Build a domain knowledge graph based on a large language model, map the entity relationship triple set to the domain knowledge graph, calculate the similarity between entities, and complete the implicit relationship to obtain a dynamic knowledge graph; Analyze the association paths of entities in the dynamic knowledge graph, extract business rules, identify event sequence patterns through time series analysis, and generate a business logic rule base; Through the Transformer-based demand forecasting model, the business logic rule base and dynamic knowledge graph are analyzed to identify data missing points and business association gaps, and output potential data demand points.

[0034] In this embodiment, after acquiring a fault-tolerant data stream, the system first leverages the powerful natural language processing capabilities of a large language model to scan various data types in the data stream, including text and numerical values, using named entity recognition (NER) technology to identify meaningful entities, such as "user ID," "product model," and "order number." Numerical data is converted into corresponding entity concepts using predefined mapping rules or contextual semantics (e.g., mapping "1001" to "specific user ID"). After identifying the entities, the system further utilizes dependency parsing and semantic role labeling to analyze the semantic relationships between entities and clarify their roles in the business scenario, such as "user - purchase - product" and "order - contains - item." Each set of entities and their relationships is structured as a "(subject, predicate, object)" triple. Ultimately, this aggregated set of triples encompasses a large number of entity relationships. This set clearly presents the direct connections between entities in the data stream, providing foundational data for subsequent knowledge graph construction.

[0035] With the industry business domain as the background, with the help of the knowledge system pre-trained by the large language model, an initial domain knowledge graph framework is constructed, which covers the common entity types, relationship categories and basic business rules in the field. Subsequently, the set of entity-relationship triples generated in the previous step is accurately mapped to the corresponding nodes and lines of the domain knowledge graph according to the entity type and relationship category, realizing the transformation from scattered triple data to a structured knowledge network. During the mapping process, knowledge embedding algorithms (such as TransE and ComplEx) are used to map entities and relationships to a low-dimensional vector space, and the degree of semantic similarity between entities is quantified by calculating the cosine similarity or Euclidean distance between vectors. Based on the results of entity similarity calculations and combined with the knowledge reasoning capabilities of the large language model, implicit relationships are mined and completed for entities in the knowledge graph that are not yet clear but have potential associations. For example, through "User A purchases product X" and "Product X belongs to brand Y", we can infer that "User A may be interested in other products of brand Y" and add this implicit relationship to the knowledge graph, thereby continuously enriching and improving the content of the knowledge graph, forming a dynamic knowledge graph that can be dynamically updated and reflect the evolution of business domain knowledge.

[0036] Deep graph analysis is performed on dynamic knowledge graphs, leveraging techniques such as graph neural networks (GNNs) to explore multi-hop association paths between entities and uncover complex indirect relationships and potential influences between entities. For example, by analyzing association paths such as "user-purchase-product," "product-category-category," and "category-promotion-discount strategy," business rules such as "when a certain category is experiencing a promotion, the number of users purchasing products in that category is likely to increase" can be extracted. Furthermore, time series analysis algorithms (such as sliding windows and Fourier transforms) are employed to identify regular event sequence patterns within the data stream (e.g., order creation time and logistics status change time). For example, the typical shopping process of "user browses a product → adds it to the shopping cart → places an order" is identified. The business rules extracted from the knowledge graph are integrated, verified, and supplemented with the event sequence patterns derived from time series analysis, removing duplicate or conflicting rules. This ultimately results in a complete, accurate, and business-information-relevant business logic rule base that comprehensively reflects core knowledge within the business domain, including entity relationships, business processes, and logical constraints.

[0037] Finally, a demand forecasting model based on the Transformer architecture is constructed. This model encodes the structured rules in the business logic rule base and the semantic information in the dynamic knowledge graph into a vector representation that the model can process. Leveraging the Transformer model's powerful attention mechanism, global feature extraction and deep semantic understanding are performed on the input data. This model analyzes the alignment between current data assets and business needs across multiple dimensions, including the execution conditions of business logic rules, their impact scope, and the closeness of connections between entities in the knowledge graph. By comparing the knowledge structure of an ideal business scenario with the knowledge system constructed from existing data, this model identifies missing fields in the data (e.g., the lack of a "consumption habits" field in user profiles), uncovered entity types (e.g., "virtual goods" entities in emerging business scenarios), and unestablished business relationships (e.g., the potential connection between "user health data" and "product recommendation strategies"). These identified data gaps and business connection gaps are summarized, filtered, and prioritized, ultimately generating a list of potential data needs. This list clearly identifies the data areas and content that require further collection or supplementation to meet business development and data analysis needs, providing precise guidance for expanding the scope of data collection.

[0038] In one embodiment, performing real-time anomaly detection and adaptive compensation on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream includes: Slicing the dynamically scheduled data stream, extracting numerical features, temporal features, and semantic features, and constructing a feature vector set; Performing anomaly detection based on the feature vector set to generate an abnormal event including an abnormality type and a confidence level; Map abnormal events to a pre-built fault knowledge graph, analyze the abnormal propagation path through a graph neural network, locate the root cause of the problem, and generate a diagnostic report; An optimal compensation strategy is automatically matched based on the diagnosis report, and parallel compensation is performed on the dynamically scheduled data stream according to the optimal compensation strategy to obtain a self-fault-tolerant data stream.

[0039] In this embodiment, to achieve comprehensive monitoring and analysis of dynamically scheduled data streams, a sliding time window technique is used to slice the data stream at preset time intervals (e.g., 1 minute, 5 minutes), dividing the continuous data stream into multiple discrete data segments. For each data slice, deep feature extraction is performed from three dimensions: For numerical feature extraction, statistical indicators such as the mean, standard deviation, maximum, and minimum values of the numerical data within the data segment are calculated to quantify the central tendency and degree of dispersion of the data; for time series features, algorithms such as Fourier transform and wavelet analysis are used to explore the periodicity and trend of data changes over time and identify the frequency and amplitude characteristics of data fluctuations; in the semantic feature extraction stage, the natural language processing capabilities of the large language model are utilized to perform word segmentation, part-of-speech tagging, and named entity recognition on text data. The text is then converted into a high-dimensional semantic vector using a word vector model (e.g., BERT and Word2Vec) to capture the semantic information of the text content. The extracted numerical features, temporal features, and semantic features are normalized and integrated into multidimensional feature vectors. Each feature vector corresponds to a data slice, and a complete feature vector set is eventually constructed to provide rich analytical data for subsequent anomaly detection.

[0040] Using the constructed feature vector set as input, a deep learning-based anomaly detection model, such as an autoencoder or Transformer model, is deployed. This model first establishes a representation of the data's normal pattern by learning the feature vectors of normal data, learning the data's distribution patterns and inherent structure in the feature space. During the real-time detection phase, the new feature vector is input into the model and the reconstruction error or difference between it and the normal pattern is calculated. When this error exceeds a preset threshold (e.g., the mean plus three standard deviations), an anomaly is identified. To accurately identify the anomaly type, the model classifies and analyzes the anomaly feature vectors based on pre-defined anomaly classification rules (e.g., numerical anomalies, time series anomalies, and semantic anomalies). For example, if the mean of a numerical feature deviates significantly from the historical mean, it is identified as a numerical jump anomaly; if the time series feature shows periodic data disappearance, it is identified as a time series interruption anomaly; and if the semantic feature contains contradictory expressions or illegal vocabulary, it is identified as a semantic error anomaly. Furthermore, the model calculates the similarity between the anomaly feature and the template for each anomaly type to generate a confidence score for each detected anomaly event, quantifying the reliability of the anomaly judgment. Finally, the detected anomaly type and the corresponding confidence level are integrated to generate an abnormal event record containing the anomaly type and confidence level, providing clear clues for subsequent problem diagnosis.

[0041] A fault knowledge graph is pre-built, covering various data sources, data processing steps, and common fault types. This graph uses entities (such as data source servers, data transmission protocols, and data processing modules) as nodes and relationships between entities (such as "server hosting data source" and "protocol affecting data transmission") as edges, forming a knowledge network for the fault domain. When an anomaly is detected, key information from the anomaly (such as the data source involved and the anomaly type) is mapped to the corresponding node in the fault knowledge graph. Then, leveraging the powerful graph-structured data processing capabilities of graph neural networks (GNNs), the knowledge graph is traversed and analyzed, exploring the possible propagation paths of the anomaly along the edges between nodes. For example, if a data transmission anomaly is detected, the GNN analyzes the status information of each node along the relationship path from "data transmission protocol → network device → data source server" to assess each node's impact on the anomaly. By calculating node importance scores (e.g., using the PageRank algorithm), the node that played a key role in the anomaly's occurrence is identified, indicating the root cause of the problem (e.g., a server CPU overload causing data processing delays). Finally, the basic information of the abnormal event, the analysis process and the root cause of the problem located are summarized to generate a detailed diagnostic report, which clearly presents the cause of the abnormality, the scope of impact and related evidence, and provides an accurate basis for the formulation of subsequent compensation strategies.

[0042] In this embodiment, a compensation strategy library is built in, which pre-stores a variety of compensation strategies for different anomaly types and problem sources. Each strategy includes specific execution steps, applicable scenarios, and expected effect evaluation indicators. When an anomaly diagnosis report is received, key information in the report (such as anomaly type, problem source) is extracted through natural language processing technology and matched with the strategies in the compensation strategy library. A rule-based matching algorithm and similarity calculation method are used to screen out the compensation strategy that best matches the current anomaly situation. If there are multiple candidate strategies, the optimal compensation strategy is determined through multi-criteria decision analysis (such as hierarchical analysis method) by comprehensively considering factors such as the strategy execution cost, recovery efficiency, and the degree of impact on the business. For example, if the diagnostic report shows that the anomaly is caused by a network interruption on the data source server, the optimal compensation strategy may be to switch to a backup data source server and adjust the data request parameters. After determining the optimal compensation strategy, a parallel compensation mechanism is activated, simultaneously implementing compensation operations at multiple levels, including the network layer, protocol layer, and data layer. At the network layer, communication links are automatically switched or network bandwidth allocation is adjusted; at the protocol layer, data transmission protocol parameters are modified or the protocol type is switched; and at the data layer, data repair rules are generated using a large language model to complete or correct erroneous or missing data. By executing compensation strategies in parallel, the impact of anomalies on the data stream is quickly eliminated, restoring normal transmission and processing. Ultimately, a self-fault-tolerant data stream is achieved, ensuring the continuity of the data collection process and the reliability of data quality.

[0043] In one embodiment, identifying protocol features of the data transmission flow of the data asset, parsing the data based on a parser mapping the protocol features to obtain protocol-independent standardized data, includes: Using a finite state automaton to scan the data transmission stream of the data asset bit by bit, and identifying the protocol type and data structure characteristics according to preset protocol state transition rules; The parser corresponding to the protocol type and data structure characteristics is retrieved from the parser rule library, and the data is decapsulated and field mapped to obtain protocol-independent standardized data including data source labels, timestamps and business fields.

[0044] In this embodiment, a finite state automaton (FSA) is deployed as a core identification tool when processing data transmission flows within data assets. Based on a predefined set of states and state transition rules, the FSA sequentially scans the data transmission flow byte by byte and bit by bit. During the scanning process, the automaton starts from the initial state and matches the received data bytes against the predefined rules, triggering corresponding state transitions. For example, when identifying the HTTP protocol, if the automaton scans specific strings such as "GET" and "POST," it transitions from the "initial state" to the "request line parsing state." If it subsequently receives fields that conform to the HTTP header format (such as "Host:" and "Content-Length:"), it continues to the "header parsing state." If a blank line delimiter is detected, it enters the "body parsing state." Through this state transition mechanism, the FSA can accurately determine the protocol type (such as HTTP, MQTT, Modbus, etc.) used in the data transmission flow and simultaneously identify data structure characteristics, including data frame format, field order, delimiter type, and other information. In addition, for custom protocols or unknown protocols, by extending the state transition rule base and combining heuristic algorithms for pattern matching, the recognition of complex protocol types and structures is achieved, and finally an identification result containing the protocol type identifier and data structure description is generated, providing a basis for subsequent data analysis.

[0045] After obtaining the protocol type and data structure characteristics of the data transmission stream, the system accesses a pre-built parser rule library. This rule library stores specialized parsers for various standard protocols (such as TCP / IP and CoAP) as well as custom protocols. Each parser contains decapsulation logic and field mapping rules tailored to the protocol type and data structure. Based on the identification results from the first step, the corresponding parser is precisely retrieved from the rule library. For example, for a data stream identified as MQTT, the MQTT parser is invoked. The parser first decapsulates the data, stripping away the transport layer (such as TCP / UDP) and application layer protocol headers (such as the fixed and variable headers of MQTT), retaining only the data payload. Next, based on pre-defined field mapping rules, it converts protocol-specific fields (such as register addresses in Modbus and status codes in HTTP) into common business fields (such as "device parameter number" and "request and response status"). During this process, metadata is automatically added to the parsed data, including a data source tag (identifying the device or system from which the data originated), a timestamp (recording the exact time the data was collected or generated), and standardized business fields (such as "temperature" and "user ID"). Through the above-mentioned protocol decapsulation, field mapping and metadata injection operations, the original data with specific protocol characteristics is converted into standardized data with a unified format and independent of the specific transmission protocol, thus achieving normalization of data of different protocols and providing standardized basic data for subsequent data cleaning, analysis and fusion.

[0046] In one embodiment, performing semantic noise filtering and cross-modal cleaning on the protocol-independent standardized data to obtain semantically pure data includes: Extracting features from the protocol-independent standardized data to obtain a neural morphological feature map, including spike timing features, synaptic weight features, and population coding features; performing noise filtering on the neuromorphic feature map to obtain an enhanced feature map; Performing entity parsing, event extraction, and causal reasoning on the enhanced feature graph to construct a cognitive graph; The cognitive graph is subjected to cross-modal semantic alignment and a meta-learning-driven cleaning process to obtain semantically pure data.

[0047] In this embodiment, when processing protocol-independent standardized data, a neuromorphic computing architecture is used for feature extraction. First, a biological neuron model is constructed through a memristor array to convert numerical data into neuron discharge frequency to form a pulse timing feature. This conversion simulates the information encoding method of the biological nervous system, converting continuous numerical changes into discrete pulse sequences, so that the data features are more consistent with the characteristics of neural processing. Next, the neural connection strength between different fields is calculated to form synaptic weight features. The degree of association of each data field in the business logic is analyzed, such as the strong correlation between "order amount" and "product quantity". By quantifying this correlation strength, a connection weight network similar to a biological synapse is constructed. Finally, the activity pattern of a neuron cluster is used to represent semantic information to form a group coding feature. For text-type data, its semantic information is encoded through the coordinated discharge pattern of a neuron group, and different semantic contents correspond to different neuron discharge combinations. The extracted pulse timing features, synaptic weight features, and population coding features are integrated to construct a neuromorphic feature map. This map represents data features in the manner of a biological neural network, retaining the data's timing characteristics, semantic associations, and collaborative relationships, providing a richer and more cognitively consistent feature representation for subsequent noise filtering and semantic analysis.

[0048] A quantum-enhanced noise filtering mechanism is deployed for neuromorphic feature maps. First, data noise is mapped to a superposition state of quantum bits, leveraging the uncertainty principle in quantum mechanics to more accurately represent the probability distribution of noise. Quantum entanglement measurement technology is used to identify incoherent noise in the data features—random interference unrelated to the data's semantics. Quantum entanglement captures subtle correlations between features, distinguishing true signal from noise. Next, quantum phase transition properties are leveraged to adaptively adjust the noise threshold. When a complex data environment and high noise levels are detected, the noise filtering threshold is automatically raised; conversely, the threshold is lowered to retain more useful information. During the filtering process, a quantum annealing algorithm is employed to find the optimal noise filtering solution. By minimizing an energy function, it determines which features are noise and which are valid signals. After quantum-enhanced noise filtering, noise in the neuromorphic feature map is effectively suppressed, and useful signals are enhanced, forming an enhanced feature map. This provides purer and more reliable feature data for subsequent cognitive analysis.

[0049] Based on an enhanced feature graph, knowledge is constructed using the powerful semantic understanding capabilities of a large language model. First, entities in the data are identified using coreference resolution techniques to address the problem of different representations of the same entity in different contexts (e.g., "apple" can refer to both fruit and a technology company). Combining named entity recognition and semantic role labeling, various entities (such as users, products, and orders) and their attributes are accurately extracted from the data. Next, business events are identified from the time series data. For example, by analyzing the sequence of order status changes ("create → payment → shipment → receipt"), complete order processing events can be extracted. A causal relationship network between events is then constructed, and a graph neural network (GNN) is used to calculate causal probabilities between them. For example, the strength of the causal relationship between "promotional activities" and "sales growth" is analyzed, representing the complex interactions between events through a graph structure. Entities, events, and their causal relationships are integrated to form a cognitive graph. This graph not only contains entity information and event relationships in the data, but also incorporates causal reasoning knowledge, enabling a deeper understanding of the inherent logic and patterns of the business domain, providing a rich knowledge foundation for cross-modal semantic alignment and data cleansing.

[0050] After obtaining the cognitive graph, cross-modal semantic alignment is performed. First, unstructured data such as images and audio are converted into vector representations using a feature extraction model. These data are then semantically matched with the structured data in the cognitive graph. For example, the visual features of a product image are aligned with the semantic features of the product description text to establish an image-text semantic association. For time series and spatial data, a spatiotemporal alignment algorithm is used to associate events in the time series with entities in three-dimensional space. For example, time series data collected by sensors can be mapped to corresponding geographic locations. During the cross-modal alignment process, an attention mechanism is used to focus on key semantic information in the data, improving alignment accuracy. After completing cross-modal semantic alignment, a meta-learning system is constructed for adaptive cleansing. Using the Model-Agnostic Meta-Learning (MAML) algorithm, common features across different data sources and types are extracted, enabling the cleansing model to quickly adapt to new data distributions. Based on historical cleansing results, cleansing parameters (such as missing value filling methods and outlier handling strategies) are dynamically adjusted. Comparative learning is used to calculate the semantic similarity of the data before and after cleansing to evaluate the cleansing effectiveness. The data was then connected to a brain-computer interface system for final verification, converting the cleaning results into neural stimulation signals. The subconscious responses of domain experts were collected through EEG, and the neural feedback from multiple experts was integrated to form the final cleaning decision. Through cross-modal semantic alignment and meta-learning-driven cleaning, semantic noise in the data was completely eliminated, and data from different modalities was semantically unified, ultimately resulting in semantically pure data, providing a high-quality, unambiguous data foundation for subsequent data analysis and intelligent decision-making.

[0051] In one embodiment, after obtaining semantically clean data, the following steps are included: Acquire attribute information of each data source, generate a character matrix based on the attribute information; draw a polyline in a preset coordinate system based on the attribute information; Extracting attribute information of the character matrix, and determining a target position in a preset Base64 encoding table based on the attribute information; Overlaying the character matrix onto a preset Base64 encoding table; wherein the center of the character matrix overlaps with the target position, and each matrix element overlaps with a cell in the preset Base64 encoding table; Performing an XOR calculation on each matrix element in the character matrix and the characters in the cells where the elements overlap, and combining the calculation results in order to obtain a first character combination string; wherein, during the XOR calculation, if the characters are of the same type, they are counted as 1, and if the characters are of different types, they are counted as 0; Superimposing the broken line on a preset Base32 encoding table according to a rule, selecting characters in all cells passed by the broken line in the preset Base32 encoding table, and combining them to obtain a second character combination string; The first character combination string and the second character combination string are interleaved and combined to obtain an encryption key to encrypt and store the semantically pure data.

[0052] In this embodiment, attribute information is first obtained for each data source. This attribute information covers key characteristics such as the data source type, update frequency, and data volume. Based on this attribute information, the system quantizes and encodes it. According to specific mapping rules, each attribute value is converted into a corresponding character. These characters are arranged according to preset row and column rules to generate a character matrix. Simultaneously, the numerical data in the data source attribute information is mapped to the coordinate axes of a preset coordinate system. The corresponding coordinate points are sequentially connected in the order of the attribute information association to form a broken line.

[0053] Next, the generated character matrix is further analyzed to extract its attribute information, including the number of rows and columns, element distribution patterns, and character frequency. Based on this matrix attribute information, a specific calculation rule is used to determine a target location within a preset Base64 encoding table. This calculation rule comprehensively considers various matrix attributes, such as using the number of rows and columns as offsets to perform positioning calculations within the Base64 encoding table, thereby accurately finding the target location.

[0054] The generated character matrix is then overlaid onto the pre-set Base64 encoding table. During the overlay process, the center of the character matrix is strictly controlled to completely overlap with the previously determined target position, ensuring that the matrix elements correspond to the cells in the Base64 encoding table.

[0055] Next, an XOR calculation is performed on each element in the character matrix and the character in the corresponding cell of the Base64 encoding table. The specific calculation rule is: if the character matrix element and the corresponding cell character are the same type (for example, both are numbers or both are letters), it is counted as 1; if the character type is different, it is counted as 0. The XOR calculation is performed sequentially according to the order of the matrix elements, and the results of each calculation are combined in sequence to finally obtain the first character combination string. This combination string contains the characteristic difference information between the character matrix and the Base64 encoding table.

[0056] At the same time, the previously generated polyline is superimposed onto the preset Base32 encoding table according to established rules. These rules include the polyline's starting position and extension direction in the encoding table. During the superposition process, the cells in the Base32 encoding table that are crossed by the polyline are detected one by one. All characters in these cells are selected and combined in the order in which the polyline crossed them, thereby generating a second character combination string. This combination string records the characteristic information of the interaction between the polyline and the Base32 encoding table.

[0057] Finally, the first character string and the second character string are interleaved and combined. This interleaving can be done by alternating characters from the two strings, for example, first taking the first character of the first string, then the first character of the second string, and so on. This interleaving and combining operation ultimately yields a complete encryption key. This encryption key incorporates multiple aspects of information, including data source attributes, character matrix characteristics, and Base64 and Base32 encoding table properties. It is used to encrypt and store semantically pure data, ensuring the security and confidentiality of the data during storage.

[0058] In one embodiment, after obtaining semantically clean data, the following steps are included: Obtain the attribute information of each data source, convert the attribute information into transfinite ordinal representation, and use Cantor's transfinite ordinal theory to construct a transfinite ordinal chain; Abstract the data structure of semantically pure data into objects in the category, and abstract the data relationship into morphisms. Generate dual categories through the duality principle of category theory, convert the semantically pure data into a dual representation, and obtain a dual data stream. Perform stacked tensor product processing on the transfinite ordinal chain and the dual data stream to generate a stacked tensor product matrix; A logical paradox structure is embedded in the stacked tensor product matrix, and contradiction points are generated through self-referential operations. The position coordinates of the contradiction points in the matrix are extracted, and the coordinate sequence is used to generate an encryption key for encrypting semantically pure data.

[0059] In this embodiment, first, the attribute information of each data source is collected, including key parameters such as data type, update frequency, and security level. Subsequently, this attribute information is converted into a transfinite ordinal representation. Transfinite ordinals are a mathematical concept used in Cantor set theory to describe infinite ordinals. The system maps the attributes to equivalent transfinite ordinal expressions based on their numerical value, importance, and other characteristics. For example, data source attributes that are updated frequently are mapped to high-order transfinite ordinals, while those that are updated infrequently are mapped to low-order ones. On this basis, the operation rules and order relationships of transfinite ordinals are used to arrange the transfinite ordinals corresponding to all attributes in sequence according to the association logic of the data source, constructing a transfinite ordinal chain that reflects the attribute hierarchy and dependency relationship. This ordinal chain not only retains the original information of the data source attributes, but also forms a data encoding with complex hierarchical relationships through the unique mathematical structure of transfinite ordinals.

[0060] Next, the semantically pure data is abstracted using category theory. Basic units within the data, such as records and fields, are abstracted into "objects" within the category; relationships between data units, such as parent-child relationships and reference relationships, are abstracted into "morphisms" within the category. In this way, the originally structured data is transformed into an abstract mathematical model within the framework of category theory. Then, applying the principle of duality in category theory, the constructed categories are reversed. Specifically, the directions of all morphisms are reversed, and the logic of the relationships between objects is adjusted to generate new categories that are dual to the original data structure. In this process, the logical relationships of the data are reorganized, forming a semantically equivalent but structurally completely different dual representation, namely a dual data flow. This dual transformation disrupts the conventional structure of the original data, increasing its complexity and confidentiality.

[0061] Subsequently, the transfinite ordinal chain and the dual data stream are subjected to stacked tensor product processing. Tensor product is an operation method that can combine different mathematical objects. The system regards the transfinite ordinal chain and the dual data stream as elements in the tensor space respectively. Through multi-level tensor product operations, the structure and information of the two are deeply integrated. In each layer of tensor product operation, the tensor basis is dynamically adjusted according to the characteristics of the data, so that operations at different levels can capture different dimensional information of the data. For example, the basic information of the data source attributes is integrated in the first layer of tensor product, and the complex relationship of the data structure is gradually integrated in the subsequent layers. After multiple layers of tensor product operations, a stacked tensor product matrix containing rich semantic and structural information is generated. This matrix highly integrates the data source attributes and data structure from a mathematical level, forming the core data carrier of encryption processing.

[0062] Finally, a logical paradox structure is embedded in the stacked tensor product matrix. A logical paradox is a proposition that contains self-referential statements and leads to contradictions. Using a pre-set algorithm, a self-referential logical structure similar to the "liar paradox" is constructed in the matrix. For example, an element at a specific position in the matrix has a value that depends on the results of operations on other elements of the matrix, and these operations in turn affect the element, forming a logical loop. By operating on this self-referential structure, the system induces contradictions—that is, positions of matrix elements that cannot meet logical consistency. The row and column coordinates of these contradictions in the matrix are extracted, and the coordinate sequence is converted into a character or number combination according to specific rules, ultimately generating an encryption key. This key combines the data source attributes, the dual transformation of the data structure, and the contradictory characteristics of the logical paradox. It is used to encrypt semantically pure data, ensuring the security of the data during encrypted storage and transmission.

[0063] In one embodiment, after obtaining semantically clean data, the following steps are included: Obtaining attribute information of each data source, mapping the attribute information to the control point coordinates of the Bezier curve to generate the Bezier curve; Taking the line connecting the starting point and the end point of the Bezier curve as the axis, recursively split the area enclosed by the curve into a binary tree structure to form a curve binary tree; Based on the curve binary tree, a digital sequence is generated according to rules; The arrangement order of the digital sequence is determined according to the overall bending direction of the Bezier curve, and the grouping method of the digital sequence is controlled by the number of inflection points of the curve. Modulo operation and character mapping are performed on the grouped digital sequence, and the results are combined in the order of layer-by-layer traversal of the curve binary tree to generate an encryption key for encrypting semantically pure data.

[0064] In this embodiment, first, the attribute information of each data source is obtained. This information covers key parameters such as data source type, data level, and update frequency. The system digitizes this attribute information and maps it into the coordinates of the control points of a Bezier curve. Specifically, each data source attribute corresponds to one or more control points on the curve. The numerical value of the attribute determines the position of the control point in the coordinate system, and the logical relationship between the attributes affects the distribution pattern of the control points. Through precise calculation and mapping, the system generates a Bezier curve that can fully reflect the attribute characteristics of the data source. The shape, curvature, and trend of the curve are all determined by the data source attribute information, which becomes the core foundation of subsequent encryption processing.

[0065] Next, the area enclosed by the generated Bézier curve is recursively segmented, using the line connecting the start and end points of the curve as the splitting axis. This segmentation process employs a binary tree structure. Starting from the root node, the area enclosed by the curve is divided into two left and right sub-regions along the splitting axis. Each sub-region serves as a child node for the next round of segmentation until a pre-defined segmentation termination condition is met (e.g., the area of the sub-region falls below a threshold or the segmentation level reaches an upper limit). During the segmentation process, the position and direction of each segmentation, as well as the characteristic information of the sub-region, are recorded, ultimately forming a complete binary tree of the curve. This curve binary tree not only reflects the spatial division relationship of the Bézier curve region, but also contains the distribution characteristics of the data source attributes in geometric space.

[0066] Subsequently, based on the constructed binary tree of the curve, a numerical sequence is generated according to pre-set rules. The rule design is closely linked to the node attributes and structural characteristics of the binary tree. For example, the number generation is based on information such as the node's level number, the identifiers of the left and right subtrees, and the geometric parameters of the area represented by the node (such as area and perimeter). Each node is traversed, and the corresponding number is calculated based on the node's specific attribute value using a specific algorithm. These numbers are then arranged in a certain order to form the initial numerical sequence. This numerical sequence contains the structural information of the binary tree of the curve and the geometric characteristics of the Bezier curve area.

[0067] Finally, the order of the digital sequence is determined based on the overall curvature of the Bezier curve. If the curve curves clockwise, the digital sequence is arranged from left to right; if it curves counterclockwise, it is arranged from right to left. Furthermore, the number of inflection points of the Bezier curve is counted to control the grouping of the digital sequence. For example, each inflection point corresponds to a group of numbers, and the number of inflection points determines the number of groups. After grouping, a modular operation is performed on each group of digital sequences, and an appropriate modulus is selected to map the numbers to a specific range. The result of the modular operation is then converted into characters using a character map. Finally, the converted characters are combined in sequence according to the layer-by-layer traversal of the curve binary tree to generate the final encryption key. This key combines the attribute information of the data source, the geometric characteristics of the Bezier curve, and the structural characteristics of the curve binary tree. It is used to encrypt semantically pure data, ensuring the security and confidentiality of the data during storage.

[0068] Reference Figure 2 Another embodiment of the present invention further provides a multi-source data automatic acquisition device based on a large language model, comprising: The analysis unit is used to analyze the characteristic distribution of historically collected data and dynamically generate the optimal collection strategy based on the current state of the data source; A compensation unit is configured to schedule acquisition tasks in real time based on the optimal acquisition strategy to obtain a dynamically scheduled data stream; perform real-time anomaly detection and adaptive compensation on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream; An expansion unit is used to build a domain knowledge graph based on a large language model, analyze the entity associations and business logic of the self-fault-tolerant data stream, and predict potential data demand points; based on the predicted potential data demand points, automatically expand the data collection scope to obtain knowledge-enhanced data assets; A cleaning unit is used to identify the protocol characteristics of the data transmission flow of the data asset, parse the data based on the parser of the protocol feature mapping to obtain protocol-independent standardized data; and perform semantic noise filtering and cross-modal cleaning on the protocol-independent standardized data to obtain semantically pure data.

[0069] In this embodiment, for the specific implementation of each unit in the above device embodiment, please refer to the above method embodiment, which will not be repeated here.

[0070] Reference Figure 3 In an embodiment of the present invention, a computer device is also provided. The computer device may be a server, and its internal structure may be as follows: Figure 3 As shown. The computer device includes a processor, memory, display screen, input device, network interface and database connected via a system bus. The processor of the computer design is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The database of the computer device is used to store the corresponding data in this embodiment. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, the above method is implemented.

[0071] Those skilled in the art will understand that Figure 3 The structure shown in the figure is merely a block diagram of a portion of the structure related to the solution of the present invention and does not constitute a limitation on the computer device to which the solution of the present invention is applied.

[0072] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which implements the above-described method when executed by a processor. It is understood that the computer-readable storage medium in this embodiment can be a volatile readable storage medium or a non-volatile readable storage medium.

[0073] In summary, the multi-source data automatic collection method based on the large language model provided in the embodiment of the present invention includes: analyzing the characteristic distribution of historical collected data, combining the status of the current data source, and dynamically generating the optimal collection strategy; scheduling the collection tasks in real time based on the optimal collection strategy to obtain a dynamically scheduled data stream; performing real-time anomaly detection and adaptive compensation on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream; constructing a domain knowledge graph based on the large language model, analyzing the entity association and business logic of the self-fault-tolerant data stream, and predicting potential data demand points; automatically expanding the data collection scope based on the predicted potential data demand points to obtain knowledge-enhanced data assets; identifying the protocol characteristics of the data transmission stream of the data asset, and parsing the data based on the parser of the protocol feature mapping to obtain protocol-independent standardized data; performing semantic noise filtering and cross-modal cleaning on the protocol-independent standardized data to obtain semantically pure data. In the present invention, by analyzing the characteristic distribution of historically collected data and combining it with the status of the current data source, the optimal collection strategy is dynamically generated. At the same time, the potential data demand points are predicted to automatically expand the data collection scope; then, the data is parsed by a parser based on protocol feature mapping to obtain protocol-independent standardized data; this overcomes the defect that existing multi-source data collection technology cannot perform protocol adaptation and adaptive strategy optimization.

[0074] Those skilled in the art will appreciate that all or part of the processes in the above-described method embodiments can be implemented by instructing the relevant hardware using a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the above-described method embodiments. Any reference to memory, storage, database, or other media provided herein and used in the embodiments may include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double-speed SDRAM (SSRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct RAMbus dynamic RAM (DRDRAM), and RAMbus dynamic RAM.

[0075] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, apparatus, article, or method comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, apparatus, article, or method. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, apparatus, article, or method comprising the element.

[0076] The above description is only a preferred embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made by using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. A multi-source data automatic collection method based on a large language model, characterized in that: The following steps are involved: Analyze the characteristic distribution of historically collected data and dynamically generate the optimal collection strategy based on the current state of the data source; Scheduling acquisition tasks in real time based on the optimal acquisition strategy to obtain dynamically scheduled data streams; Performing real-time anomaly detection and adaptive compensation on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream; Building a domain knowledge graph based on a large language model, analyzing the entity associations and business logic of the self-fault-tolerant data stream, and predicting potential data demand points; based on the predicted potential data demand points, automatically expanding the data collection scope to obtain knowledge-enhanced data assets; Identify the protocol characteristics of the data transmission flow of the data asset, and parse the data based on the parser of the protocol characteristic mapping to obtain protocol-independent standardized data; Semantic noise filtering and cross-modal cleaning are performed on the protocol-independent standardized data to obtain semantically pure data.

2. The multi-source data automatic collection method based on a large language model according to claim 1 is characterized in that: The protocol-independent standardized data is data that eliminates the communication protocol characteristics that the original data relies on and is uniformly converted into an independent data format.

3. The method for automatic multi-source data collection based on a large language model according to claim 1, characterized in that: The acquisition strategy includes sampling frequency, batch size and priority sorting.

4. The method for automatic multi-source data collection based on a large language model according to claim 1, characterized in that: Analyze the characteristic distribution of historically collected data and dynamically generate the optimal collection strategy based on the current state of the data source, including: Slice the historically collected data and extract the data volume fluctuation characteristics, value density distribution, and field integrity characteristics of each slice to obtain the data feature vector; Monitor the connection success rate, response timeliness, and server load pressure of each data source, and conduct hierarchical evaluation to obtain a data source status matrix; A reinforcement learning model is constructed, and multi-objective reinforcement learning optimization processing is performed on the data feature vector and the data source state matrix to generate an optimal parameter vector, and an optimal acquisition strategy is obtained based on the optimal parameter vector.

5. The method for automatic multi-source data collection based on a large language model according to claim 1, characterized in that: Build a domain knowledge graph based on a large language model, analyze the entity associations and business logic of the self-fault-tolerant data flow, and predict potential data demand points, including: Identifying entities in the self-fault-tolerant data stream, building semantic associations between entities, and generating a set of entity relationship triples; Build a domain knowledge graph based on a large language model, map the entity relationship triple set to the domain knowledge graph, calculate the similarity between entities, and complete the implicit relationship to obtain a dynamic knowledge graph; Analyze the association paths of entities in the dynamic knowledge graph, extract business rules, identify event sequence patterns through time series analysis, and generate a business logic rule base; Through the Transformer-based demand forecasting model, the business logic rule base and dynamic knowledge graph are analyzed to identify data missing points and business association gaps, and output potential data demand points.

6. The method for automatic multi-source data collection based on a large language model according to claim 1, characterized in that: Performing real-time anomaly detection and adaptive compensation on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream, including: Slicing the dynamically scheduled data stream, extracting numerical features, temporal features, and semantic features, and constructing a feature vector set; Performing anomaly detection based on the feature vector set to generate an abnormal event including an abnormality type and a confidence level; Map abnormal events to a pre-built fault knowledge graph, analyze the abnormal propagation path through a graph neural network, locate the root cause of the problem, and generate a diagnostic report; An optimal compensation strategy is automatically matched based on the diagnosis report, and parallel compensation is performed on the dynamically scheduled data stream according to the optimal compensation strategy to obtain a self-fault-tolerant data stream.

7. The method for automatic multi-source data collection based on a large language model according to claim 1, characterized in that: Identify the protocol characteristics of the data transmission flow of the data asset, parse the data based on the parser of the protocol characteristic mapping, and obtain protocol-independent standardized data, including: Using a finite state automaton to scan the data transmission stream of the data asset bit by bit, and identifying the protocol type and data structure characteristics according to preset protocol state transition rules; The parser corresponding to the protocol type and data structure characteristics is retrieved from the parser rule library, and the data is decapsulated and field mapped to obtain protocol-independent standardized data including data source labels, timestamps and business fields.

8. The method for automatic multi-source data collection based on a large language model according to claim 1, characterized in that: The protocol-independent standardized data is subjected to semantic noise filtering and cross-modal cleaning to obtain semantically pure data, including: Extracting features from the protocol-independent standardized data to obtain a neural morphological feature map, including spike timing features, synaptic weight features, and population coding features; performing noise filtering on the neuromorphic feature map to obtain an enhanced feature map; Performing entity parsing, event extraction, and causal reasoning on the enhanced feature graph to construct a cognitive graph; The cognitive graph is subjected to cross-modal semantic alignment and a meta-learning-driven cleaning process to obtain semantically pure data.

9. A multi-source data automatic acquisition device based on a large language model, characterized in that: include: The analysis unit is used to analyze the characteristic distribution of historically collected data and dynamically generate the optimal collection strategy based on the current state of the data source; A compensation unit, configured to schedule acquisition tasks in real time based on the optimal acquisition strategy to obtain a dynamically scheduled data stream; Performing real-time anomaly detection and adaptive compensation on the dynamically scheduled data stream to obtain a self-fault-tolerant data stream; An expansion unit is used to build a domain knowledge graph based on a large language model, analyze the entity associations and business logic of the self-fault-tolerant data stream, and predict potential data demand points; based on the predicted potential data demand points, automatically expand the data collection scope to obtain knowledge-enhanced data assets; A cleaning unit, configured to identify protocol features of the data transmission flow of the data asset, and parse the data based on a parser using protocol feature mapping to obtain protocol-independent standardized data; Semantic noise filtering and cross-modal cleaning are performed on the protocol-independent standardized data to obtain semantically pure data.

10. A computer device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Data integration method, device and equipment and storage medium thereof

    CN117370558A

  • Big data file analysis processing method and system in cloud computing environment

    CN118535577A

  • Data acquisition method and system for multiple data sources

    CN119884670A

  • Server load state evaluation method based on dynamic evaluation algorithm

    CN119902905A

Cited By

  • Distribution network fault scheduling decision generation method fusing knowledge graph

    CN120911584A

  • AI-based enterprise data acquisition method, equipment and medium

    CN121524243A

  • Industrial IoT-oriented equipment adaptive access method and system

    CN121567791A