A fusion method for automatic collaborative processing of multi-source heterogeneous data
By acquiring the metadata features of multi-source heterogeneous data, generating a standardized description set and dynamically updating the protocol template, combined with distributed message queue technology, it solves the communication barriers in the collaborative processing of multi-source heterogeneous data and achieves efficient data flow and information fusion.
Patent Information
- Application Number
- CN202510938483.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-08
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-08
AI Technical Summary
Existing technologies are difficult to adapt to the communication needs between different systems in the collaborative processing of multi-source heterogeneous data, resulting in inefficient data flow and the possible emergence of data silos. Especially when faced with dynamically changing heterogeneous data, protocol design is difficult to balance universality and efficiency.
By parsing the format and semantics of multi-source heterogeneous data to obtain metadata features, a standardized metadata description set is generated. The initial unified exchange protocol is generated using the protocol template library, and the protocol template is dynamically updated through machine learning algorithms. Combined with distributed message queue technology, a data flow channel is constructed to optimize protocol performance for efficient transmission.
It effectively eliminates the phenomenon of data silos, improves the collaborative processing capabilities of multi-source heterogeneous data, and provides efficient and reliable information fusion and analysis support.
Smart Images

Figure CN120469967B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of information technology, and in particular to a fusion method for automatic collaborative processing of multi-source heterogeneous data. Background Art
[0002] Automatic collaborative processing of multi-source heterogeneous data is an important research direction in the fields of big data and artificial intelligence. Its core lies in achieving efficient integration and collaboration of diverse data sources, providing key support for intelligent decision-making and business optimization. In complex scenarios such as smart cities and the Industrial Internet of Things, data comes from a wide range of sources and in various formats. Efficient collaborative data processing can significantly improve system performance and decision-making quality. However, existing methods for coordinating heterogeneous data often rely on customized protocols for specific fields or simple data format conversion. When faced with large-scale, dynamically changing heterogeneous data, this approach has difficulty adapting to the communication needs of different systems, resulting in inefficient data flow and even the emergence of data silos.
[0003] In the collaborative processing of multi-source heterogeneous data, data communication is the core link for achieving seamless integration, but the technical challenges it faces cannot be ignored. Due to differences in format, protocol, and semantics, it is difficult to directly establish an effective communication mechanism for different data sources. A unified communication mechanism requires the design of a highly compatible data exchange protocol to bridge the communication barriers between heterogeneous systems. However, the design of the protocol must balance universality and efficiency. If the protocol is too complex, it may increase the system's computing burden and reduce real-time processing capabilities. More importantly, the dynamic changes in heterogeneous data require the protocol to have adaptive capabilities to cope with frequent updates and additions to data sources. The above challenges are progressive: heterogeneous differences lead to communication barriers, which require the protocol to balance universality and efficiency, and dynamic changes further increase the technical difficulty of protocol adaptation. Summary of the Invention
[0004] The purpose of the present invention is to solve the above problems and provide a fusion method for automatic collaborative processing of multi-source heterogeneous data.
[0005] The technical solution of the present invention is achieved as follows:
[0006] The present invention provides a fusion method for automatic collaborative processing of multi-source heterogeneous data, the method comprising:
[0007] By analyzing the format and semantics of multi-source heterogeneous data, the metadata characteristics of each data source are obtained. Metadata characteristics include data type, encoding method and semantic label, and a standardized metadata description set is generated;
[0008] Based on the metadata description set, a pre-established protocol template library is used to match applicable templates to generate an initial unified exchange protocol, wherein the initial unified exchange protocol includes format conversion rules and semantic mapping relationships;
[0009] If the format conversion rules of the initial unified exchange protocol cannot adapt to the metadata characteristics of the newly added data source, the protocol template is dynamically updated through a machine learning algorithm to obtain an exchange protocol that adapts to the new data source;
[0010] For the updated exchange protocol, the computational complexity and communication delay parameters of the exchange protocol are obtained to determine whether the preset real-time processing threshold is met, so as to determine the protocol configuration with high communication performance;
[0011] If the computational complexity of the exchange protocol exceeds the threshold, a streamlined exchange protocol is obtained by compressing redundant rules in the protocol and optimizing the format conversion and semantic mapping logic.
[0012] Based on the streamlined exchange protocol, distributed message queue technology is used to build data flow channels and generate unified communication data streams.
[0013] The advantages or beneficial effects of the above technical solution include at least:
[0014] By parsing multi-source heterogeneous data to obtain metadata features, a standardized metadata description set is generated, and an initial unified exchange protocol is generated using a protocol template library. For newly added data sources, the present invention utilizes machine learning algorithms to dynamically update protocol templates and optimizes protocol performance by compressing redundant rules. Distributed message queue technology is used to construct data flow channels for efficient transmission. The present invention also monitors data flow characteristics and utilizes incremental learning algorithms to update protocol rules, enabling adaptive adjustment of the protocol. This method can effectively eliminate data silos, improve the collaborative processing capabilities of multi-source heterogeneous data, and provide efficient and reliable technical support for information fusion and analysis in complex data environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings illustrate exemplary embodiments of the invention and together with the description serve to explain the principles of the invention. These drawings are included to provide a further understanding of the invention and are incorporated in and constitute a part of this specification.
[0016] Figure 1 This is a flow chart of a fusion method for automatic collaborative processing of multi-source heterogeneous data according to the present invention. DETAILED DESCRIPTION
[0017] Embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although certain embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments described herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of the present invention. It should be understood that the drawings and embodiments of the present invention are for illustrative purposes only and are not intended to limit the scope of protection of the present invention.
[0018] It should be noted that, in the absence of conflict, the embodiments and features of the embodiments of the present invention can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0019] It should be understood that the term "including" and its variations used in this document are open inclusions, that is, "including but not limited to". The term "based on" means "based at least in part on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one other embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description. It should be noted that the concepts of "first", "second", etc. mentioned in the present invention are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0020] A fusion method for automatic collaborative processing of multi-source heterogeneous data specifically includes:
[0021] Step S101 , by parsing the format and semantics of multi-source heterogeneous data, metadata features of each data source are obtained, including data type, encoding method and semantic label, and a standardized metadata description set is generated.
[0022] The original data stream is obtained through the multi-source data interface, and the data source and data format are determined using a streaming parsing method. If the data format conforms to the preset structured standard, the data type and encoding method are extracted to obtain metadata features; if not, a standardized data format is generated through a format conversion tool. Based on the metadata features, a semantic parsing algorithm is used to determine the semantic tags and generate a preliminary tag mapping. Through the tag mapping, a clustering algorithm is used to group the semantic tags to obtain a classified tag set. For the classified tag set, a template matching method is used to generate a description set in a unified format. If the completeness of the description set reaches the preset threshold, the final metadata description set is output; if not, the missing metadata features are obtained through a supplementary parsing process. Based on the final metadata description set, a verification tool is used to determine the accuracy of the description set and a standardized metadata description set is generated.
[0023] For example, when parsing multi-source heterogeneous data, the data format is first matched using regular expressions. For example, field names and data types are extracted from CSV files, and "^[a-zA-Z0-9_]+$" is used to match field names. The data type is then inferred using the "infer_dtype" method of the "pandas" library, such as identifying "2023-05-01" as a date type. For JSON data, a recursive parsing algorithm is used to extract nested structures. For example, when parsing "{"sensor": {"id": 123, "value":45.6}}", the extraction path "sensor.id" is an integer, and "sensor.value" is a floating point type. Semantic labels are generated using the pre-trained BERT model. The input field name "temperature" outputs the label "physical quantity / temperature" with a confidence level of 0.92. The "chardet" library is used for encoding detection. For example, the probability of identifying a file encoded as UTF-8 is 95%. During the standardization phase, metadata is mapped to a unified model, such as using RDF triples to generate "<field 1, data type, string>" and "<field 2, semantic label, geographic coordinates>." This description set is hashed and verified using the SHA-256 algorithm to ensure consistency. For image data, an EXIF parser extracts the resolution (1920x1080) and color space (sRGB). For audio files, FFT analysis is used to determine the sampling rate (44.1kHz) and bit depth (16-bit). Ultimately, all metadata is stored in a Neo4j graph database, and relationships between fields are established, such as the "hasAttribute" edge between the "temperature" field and the "sensor" entity.
[0024] Step S102: Based on the metadata description set, a pre-established protocol template library is used to match applicable templates to generate an initial unified exchange protocol, which includes format conversion rules and semantic mapping relationships.
[0025] Extract the data set from the metadata description, and obtain the attribute characteristics of the data set by parsing the structure and content of the metadata description. Based on the attribute characteristics of the data set, use the template matching algorithm to retrieve the applicable template from the protocol template library and determine the matching template identifier. Use the matching template identifier to obtain the template content in the protocol template library, generate the initial unified exchange protocol, and obtain the protocol framework. Based on the data fields in the protocol framework, use the format conversion rules to map the attribute characteristics of the data set to the protocol framework, and generate the protocol content after format conversion. Extract semantic information from the protocol content after format conversion, and align the protocol content with the semantic standard of the target system through the semantic mapping relationship to obtain a semantically aligned protocol. If there are missing fields in the semantically aligned protocol, use the data exchange mechanism to obtain supplementary data from an external data source to generate a complete unified exchange protocol. Based on the complete unified exchange protocol, use the data verification algorithm to determine the integrity and consistency of the protocol content to obtain the final unified exchange protocol.
[0026] For example, in the metadata description set, the structure of the metadata is first parsed to extract key fields such as "data type", "data format", "data source", etc.
[0027] For example, suppose the metadata includes "Data Type" as "Numeric," "Data Format" as "CSV," and "Data Source" as "Database A." Next, based on these fields, the system matches an applicable template from a pre-established protocol template library. Assume that a template exists in the template library for "numeric" data. This template defines the data format conversion rule as "CSV to JSON" and provides semantic mappings, such as mapping the "ID" field in "Database A" to "User Number." Based on the matched template, the system generates an initial unified exchange agreement, which includes specific format conversion rules and semantic mappings.
[0028] For example, the protocol details how to convert each row of data in a CSV file into JSON format, ensuring that the "ID" field corresponds to the "User ID" after conversion. Furthermore, the system optimizes the conversion algorithm based on other metadata information, such as data size. For example, if the data volume is 1 million, the system will select an efficient batch processing algorithm to ensure that the conversion process is completed within a reasonable time. Ultimately, the resulting unified exchange protocol will be directly applied to the data exchange process, ensuring the efficient and accurate transmission of data between different systems.
[0029] In step S103, if the format conversion rules of the initial unified exchange protocol cannot cover the metadata features of the newly added data source, the protocol template is dynamically updated through a machine learning algorithm to obtain an exchange protocol adapted to the new data source.
[0030] By parsing the metadata of the newly added data source, a metadata feature set is obtained, and the degree of match between the feature set and the unified exchange protocol is determined. If the match falls below a preset threshold, key features are extracted from the metadata feature set to generate a feature vector. A random forest algorithm is used to classify the feature vectors, determine the adjustment direction of the required protocol template, and obtain the template adjustment parameters. Based on the template adjustment parameters, the format conversion rules of the protocol template are updated, and a draft exchange protocol adapted for the newly added data source is generated. The adaptability of the draft exchange protocol is verified by simulating data exchange, and a verification result is obtained. If the verification result does not meet the preset adaptability threshold, the draft protocol is optimized using a gradient boosting algorithm to obtain an optimized exchange protocol. By deploying the optimized exchange protocol, the adaptation of the data source to the unified exchange protocol is completed, and the final exchange protocol is generated.
[0031] For example, if the format conversion rules of the initial unified exchange protocol cannot cover the metadata features of a newly added data source, the data collection module first obtains metadata information for the newly added data source, such as data field name, data type, and data length. For example, suppose the newly added data source contains a field named "transaction amount" with a floating-point data type and 10-digit precision. Next, the metadata is processed using a feature extraction algorithm. For example, the field name is vectorized using the TF-IDF algorithm to generate a feature vector. The feature vector is then analyzed using a clustering algorithm such as K-means. Assuming a K value of 3, the metadata is classified into three categories, identifying differences between the newly added data source and the existing protocol template. Based on the clustering results, a decision tree algorithm is used to dynamically update the protocol template. For example, a new node is added to the decision tree to handle the conversion rules for the "transaction amount" field. Finally, the updated protocol template is applied to the data exchange process to ensure smooth integration of the new data source into the system and complete data exchange. This entire process is automated using a machine learning algorithm, eliminating the need for human intervention and improving the system's adaptability and scalability.
[0032] Step S104: for the updated exchange protocol, the computational complexity and communication delay parameters of the protocol are obtained, and it is determined whether a preset real-time processing threshold is met to determine the efficient communication performance of the protocol.
[0033] Obtain the updated exchange protocol configuration and extract computational complexity and communication delay parameters from the protocol. Use static analysis tools to parse the protocol code, count the number of computational steps and data transmission times, and obtain initial values for computational complexity and communication delay. Based on the extracted computational complexity, analyze the algorithm types and data structures involved in the protocol. If the computational steps include nested loops, use loop unrolling to reduce the computational effort. If the data structure is high-complexity, replace it with a low-complexity structure to obtain the optimized computational complexity. Based on the optimized computational complexity, simulate the protocol's runtime under different loads. Use simulation tools to generate load test scenarios, record the protocol processing time, and determine whether it meets the preset real-time processing threshold. Obtain communication delay parameters and analyze the data transmission path and packet size in the protocol. If the transmission path includes multiple routes, use a path optimization algorithm to adjust to the shortest path. If the packet size exceeds the bandwidth limit, use data compression to obtain the optimized communication delay. Based on the optimized communication delay, deploy the protocol in a test network environment. Use a network simulation platform to simulate actual communication scenarios, collect delay data, and determine whether it meets the preset real-time processing threshold. Based on the results of the computational complexity and communication latency assessments, the performance indicators of both are integrated. A weighted average method is used to calculate an overall performance score to determine whether the protocol achieves efficient communication performance. If the overall performance score does not meet the preset threshold, the protocol configuration is iteratively optimized. The algorithm parameters or transmission strategy are adjusted, and the above analysis and testing steps are repeated to obtain a protocol configuration that meets efficient communication performance.
[0034] For example, when evaluating the performance of an updated exchange protocol, its efficiency is first determined through computational complexity analysis. For example, when using a key exchange algorithm based on RSA-2048, a single key generation takes approximately 120ms, while using the optimized ECC-256 algorithm reduces this time to 15ms. By comparing the algorithm's time complexity of O(n^2) and O(n log n), combined with actual server data (e.g., CPU utilization dropping from 70% to 30%), the protocol's improved performance can be quantified. Regarding communication latency parameters, transmission performance under different network environments must be simulated. For example, in a network with 100Mbps bandwidth and 5ms base latency, the protocol's newly added checksum mechanism increases the round-trip time of packets by 2ms. However, using header compression to reduce a 1KB packet to 600 bytes reduces the overall latency from 25ms to 18ms. At this point, the protocol is compared against a preset real-time processing threshold of 50ms. If the protocol's latency is below 40ms (standard deviation ±3ms) in 90% of test cases, it is considered to meet the real-time requirements. To verify efficient communication performance, a multi-node test environment was constructed, injecting a load of 1,000 requests per second. The protocol throughput was monitored to ensure stability above 950 TPS. Wireshark was used to capture packets and analyze the frame interval distribution, confirming that 95% of the data frame intervals were less than 1ms. Combined with a packet loss rate of 0.05% and a retransmission rate of 0.8%, the protocol was determined to meet efficient communication standards. During this process, the congestion window size was dynamically adjusted (for example, gradually increasing from an initial 4KB to 16KB) and the bandwidth utilization curve under the TCP BBR algorithm was recorded to ensure it remained within the optimized range of 85%-92%.
[0035] Step S105: If the computational complexity of the protocol exceeds a threshold, a simplified exchange protocol is obtained by compressing redundant rules in the protocol and optimizing the format conversion and semantic mapping logic.
[0036] If the computational complexity of the protocol exceeds the preset threshold, the protocol rule set is analyzed to obtain a redundant rule set, thereby obtaining a redundant rule set. Based on the redundant rule set, a compression algorithm is used to deduplicate and merge the rules to generate a compressed rule set. Using the compressed rule set, the format conversion logic is optimized to generate a streamlined format conversion table. Based on the streamlined format conversion table, a semantic mapping algorithm is used to generate optimized semantic mapping logic. If the computational complexity of the optimized semantic mapping logic is lower than the preset threshold, a streamlined exchange protocol is generated by integrating the compression rule set, format conversion table, and semantic mapping logic. Based on the streamlined exchange protocol, a protocol verification tool is used to determine the integrity and consistency of the protocol to obtain a verified protocol. Based on the verified protocol, the final protocol deployment file is generated to determine the deployment format.
[0037] For example, when the computational complexity of a protocol exceeds a preset threshold (e.g., a single interaction takes longer than 50ms or CPU utilization exceeds 70%), the system automatically triggers the redundant rule compression process. First, the protocol rule base is scanned using the Apriori algorithm (minimum support set to 0.3), based on frequent itemset mining, to identify frequently repeated rule groups. For example, if five validation rule groups all contain the same field range check (e.g., the numeric interval [0, 100]), they are merged into a single rule and assigned a weight. Next, a format conversion optimization module converts the original multi-layered nested JSON structure (average depth of 4 layers) into a flattened ProtoBuf format, using variable-length integer encoding to compress the field headers to 40% of their original size. The semantic mapping phase uses a pre-trained BERT model (hidden layer dimension 768) to perform vectorized clustering of protocol fields. If the cosine similarity between two fields exceeds 0.85 (e.g., "userID" and "clientID"), an alias mapping relationship table is established. The resulting streamlined protocol is re-evaluated by a complexity validator, requiring its parsing latency to be reduced to less than 20ms and its memory usage to be no more than 50MB. Otherwise, the above process is repeated until the requirements are met. Throughout this process, conflict detection during rule merging uses the Raft consensus algorithm to ensure consistency across distributed nodes. Each compression operation generates a version hash (SHA-256) for tracking the change chain.
[0038] Step S106: Based on the streamlined exchange protocol, a distributed message queue technology is adopted to build a data flow channel, realize efficient transmission of multi-source heterogeneous data, and generate a unified communication data stream.
[0039] Through a distributed message queue deployment, heterogeneous data from multiple sources is acquired to generate an initial data stream. If the data sources have inconsistent formats, they are consolidated into a unified format using pre-set format conversion rules, resulting in a standardized data stream. The message queue's asynchronous transmission mechanism is used to process the standardized data stream and determine data distribution priorities. If the data priority exceeds a preset threshold, it is immediately pushed to the target channel, resulting in a high-priority data stream. Based on the high-priority data stream, a real-time processing channel is established to obtain transmission status information. If the transmission status indicates congestion, the message queue's concurrent capacity is adjusted to obtain an optimized data stream. Based on the optimized data stream, a data compression algorithm is applied to generate compressed data packets. If the compressed data packet size is below a preset threshold, it is directly transmitted to the target system, confirming the transfer is complete. Transmission logs of the compressed data packets are collected and analyzed for data flow efficiency. If the efficiency falls below the preset standard, the node allocation of the distributed architecture is adjusted to obtain an improved data stream. Based on the improved data stream, the message queue's routing policy is updated to generate the final communication data stream. The final communication data stream is used to verify data integrity and determine the success of the transmission.
[0040] For example, when building a data flow channel, a distributed message queue technology (such as Apache Kafka) is first used to build the data flow channel. The number of partitions is set to 8, and the replication factor of each partition is configured to 3 to ensure high data availability and fault tolerance. Using Kafka's producer API, multi-source heterogeneous data is serialized in JSON format, and the message compression method is set to Snappy to reduce network transmission overhead. During data transmission, a polling strategy is used to evenly distribute messages to each partition to ensure load balancing. To improve transmission efficiency, the producer batch size is set to 16KB and the latency is set to 100ms to optimize network I / O. On the consumer side, a multi-threaded concurrent processing mechanism is used, with each thread processing data from a partition. The consumer group size is set to 4 to ensure parallel processing capabilities. Kafka's offset management mechanism records the consumption progress of each partition to prevent duplicate consumption or data loss. During the data flow process, a data verification mechanism is introduced, using the CRC32 algorithm to verify each message to ensure data integrity. To uniformly process heterogeneous data, we employed an ETL (Extract, Transform, Load) process, using the Apache NiFi tool for data extraction. We set the number of concurrent threads to 10 and the processing rate to 1,000 records per second. During the data transformation phase, we used regular expressions to cleanse the data and remove invalid characters. We then aggregated the data using the MapReduce algorithm to generate a unified communication data stream. Finally, we wrote the processed data to the target storage system, such as HDFS, using the Kafka consumer API. We set the block size to 128MB to ensure efficient and scalable data storage.
[0041] Step S107 , by monitoring the transmission rate and error rate of the communication data stream, it is determined whether the data flow efficiency reaches a preset threshold, and the real-time change characteristics of the dynamic data source are obtained.
[0042] By collecting the transmission rate and error rate of the communication data stream, the data flow efficiency is calculated to obtain an efficiency evaluation value. If the efficiency evaluation value falls below a preset threshold, an alarm is triggered, and an abnormal state is identified. Based on the abnormal state, dynamic data change characteristics are extracted from the communication data stream to obtain a feature dataset. A time series analysis algorithm is used to analyze the real-time changes in the feature dataset and determine the trend of change. Based on this trend, the transmission rate and error rate of the communication data stream in the future time period are predicted to obtain a prediction result. If the prediction result shows that the efficiency falls below the preset threshold, the data flow parameters are adjusted to optimize the transmission rate and obtain an optimized data stream. A sliding window method is used to monitor the transmission rate and error rate of the optimized data stream in real time to obtain real-time monitoring data. Based on the real-time monitoring data, the new data flow efficiency is calculated to determine whether it meets the preset threshold. If the new data flow efficiency does not meet the preset threshold, anomalies are detected in the real-time monitoring data to obtain an anomaly feature set. An anomaly detection algorithm is used to analyze the anomaly feature set to determine the cause of the anomaly. Based on the cause of the anomaly, the monitoring system's collection frequency or data flow strategy is adjusted to obtain an adjusted monitoring plan. Through the adjusted monitoring plan, the transmission rate and error rate of the communication data stream are recollected to determine whether the data flow efficiency reaches the preset threshold.
[0043] For example, by monitoring the transmission rate and error rate of the communication data stream, it can be determined in real time whether the data flow efficiency reaches a preset threshold.
[0044] For example, assuming the preset transmission rate threshold is 100 Mbps and the error rate threshold is 0.01%, the system continuously collects the transmission rate and error rate of the data stream. Using a sliding window algorithm, the system calculates the average transmission rate and error rate over the past 10 seconds every second. If the average transmission rate falls below 100 Mbps or the error rate exceeds 0.01%, the system triggers an alarm and records the abnormal data. Simultaneously, the system uses time series analysis methods, such as the ARIMA model, to forecast the collected data and capture the real-time changing characteristics of dynamic data sources.
[0045] For example, by analyzing transmission rate data from the past hour, the ARIMA model can predict transmission rate trends for the next five minutes. If the forecast indicates that the transmission rate will fall below a threshold, the system will take proactive measures, such as adjusting network bandwidth or optimizing data transmission paths, to ensure that data flow efficiency remains above the preset threshold. Furthermore, the system incorporates machine learning algorithms, such as random forests, to analyze historical data, identify key factors influencing transmission rate and error rate, and generate optimization recommendations.
[0046] For example, the random forest model may find that network congestion is the main reason for the decrease in transmission rate during a specific time period. Based on this finding, the system will automatically adjust the data transmission strategy, such as increasing the priority of data transmission or reducing the transmission volume of non-critical data, thereby improving the overall data flow efficiency.
[0047] Step S108: If the changing characteristics of the dynamic data source exceed the protocol adaptation range, the format compatibility and semantic compatibility rules of the protocol are updated through an incremental learning algorithm to obtain an adaptively adjusted exchange protocol.
[0048] If the changing characteristics of a dynamic data source exceed the preset adaptive range, the changing characteristic data is obtained from the data source, and feature extraction techniques are used to determine the set of changing characteristics. Based on the set of changing characteristics, an incremental learning algorithm is used to update the protocol format compatibility rules, resulting in updated format compatibility rules. If the updated format compatibility rules conflict with the existing semantic compatibility rules, semantic analysis techniques are used to adjust the semantic compatibility rules, resulting in adjusted semantic compatibility rules. A preliminary adaptive exchange protocol is generated based on the updated format compatibility rules and the adjusted semantic compatibility rules. Protocol verification techniques are used to test the integrity of the preliminary adaptive exchange protocol and determine whether it meets the data exchange requirements. If the protocol meets the data exchange requirements, the adaptive exchange protocol is streamlined using protocol optimization techniques to obtain the final adaptive exchange protocol. If the protocol does not meet the data exchange requirements, inconsistent data is obtained from the protocol verification results, and an incremental learning algorithm is used to make secondary adjustments to the format compatibility rules and semantic compatibility rules, resulting in the optimized adaptive exchange protocol.
[0049] Exemplarily, when the changing characteristics of the dynamic data source exceed the protocol adaptation range, the changes in the data source are first monitored and analyzed in real time through an incremental learning algorithm.
[0050] For example, if the data source's update frequency increases from 100 to 150 records per minute, the incremental learning algorithm uses online gradient descent to model the data stream. By calculating the data stream's rate of change (Δ = 50 records / minute), the algorithm determines whether it exceeds the protocol's preset upper limit of 100 records / minute. If so, the algorithm initiates an update to the protocol format compatibility rules.
[0051] Specifically, using an XML-based protocol format, by parsing the tag frequencies in the data stream, it was found that the frequency of tag A increased from 30% to 45%, while the frequency of tag B decreased from 20% to 10%. Based on this, the protocol format compatibility rules dynamically adjusted the tag weights, increasing the weight of tag A from 0.3 to 0.45 and decreasing the weight of tag B from 0.2 to 0.1 to ensure the accuracy of data parsing. Simultaneously, the semantic compatibility rules were updated using a BERT-based semantic embedding model to perform semantic analysis on the text in the data stream, finding that the frequency of the keyword "abnormal" increased from 5% to 15%. By calculating semantic similarity, the semantic weight of "abnormal" was adjusted from 0.05 to 0.15 to ensure the accuracy of semantic parsing. Ultimately, through the incremental learning algorithm and dynamic adjustment of protocol rules, an adaptively adjusted exchange protocol was obtained, ensuring format and semantic compatibility after data source changes.
[0052] Step S109 , regenerate the communication data stream according to the adaptively adjusted exchange protocol, verify whether the data island phenomenon is eliminated, and determine the collaborative processing capability of multi-source heterogeneous data.
[0053] Through an adaptively adjusted exchange protocol, communication data streams are generated to obtain an initial data stream set. If there are interruptions or delays in the initial data stream set, a data stream reconstruction algorithm is used to adjust the data stream path to obtain an optimized data stream set. Based on the optimized data stream set, the transmission status of multi-source heterogeneous data is determined to determine whether data islands have been eliminated. Through the collaborative processing module, the multi-source heterogeneous data in the optimized data stream set is format converted to obtain a standardized data set. If there are format inconsistencies in the standardized data set, a data fusion algorithm is used to reintegrate the data to obtain a unified data set. Based on the unified data set, the collaborative processing capability of multi-source heterogeneous data is verified to obtain collaborative processing performance indicators. Performance analysis tools are used to analyze the collaborative processing performance indicators to determine the system's data processing efficiency.
[0054] For example, in the adaptively adjusted exchange protocol, a dynamic weight allocation algorithm is first used to prioritize multi-source heterogeneous data. For example, the initial weight of data source A is set to 0.3, data source B to 0.5, and data source C to 0.2. By monitoring data throughput in real time (e.g., if A's throughput drops from 200MB / s to 150MB / s), the dynamic weight adjustment mechanism is triggered, adjusting A's weight to 0.2, B to 0.6, and C to 0.2, ensuring that high-throughput data sources are prioritized. Next, a data stream pipeline is constructed based on an improved Kafka message queue, with the number of partitions set to 8, the replication factor set to 3, and CRC32 checksum enabled (checksum value 0xEDB88320). The pipeline's stability is verified through a producer throughput test (peaking at 120,000 messages per second). To address data silos, a Flink-based stream processing engine was deployed with a 5-second window size and a 1-second sliding interval. Join operations were performed to join heterogeneous data from MySQL (UUID primary key ID) and MongoDB (three-level nested document structure). Statistics showed that the cross-source join success rate increased from 78% to 95%. Finally, an XGBoost algorithm (learning rate 0.1, tree depth 6) was used to perform a binary classification evaluation on the co-processed feature matrix (128 dimensions), achieving an AUC of 0.92, demonstrating significantly enhanced heterogeneous data collaboration capabilities. During the process, Prometheus monitoring (sampling interval 15 seconds) captured CPU utilization fluctuations between 45% and 65%, while memory usage remained stable within 8GB, ensuring that system resource consumption was manageable.
[0055] It should be understood by those skilled in the art that the above embodiments are merely for the purpose of illustrating the present invention clearly, and are not intended to limit the scope of the present invention. For those skilled in the art, other changes or modifications may be made based on the above disclosure, and these changes or modifications are still within the scope of the present invention.
Claims
1. A fusion method for automatic collaborative processing of multi-source heterogeneous data, characterized by: The method comprises: By analyzing the format and semantics of multi-source heterogeneous data, the metadata characteristics of each data source are obtained. Metadata characteristics include data type, encoding method and semantic label, and a standardized metadata description set is generated; Based on the metadata description set, a pre-established protocol template library is used to match applicable templates to generate an initial unified exchange protocol, wherein the initial unified exchange protocol includes format conversion rules and semantic mapping relationships; If the format conversion rules of the initial unified exchange protocol cannot adapt to the metadata characteristics of the newly added data source, the protocol template is dynamically updated through a machine learning algorithm to obtain an exchange protocol that adapts to the new data source; For the updated exchange protocol, the computational complexity and communication delay parameters of the exchange protocol are obtained to determine whether the preset real-time processing threshold is met, so as to determine the protocol configuration with high communication performance; If the computational complexity of the exchange protocol exceeds the threshold, a streamlined exchange protocol is obtained by compressing redundant rules in the protocol and optimizing the format conversion and semantic mapping logic. Based on the streamlined exchange protocol, distributed message queue technology is used to build data flow channels and generate unified communication data streams; Among them, through machine learning algorithms, the protocol template is dynamically updated to obtain exchange protocols that adapt to new data sources, including: By parsing the metadata of the newly added data source, the metadata feature set is obtained, and the matching degree between the feature set and the unified exchange protocol is determined; if the matching degree is lower than the preset threshold, the key features are extracted from the metadata feature set to generate a feature vector; the feature vector is classified using the random forest algorithm to determine the adjustment direction of the required protocol template and obtain the template adjustment parameters; according to the template adjustment parameters, the format conversion rules of the protocol template are updated to generate a draft exchange protocol adapted to the newly added data source; by simulating data exchange, the adaptability of the draft exchange protocol is verified and the verification result is obtained; if the verification result does not reach the preset adaptability threshold, the gradient boosting algorithm is used to optimize the draft protocol to obtain the optimized exchange protocol; by deploying the optimized exchange protocol, the adaptation of the data source to the unified exchange protocol is completed and the final exchange protocol is generated.
2. The fusion method for automatic collaborative processing of multi-source heterogeneous data according to claim 1 is characterized by: Also includes: By monitoring the transmission rate and error rate of communication data streams, it is determined whether the data flow efficiency has reached the preset threshold and the real-time change characteristics of dynamic data sources are obtained; If the changing characteristics of the dynamic data source exceed the adaptive range of the protocol, the format compatibility and semantic compatibility rules of the protocol are updated through the incremental learning algorithm to obtain an adaptively adjusted exchange protocol; According to the adaptively adjusted exchange protocol, the communication data stream is regenerated to verify whether the data island phenomenon is eliminated.
3. The fusion method for automatic collaborative processing of multi-source heterogeneous data according to claim 2 is characterized by: in, Methods for generating standardized metadata description sets include: Obtain original data streams through multi-source data interfaces and use streaming parsing methods to determine data sources and formats; If the data format meets the preset structural standards, the data type and encoding method are extracted to obtain metadata features; If the data format does not conform to the preset structured standards, a format conversion tool is used to generate a standardized data format until the data format conforms to the preset structured standards; According to metadata features, semantic parsing algorithms are used to determine semantic tags and generate preliminary tag mappings; Through label mapping, a clustering algorithm is used to group semantic labels to obtain a classified label set; For the classified label set, a template matching method is used to generate a description set in a unified format; If the completeness of the description set reaches the preset threshold, the final metadata description set is output; If not, the missing metadata features are obtained through a supplementary parsing process.
4. The fusion method for automatic collaborative processing of multi-source heterogeneous data according to claim 3 is characterized by: It further includes: using a verification tool to determine the accuracy of the final metadata description set and generating a standardized metadata description set.
5. The fusion method for automatic collaborative processing of multi-source heterogeneous data according to claim 1 is characterized by: The method for generating the initial unified exchange protocol includes: Extract data sets from metadata descriptions and obtain attribute characteristics of data sets by analyzing the structure and content of metadata descriptions; According to the attribute characteristics of the data set, a template matching algorithm is used to retrieve the applicable template from the protocol template library and determine the matching template identifier; According to the matching template identifier, the template content in the protocol template library is obtained, the initial unified exchange protocol is generated, and the protocol framework is obtained; For the data fields in the protocol framework, the format conversion rules are used to map the attribute characteristics of the data set to the protocol framework to generate the protocol content after format conversion; Extract semantic information from the protocol content after format conversion, align the protocol content with the semantic standard of the target system according to the semantic mapping relationship, and obtain a semantically aligned protocol; If there are no missing fields in the semantically aligned protocol, it is determined to be a complete unified exchange protocol; If there are missing fields in the semantically aligned protocol, supplementary data is obtained from external data sources through the data exchange mechanism to generate a complete unified exchange protocol; According to the complete unified exchange protocol, a data verification algorithm is used to judge the integrity and consistency of the protocol content to obtain the final unified exchange protocol.
6. The fusion method for automatic collaborative processing of multi-source heterogeneous data according to claim 1, characterized in that: in, The method of obtaining the computational complexity and communication delay parameters of the exchange protocol and determining whether a preset real-time processing threshold is met to determine the protocol configuration with the highest communication performance includes: Obtaining the configuration of the updated exchange protocol and extracting computational complexity and communication delay parameters from the protocol; Static analysis tools are used to parse the protocol code, count the number of operation steps and data transmission times, and obtain the initial values of computational complexity and communication delay; Analyze the algorithm types and data structures involved in the protocol based on the extracted computational complexity; If the algorithm type's operation steps include loop nesting, the amount of calculation is reduced by loop unrolling technology; If the data structure is of high complexity type, it is replaced with a low complexity structure to obtain the optimized computational complexity; Based on the optimized computational complexity, simulate the protocol's runtime under different loads; Use simulation tools to generate load test scenarios, record protocol processing time, and determine whether the preset real-time processing threshold is met; Obtain communication delay parameters and analyze the data transmission path and message size in the protocol; If the transmission path contains multiple routes, it will be adjusted to the shortest path through the path optimization algorithm; If the message size exceeds the bandwidth limit, data compression technology is used to obtain optimized communication delay; For the optimized communication delay, the protocol is deployed in the test network environment. The actual communication scenario is simulated through the network simulation platform, and the delay data is collected to determine whether the preset real-time processing threshold is met. Based on the judgment results of computational complexity and communication delay, the performance indicators of the two are integrated; A weighted average method is used to calculate the comprehensive performance score to determine whether the protocol achieves efficient communication performance; If the comprehensive performance score does not reach the preset threshold, the protocol configuration is iteratively optimized; Adjust the algorithm parameters or transmission strategy and repeat the above analysis and testing steps to obtain a protocol configuration that meets efficient communication performance.
7. The fusion method for automatic collaborative processing of multi-source heterogeneous data according to claim 1 is characterized by: If the computational complexity of the protocol exceeds the threshold, a streamlined exchange protocol is obtained by compressing redundant rules in the protocol and optimizing the format conversion and semantic mapping logic, including: If the computational complexity of the protocol exceeds a preset threshold, a redundant rule set is obtained by analyzing the protocol rule set to obtain a redundant rule set; Based on the redundant rule set, a compression algorithm is used to remove duplicates and merge the rules to generate a compressed rule set; Through the compressed rule set, the format conversion logic is optimized and a streamlined format conversion table is generated; Based on the simplified format conversion table, the semantic mapping algorithm is used to generate optimized semantic mapping logic; If the computational complexity of the optimized semantic mapping logic is lower than a preset threshold, a streamlined exchange protocol is generated by integrating the compression rule set, the format conversion table, and the semantic mapping logic; By streamlining the exchange protocol and using protocol verification tools to determine the integrity and consistency of the protocol, a verified protocol was obtained. Generate the final protocol deployment file based on the verified protocol and determine the deployment format.
8. The fusion method for automatic collaborative processing of multi-source heterogeneous data according to claim 2, characterized in that: in, If the changing characteristics of the dynamic data source exceed the adaptive range of the protocol, the method of obtaining an adaptively adjusted exchange protocol by updating the format compatibility and semantic compatibility rules of the protocol through an incremental learning algorithm includes the following: If the change characteristics of the dynamic data source exceed the preset adaptive range, the change characteristic data is obtained from the data source and the change characteristic set is determined using feature extraction technology; According to the set of changed features, the incremental learning algorithm is used to update the protocol format compatibility rules to obtain the updated format compatibility rules; If the updated format compatibility rule conflicts with the existing semantic compatibility rule, the semantic compatibility rule is adjusted using semantic analysis technology to obtain an adjusted semantic compatibility rule; Generate a preliminary adaptive exchange protocol based on the updated format compatibility rules and the adjusted semantic compatibility rules; Use protocol verification technology to test the integrity of the preliminary adaptive exchange protocol and determine whether the protocol meets the data exchange requirements; If the protocol meets the data exchange requirements, the adaptive exchange protocol is streamlined through protocol optimization technology to obtain the final adaptive exchange protocol.
9. The fusion method for automatic collaborative processing of multi-source heterogeneous data according to claim 8, characterized in that: If the protocol does not meet the data exchange requirements, inconsistent data is obtained from the protocol verification results, and an incremental learning algorithm is used to perform secondary adjustments on the format compatibility rules and semantic compatibility rules to obtain an optimized adaptive exchange protocol.
Citation Information
Patent Citations
Multi-protocol intelligent interaction heterogeneous device full interconnection method and system
CN119324919A
Electric heating gas state information transmission cooperative power grid energy disturbance suppression method
CN119561090A