Data management system and method based on heterogeneous platform
By building dynamic knowledge graphs and real-time monitoring, and dynamically adjusting data transmission paths and resource allocation, the data integration efficiency and critical task delay of multi-source heterogeneous platforms in the big data service platform are solved, and efficient and reliable data management is achieved.
Patent Information
- Application Number
- CN202510465537.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-15
AI Technical Summary
In the big data service platform, the data analysis and integration of multi-source heterogeneous platforms are inefficient, and the static scheduling strategy leads to delays in critical tasks and is difficult to deal with dynamic changes in node loads and data exceptions.
Build a dynamic knowledge graph, classify and prioritize scheduling based on protocol features and task metadata, monitor platform load and resource usage in real time, dynamically adjust transmission paths and resource allocation, and build an abnormal data monitoring model to optimize scheduling strategies.
It improves data processing efficiency and accuracy, ensures priority transmission of high real-time data, avoids data congestion and resource waste, realizes system adaptability and stability, and continuously optimizes data management.
Smart Images

Figure CN120337076A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data management in a big data service platform, and in particular, to a data management system and method based on a heterogeneous platform. Background Art
[0002] In the modern big data service platform architecture, each platform usually integrates multiple sources and types of nodes, including data collection nodes, computing nodes, storage nodes, etc. These nodes perform data interaction through different protocols (such as REST API, RPC, message queues, etc.), and the data formats are diverse (such as JSON, Parquet, image / video streams, etc.). This leads to many challenges in the prior art when dealing with multi-source heterogeneous data: First, due to significant differences in the data formats, transmission protocols, and communication mechanisms of different nodes, it takes a long time and is prone to errors when parsing and integrating cross-node data, resulting in low data integration efficiency. For example, the processing flows of real-time data streams and batch data are very different, and traditional static configuration methods are difficult to dynamically adapt. Second, when high-priority tasks and low-priority tasks share limited computing and network resources, static scheduling strategies are likely to cause critical task delays. For example, real-time analysis tasks may not meet the timeliness requirements due to network congestion or computing node overload. In addition, the prior art usually relies on manual configuration of routing paths and resource allocation strategies, and it is difficult to cope with dynamic changes in node loads and data anomalies. When a computing node is overloaded, data may be lost or delayed because it is not routed to an idle node in time. Summary of the Invention
[0003] The purpose of the present invention is to provide a data management system and method based on a heterogeneous platform to solve the problems raised in the above background art.
[0004] To solve the above technical problems, the present invention provides the following technical solution: A data management method based on a heterogeneous platform, characterized in that: the method includes: Step S100: Collect the protocol characteristics and task metadata of multi-source heterogeneous platforms in the big data service platform, and construct a dynamic knowledge graph according to the interaction relationships between different heterogeneous platforms; Step S200: Analyze the data uploaded by different heterogeneous platforms according to the knowledge graph, classify the data according to the data format and task priority, and mark the real-time level. For different categories of data, formulate a priority scheduling strategy; during the scheduling process, introduce a dynamic weight adjustment mechanism to dynamically adjust the weights of task priority and real-time in the scheduling algorithm, so that high-real-time and high-priority data are processed first; Step S300: Monitor the data traffic, processor usage rate, and memory occupancy rate of each heterogeneous platform in real time to calculate the load index of each platform, dynamically adjust the data transmission path in combination with the data real-time level and the priority scheduling strategy, and adjust the resource allocation according to the load condition; when the load of the platform is too high and the transmitted data is high-real-time data, screen out the platform with the lowest load as the alternative routing target by calculating the load index; Step S400: Based on the adjustment of the transmission path and resource allocation, monitor and identify abnormal data that appears in each heterogeneous platform during the transmission process in real time. Set an abnormal monitoring threshold based on protocol characteristics and task metadata, and determine the data as abnormal when the data index exceeds the threshold; extract various abnormal characteristics from the identified abnormal data, construct an abnormal data monitoring model, and dynamically update the knowledge graph according to the abnormal information fed back by the model, so as to continuously optimize the data classification and scheduling strategy.
[0005] Further, the step S100 includes: Step S101: Set the heterogeneous platform set as P = {P1, P2,..., Pi}, where Pi represents the i-th heterogeneous platform, and each heterogeneous platform includes a data interface, a processing module, and a communication module; for the i-th heterogeneous platform Pi, collect the corresponding protocol feature Fc_i and task metadata Fr_i; the protocol features include data format Ff_i, transmission rate Fv_i, and communication protocol type Ft_i; the task metadata includes sampling frequency Fl_i, task priority Fp_i, and signal strength Fs_i; The data format Ff_i represents the organizational structure of the data of the i-th heterogeneous platform during transmission; the transmission rate Fv_i represents the data transmission speed of the i-th heterogeneous platform under different loads; the communication protocol type Ft_i represents the communication protocol used by the i-th heterogeneous platform and the corresponding communication mechanism; the sampling frequency Fl_i represents the number of times each sensor in the i-th heterogeneous platform collects data per second; the task priority Fp_i represents the priority level of the tasks in the i-th heterogeneous platform, including high priority, medium priority, and low priority; the signal strength Fs_i is for the communication module of the i-th heterogeneous platform and is used to quantitatively represent the signal strength in the communication link; Step S102: Select the Neo4j graph database to construct the knowledge graph G(V, E), where the nodes V represent platform types, protocol features, and task metadata, and the edges E represent the association relationships between the nodes; for the i-th heterogeneous platform Pi, create a corresponding node vi in the knowledge graph G, and store the corresponding protocol features and task metadata as the attributes of the node. The attribute set of the node vi is represented as Ai={Fc_i, Fr_i, Ff_i, Fv_i, Ft_i, Fl_i, Fp_i, Fs_i}; according to the data transmission between platforms, establish an edge eij to connect the nodes vi and vj, where vj represents the node corresponding to the j-th heterogeneous platform Pj, and set the weight Wij of the edge to represent the data interaction frequency between the platforms Pi and Pj. The data interaction frequency is calculated by counting the amount of data transmitted between the two platforms within a certain time window: set the time window T, and within the time window T, count the amount of data Dij transmitted from the platform Pi to the platform Pj, and the amount of data Dji transmitted from the platform Pj to the platform Pi. The weight Wij of the edge eij is represented as: Wij=(Dij+Dji) / T; real-time monitor the sampling frequency Fl_i of the platform Pi, and when the sampling frequency changes, update the new sampling frequency value to the attribute Ai of the node vi.
[0006] In the above technical solution, the scattered platform information in the big data service platform is integrated to construct the knowledge system of the service platform, providing a data basis and association basis for subsequent data classification, scheduling, etc. Further, the step S200 includes: Step S201: From the dynamic knowledge graph G(V, E), identify the same type of data through the protocol features and task metadata represented by the nodes V; for the node vi corresponding to each heterogeneous platform Pi and its attribute set Ai, determine the data category and mark the real-time level according to the data format and task priority; set the category mapping function C, for the data uploaded by the heterogeneous platform Pi, denoted as Di, obtain the data category C(Di) through the mapping function C; by traversing the data uploaded by all platforms, group the data with the same value into the same category; gather the identified same-type data into a unified data buffer B. For the m-th type of data, the data uploaded by the platform Pi is denoted as Dim, and all Dim are gathered into the corresponding m-th sub-buffer Bm in the buffer B; at the same time, mark the real-time level Ri for the data according to the task priority, where the high-real-time data is marked as Ri=1, and the low-real-time data is marked as Ri=0. Step S202: Set a priority scheduling policy for each category of data m according to the real-time level Ri of the data; in the data buffer B, perform priority sorting on the data in each sub-buffer Bm, with high-real-time data being processed and transmitted first, and low-real-time data being processed in order; adopt the priority scheduling algorithm: Q(Dim)=α*Ri+β*Fp_i; where Q(Dim) represents the priority value of the data Dim, and α and β are the set real-time and priority weight coefficients respectively, and α+β=1.
[0007] In the above technical solution, based on the knowledge graph, data is classified according to data characteristics, and resource allocation is combined with real-time and task priorities to balance the data requirements of different tasks, providing a basis for subsequent optimization of the data transmission path; The said step S300 includes: Step S301: Monitor the load conditions of each platform in real time, including data traffic, processor usage rate, and memory occupancy rate; calculate the load index Li of each platform: Li=w1*Fi+w2*Ui+W3*Ni; where Fi represents the data traffic of the i-th heterogeneous platform, Ui represents the processor usage rate of the i-th heterogeneous platform, and Ni represents the memory occupancy rate of the i-th platform; w1, w2, and w3 are the weight coefficients of data traffic, processor usage rate, and memory occupancy rate respectively, and w1+w2+w3=1; Step S302 Dynamically adjust the data transmission path according to the load index Li and the data real-time level. When the load of platform Pi is too high and the transmitted data is high-real-time data, screen out the platform with the lowest load as the alternative routing target by calculating the load index Li: R(Dim)=argmin Pj (Lj+γ*(1-Ri)); where R(Dim) represents the routing target selected for the data Dim transmitted from platform Pi to platform Pj, Lj represents the load index of the j-th heterogeneous platform Pj, and γ is the adjustment coefficient. By adjusting the value of γ, the platform load and data real-time are comprehensively considered to find the next routing target platform.
[0008] In the above technical solution, aiming at the dynamically changing load conditions of each platform, the data flow direction is decided in real time, improving the adaptive ability of the system and the reliability of high-real-time data transmission; Furthermore, the said step S400 includes: Step S401: After determining the scheduling strategy and transmission route, monitor the data transmitted by each heterogeneous platform in real time. Based on the data classification in Step S200, set monitoring metrics Mm = {Mm1, Mm2,..., Mmn} for each category of data based on protocol characteristics and task metadata, where Mmn represents the nth monitoring metric for the mth category of data; in the data buffer B, monitor the data in each sub-buffer in real time. For the data Dim uploaded by platform Pi to Bm, extract its corresponding protocol characteristics and task metadata values, and compare them with the monitoring metrics Mm; if a certain characteristic value of the data Dim exceeds the corresponding monitoring metric range, determine that the data is abnormal data; Step S402: Extract abnormal characteristics from all the data marked as abnormal. For each category of data m, summarize the abnormal characteristics of all abnormal data to form an abnormal characteristic set Em = {Em1, Em2,..., Ems}, where s is the number of abnormal data, and Ems represents the sth abnormal data characteristic of the mth category of abnormal data; perform statistical analysis on the abnormal characteristic set Em to calculate the frequency of occurrence of each characteristic; let the number of times the characteristic k appears in the abnormal characteristic set Em be Nmk, then the occurrence frequency Ymk = Nmk / s; set a screening threshold μ, sort all the characteristics in the abnormal characteristic set Em in descending order of occurrence frequency, and select the characteristics with Ymk > μ as the key abnormal characteristics, denoted as Km = {Km1, Km2,..., Kmt}, where t is the number of key abnormal characteristics, and Kmt represents the tth characteristic in the key characteristic set; Step S403: Use the key abnormal characteristic set Km obtained in Step S402 as the input characteristics of the model; at the same time, use the abnormal data and normal data marked in Step S401 as training samples; for each data sample, extract its corresponding key abnormal characteristic values, and perform data normalization operation. Divide the preprocessed data into a training set and a test set according to the ratio of 8:2; construct an abnormal monitoring model through the random forest algorithm and use the training set for training; through multiple random samplings and feature selections of the training data, construct multiple decision trees. During the training process, adjust the number of decision trees in the model and the maximum depth of each decision tree, and finally determine the hyperparameter combination to complete the training of the model; receive new transmitted data input into the model in real time. When abnormal data is detected, output the corresponding abnormal location and abnormal type; the abnormal location is located to the specific heterogeneous platform Pi; the abnormal type is the data in the protocol characteristics and task metadata; according to the abnormal information output by the model, return to Step S102 to update the attribute set Ai of the corresponding node vi and the corresponding edge weight data in the knowledge graph.
[0009] In the above technical solution, closely associated with all the previous steps, it monitors the data during the data transmission process, from data monitoring, anomaly analysis to knowledge graph correction, ensuring that the system continuously improves itself according to the actual situation during long-term operation and maintaining the accuracy and effectiveness of data management; A data management system based on a heterogeneous platform, the system comprising: a heterogeneous platform data collection module, a data priority scheduling module, a transmission path dynamic adjustment module, and a data anomaly monitoring module; The heterogeneous platform data collection module is used to collect the protocol features and task metadata of multi-source heterogeneous platforms in the big data service platform, and construct a dynamic knowledge graph according to the interaction relationships between the heterogeneous platforms; The data priority scheduling module analyzes the data uploaded by the heterogeneous platforms according to the knowledge graph, classifies the data according to the data format and task priority, and marks the real-time level. For different types of data, it formulates a priority scheduling strategy; during the scheduling process, a dynamic weight adjustment mechanism is introduced to dynamically adjust the weights of task priority and real-time in the scheduling algorithm, so that high-real-time and high-priority data are processed first; The transmission path dynamic adjustment module calculates the load index of each platform by real-time monitoring the data traffic, processor usage rate and memory occupancy rate of the heterogeneous platforms, and dynamically adjusts the data transmission path in combination with the data real-time level and the priority scheduling strategy, and at the same time adjusts the resource allocation according to the load condition; when the load of the platform is too high and the transmitted data is high-real-time data, the platform with the lowest load is selected as the alternative routing target by calculating the load index; Based on the adjustment of the transmission path and resource allocation, the data anomaly monitoring module real-time monitors and identifies the abnormal data that appears in each heterogeneous platform during the transmission process. Based on the protocol features and task metadata, it sets an anomaly monitoring threshold, and determines the data as abnormal when the data index exceeds the threshold; extracts various abnormal features from the identified abnormal data, constructs an abnormal data monitoring model, and dynamically updates the knowledge graph according to the abnormal information fed back by the model, so as to continuously optimize the data classification and scheduling strategy.
[0010] Compared with the prior art, the beneficial effects achieved by the present invention are: The present invention realizes the unified management and dynamic update of the data of multi-source heterogeneous platforms in the big data service platform by constructing a dynamic knowledge graph, and through the data classification and priority scheduling strategy, ensures the priority transmission and processing of high-real-time data, and improves the efficiency and accuracy of data processing; The present invention effectively avoids data congestion and resource waste by real-time monitoring the load conditions of each heterogeneous platform, dynamically adjusting the data transmission path and resource allocation, and ensures the smoothness and stability of data transmission; at the same time, reasonably allocates resources according to the load condition, further improving the overall performance of the system and resource utilization rate; By setting monitoring indicators, the present invention monitors the transmitted data in real time, accurately identifies abnormal data based on protocol characteristics and task metadata, extracts abnormal features to construct an abnormal data monitoring model, realizes real-time monitoring and accurate positioning of abnormal data, and dynamically updates the knowledge graph according to the abnormal information fed back by the model, thereby realizing continuous optimization and upgrading of the platform. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] The drawings are used to provide further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention, but do not constitute a limitation to the present invention. In the drawings: Figure 1 is a flowchart of a method for a data management method based on a heterogeneous platform. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0012] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0013] Please refer to Figure 1 , the present invention provides a technical solution: a data management method based on a heterogeneous platform, characterized in that: the method includes: Step S100: Collect protocol characteristics and task metadata of multi-source heterogeneous platforms in a big data service platform, and construct a dynamic knowledge graph according to the interaction relationships between different heterogeneous platforms; Step S200: Analyze the data uploaded by different heterogeneous platforms according to the knowledge graph, classify the data according to the data format and task priority, and mark the real-time level. For different types of data, formulate a priority scheduling strategy; during the scheduling process, introduce a dynamic weight adjustment mechanism to dynamically adjust the weights of task priority and real-time in the scheduling algorithm, so that high-real-time and high-priority data are processed first; Step S300: Monitor the data traffic, processor usage rate, and memory occupancy rate of different heterogeneous platforms in real time to calculate the load index of each platform, and dynamically adjust the data transmission path in combination with the data real-time level and the priority scheduling strategy. At the same time, adjust the resource allocation according to the load condition; when the load of the platform is too high and the transmitted data is high-real-time data, select the platform with the lowest load as the alternative routing target by calculating the load index; Step S400: Based on the adjustment of the transmission path and resource allocation, monitor and identify abnormal data that appears in different heterogeneous platforms during the transmission process in real time. Set an abnormal monitoring threshold based on protocol characteristics and task metadata. When the data metrics exceed the threshold, it is determined as abnormal data. Extract various abnormal characteristics from the identified abnormal data, construct an abnormal data monitoring model, and dynamically update the knowledge graph according to the abnormal information fed back by the model, so as to continuously optimize the data classification and scheduling strategy.
[0014] Further, the step S100 includes: Step S101: Set the heterogeneous platform set as P = {P1, P2,..., Pi}, where Pi represents the i-th heterogeneous platform. Each heterogeneous platform includes a data interface, a processing module, and a communication module. For the i-th heterogeneous platform Pi, collect the corresponding protocol characteristics Fc_i and task metadata Fr_i. The protocol characteristics include data format Ff_i, transmission rate Fv_i, and communication protocol type Ft_i. The task metadata includes sampling frequency Fl_i, task priority Fp_i, and signal strength Fs_i. The data format Ff_i represents the organizational structure of the data of the i-th heterogeneous platform during the transmission process. The transmission rate Fv_i represents the data transmission speed of the i-th heterogeneous platform under different loads. The communication protocol type Ft_i represents the communication protocol used by the i-th heterogeneous platform and the corresponding communication mechanism. The sampling frequency Fl_i represents the number of times each sensor in the i-th heterogeneous platform collects data per second. The task priority Fp_i represents the priority level of the tasks in the i-th heterogeneous platform, including high priority, medium priority, and low priority. The signal strength Fs_i is for the communication module of the i-th heterogeneous platform and is used to quantitatively represent the strength of the signal in the communication link. Step S102: Select the Neo4j graph database to construct the knowledge graph G(V, E), where the nodes V represent platform types, protocol features, and task metadata, and the edges E represent the association relationships between the nodes; for the i-th heterogeneous platform Pi, create a corresponding node vi in the knowledge graph G, and store the corresponding protocol features and task metadata as the attributes of the node. The attribute set of the node vi is represented as Ai={Fc_i, Fr_i, Ff_i, Fv_i, Ft_i, Fl_i, Fp_i, Fs_i}; according to the data transmission between platforms, establish an edge eij to connect the nodes vi and vj, where vj represents the node corresponding to the j-th heterogeneous platform Pj, and set the weight Wij of the edge to represent the data interaction frequency between the platforms Pi and Pj. The data interaction frequency is calculated by counting the amount of data transmitted between the two platforms within a certain time window: set the time window T, and within the time window T, count the amount of data Dij transmitted from the platform Pi to the platform Pj, and the amount of data Dji transmitted from the platform Pj to the platform Pi. The weight Wij of the edge eij is expressed as: Wij=(Dij+Dji) / T; monitor the sampling frequency Fl_i of the platform Pi in real time, and when the sampling frequency changes, update the new sampling frequency value to the attribute Ai of the node vi.
[0015] In the above technical solution, the scattered platform information in the big data service platform is integrated to construct the knowledge system of the service platform, providing a data basis and association basis for subsequent data classification, scheduling, etc. Further, the step S200 includes: Step S201: From the dynamic knowledge graph G(V, E), identify the same type of data through the protocol features and task metadata represented by the nodes V; for the node vi corresponding to each heterogeneous platform Pi and its attribute set Ai, determine the data category and mark the real-time level according to the data format and task priority; set the category mapping function C, for the data uploaded by the heterogeneous platform Pi, denoted as Di, obtain the data category C(Di) through the mapping function C; by traversing the data uploaded by all platforms, group the data with the same value into the same category; collect the identified same type of data into a unified data buffer B. For the m-th type of data, the data uploaded by the platform Pi is denoted as Dim, and all Dim are collected into the corresponding m-th sub-buffer Bm in the buffer B; at the same time, mark the real-time level Ri for the data according to the task priority, where the high-real-time data is marked as Ri=1, and the low-real-time data is marked as Ri=0. Step S202: Set a priority scheduling policy for each category of data m according to the real-time level Ri of the data; in the data buffer B, perform priority sorting on the data in each sub-buffer Bm, with high-real-time data being processed and transmitted first, and low-real-time data being processed in order; adopt the priority scheduling algorithm: Q(Dim)=α*Ri+β*Fp_i; where Q(Dim) represents the priority value of the data Dim, α and β are the set real-time and priority weight coefficients respectively, and α+β=1.
[0016] In the above technical solution, based on the knowledge graph, data is classified according to data characteristics, and resource allocation is combined with real-time and task priorities to balance the data requirements of different tasks, providing a basis for subsequent optimization of the data transmission path; The said step S300 includes: Step S301: Monitor the load conditions of each platform in real time, including data traffic, processor usage rate, and memory occupancy rate; calculate the load index Li of each platform: Li=w1*Fi+w2*Ui+W3*Ni; where Fi represents the data traffic of the i-th heterogeneous platform, Ui represents the processor usage rate of the i-th heterogeneous platform, Ni represents the memory occupancy rate of the i-th platform; w1, w2, w3 are the weight coefficients of data traffic, processor usage rate, and memory occupancy rate respectively, and w1+w2+w3=1; Step S302 Dynamically adjust the data transmission path according to the load index Li and the data real-time level. When the load of platform Pi is too high and the transmitted data is high-real-time data, screen out the platform with the lowest load as the alternative routing target by calculating the load index Li: R(Dim)=argmin Pj (Lj+γ*(1-Ri)); where R(Dim) represents the routing target selected for the data Dim transmitted from platform Pi to platform Pj, Lj represents the load index of the j-th heterogeneous platform Pj, γ is the adjustment coefficient, and by adjusting the value of γ, the platform load and data real-time are comprehensively considered to find the next routing target platform.
[0017] In the above technical solution, aiming at the dynamically changing load conditions of each platform, the data flow direction is decided in real time, improving the adaptive ability of the system and the reliability of high-real-time data transmission; Furthermore, the said step S400 includes: Step S401: After determining the scheduling policy and transmission route, monitor the data transmitted by each heterogeneous platform in real time. Based on the data classification in Step S200, set monitoring metrics Mm = {Mm1, Mm2,..., Mmn} for each category of data based on protocol features and task metadata, where Mmn represents the nth monitoring metric for the mth category of data; in the data buffer B, monitor the data in each sub-buffer in real time. For the data Dim uploaded by platform Pi to Bm, extract its corresponding protocol features and task metadata values, and compare them with the monitoring metrics Mm; if a certain feature value of the data Dim exceeds the corresponding monitoring metric range, determine that the data is abnormal data; Step S402: Extract abnormal features from all the data marked as abnormal. For each category of data m, summarize the abnormal features of all the abnormal data to form an abnormal feature set Em = {Em1, Em2,..., Ems}, where s is the number of abnormal data, and Ems represents the sth abnormal data feature of the mth category of abnormal data; perform statistical analysis on the abnormal feature set Em to calculate the frequency of occurrence of each feature; let the number of times the feature k appears in the abnormal feature set Em be Nmk, then the frequency of occurrence of the feature k is Ymk = Nmk / s; set a screening threshold μ, sort all the features in the abnormal feature set Em in descending order of frequency of occurrence, and screen out the features with Ymk > μ as the key abnormal features, denoted as Km = {Km1, Km2,..., Kmt}, where t is the number of key abnormal features, and Kmt represents the tth feature in the key feature set; Step S403: Use the key abnormal feature set Km obtained in Step S402 as the input features of the model; at the same time, use the abnormal data and normal data marked in Step S401 as training samples; for each data sample, extract its corresponding key abnormal feature values and perform data normalization operations, and divide the preprocessed data into a training set and a test set in a ratio of 8:2; construct an abnormal monitoring model through the random forest algorithm and use the training set for training; through multiple random samplings and feature selections of the training data, construct multiple decision trees. During the training process, adjust the number of decision trees in the model and the maximum depth of each decision tree, and finally determine the hyperparameter combination to complete the training of the model; receive new transmitted data and input it into the model in real time. When abnormal data is detected, output the corresponding abnormal location and abnormal type; the abnormal location is located to the specific heterogeneous platform Pi; the abnormal type is the data in the protocol features and task metadata; according to the abnormal information output by the model, return to Step S102 to update the attribute set Ai of the corresponding node vi and the corresponding edge weight data in the knowledge graph.
[0018] In the above technical solution, it is closely related to all the previous steps, monitors the data during the data transmission process, and from data monitoring, anomaly analysis to knowledge graph correction, ensures that the system continuously improves itself according to the actual situation during long-term operation, and maintains the accuracy and effectiveness of data management; A data management system based on a heterogeneous platform, the system includes: a heterogeneous platform data acquisition module, a data priority scheduling module, a transmission path dynamic adjustment module, and a data anomaly monitoring module; The heterogeneous platform data acquisition module is used to collect the protocol features and task metadata of multi-source heterogeneous platforms in the big data service platform, and construct a dynamic knowledge graph according to the interaction relationship between different heterogeneous platforms; The data priority scheduling module analyzes the data uploaded by different heterogeneous platforms according to the knowledge graph, classifies and marks the real-time level of the data according to the data format and task priority, and formulates a priority scheduling strategy for different types of data; during the scheduling process, a dynamic weight adjustment mechanism is introduced to dynamically adjust the weights of task priority and real-time in the scheduling algorithm, so that high-real-time and high-priority data are processed first; The transmission path dynamic adjustment module calculates the load index of each platform by real-time monitoring the data traffic, processor usage rate and memory occupancy rate of different heterogeneous platforms, and dynamically adjusts the data transmission path in combination with the data real-time level and priority scheduling strategy, and at the same time adjusts the resource allocation according to the load condition; when the load of the platform is too high and the transmitted data is high-real-time data, the platform with the lowest load is selected as the alternative routing target by calculating the load index; The data anomaly monitoring module, based on the adjustment of the transmission path and resource allocation, monitors and identifies abnormal data that appears in different heterogeneous platforms during the transmission process, sets an anomaly monitoring threshold based on the protocol features and task metadata, and determines the data as abnormal when the data index exceeds the threshold; extracts various anomaly features from the identified abnormal data, constructs an abnormal data monitoring model, and dynamically updates the knowledge graph according to the anomaly information fed back by the model, so as to continuously optimize the data classification and scheduling strategy.
[0019] Embodiment of the present invention: Step S100: Set the heterogeneous platform set of the big data service platform as P={P1, P2, P3}, where P1 represents the data acquisition platform, P2 represents the data calculation platform, and P3 represents the data storage platform; For the data acquisition platform P1: Protocol Features: ① Data Format Ff_1: JSON format is adopted, and the data includes timestamp, data source identifier, and specific data values; ② Transmission Rate Fv_1: At the initial stage of data collection, when the load is low, it is 100 Mbps; when a large amount of data is collected concurrently, the transmission rate is increased to 1 Gbps; ③ Communication Protocol Type Ft_1: REST API protocol, data is transmitted through HTTP requests, supporting methods such as GET and POST; Task Metadata: ① Sampling Frequency Fl_1: For real-time sensor data, it is sampled once every 1 ms; for non-real-time data, it is sampled once every 1 hour; ② Task Priority Fp_1: Real-time sensor data collection is of high priority; historical data supplementation is of low priority; ③ Signal Strength Fs_1: If wireless communication is adopted, in a well-signaled indoor environment, the signal strength is stable at about -70 dBm; Data Computing Platform P2: Protocol Features: ① Data Format Ff_2: Parquet format is adopted, which is suitable for large-scale data storage and efficient computing, with columnar storage and compression functions; ② Transmission Rate Fv_2: During daily computing tasks, the transmission rate is 500 Mbps; during complex batch computing, it can be increased to 5 Gbps; ③ Communication Protocol Type Ft_2: RPC protocol, used for efficient remote procedure calls to achieve communication between different computing nodes; Task Metadata: ① Sampling Frequency Fl_2: According to the requirements of the computing task, the computing status is sampled once every 100 ms; ② Task Priority Fp_2: Real-time data analysis tasks are of high priority; batch data processing tasks are of medium priority; data cleaning and archiving tasks are of low priority; ③ Signal Strength Fs_2: Internal communication uses a wired network, and the signal strength is stable; Data Storage Platform P3: Protocol Features: ① Data Format Ff_3: HDFS block storage format is adopted, and the data is split into fixed-size blocks for distributed storage; ② Transmission Rate Fv_3: When writing data, the transmission rate is 200 Mbps; when reading data, depending on the concurrent reading situation, it can reach up to 10 Gbps; ③ Communication Protocol Type Ft_3: HDFS protocol, used to interact with the Hadoop distributed file system; Task Metadata: ① Sampling Frequency Fl_3: The storage status, including disk usage rate, number of files, etc., is sampled once every 5 minutes; ② Task Priority Fp_3: Data backup tasks are of high priority; data migration tasks are of medium priority; data deletion tasks are of low priority; ③ Signal Strength Fs_3: Internal communication uses a wired network, and the signal strength is stable; Select the Neo4j graph database to construct the knowledge graph G(V, E), where the nodes V represent platform types, protocol features, and task metadata, and the edges E represent the association relationships between the nodes; taking P1 as an example, A1 = {Fc_1, Fr_1, Ff_1, Fv_1, Ft_1, Fl_1, Fp_1, Fs_1}; Calculate the edge weights: Set the time window T = 60s. Within this minute, it is statistically found that P1 transfers 500MB of data D12 = 500MB to P2, and P2 transfers 100MB of data D21 = 100MB to P1. Then the weight W12 of the edge e12 = [(500 + 100) * 8] / (60 * 100) ≈ 0.08Mbps; Calculate the weights of other edges in the same way to construct the inter-platform association relationships; Real-time monitoring sampling frequency: If the real-time sensor data sampling frequency of P1 is temporarily increased from every 1ms to every 5ms due to excessive data volume, immediately update the new data to the attribute Fl1 of the node v1; Step S200: Analyze from the dynamic knowledge graph G(V, E). For the data uploaded by P1: If the data format is JSON real-time sensor data and the task priority is high, it is determined as the real-time critical data class through the category mapping function C and marked as the first class, with the real-time level R1 = 1; If it is non-real-time historical data and the priority is low, it is classified as the ordinary data class and marked as the second class, R2 = 0; Classify the data of P2 and P3 in the same way, and store various types of data in the data buffer B. The real-time critical data D11 of P1 is stored in the first sub-buffer B1 of the buffer B; Set the real-time weight coefficient α = 0.6 and the priority weight coefficient β = 0.4; In the buffer B, for the high-real-time real-time critical data in B1, according to the algorithm Q(D11) = 0.6 * 1 + 0.4 * Fp1; where Fp_1 is the corresponding value 3 for high priority, calculate to obtain a higher priority value and process and transmit it first; While the ordinary data in B2 waits to be processed in order to ensure quick response to critical data; Step S300: Real-time monitor the load of each platform: The data traffic F1 of the data acquisition platform P1 = 800Mbps, the processor usage rate U1 = 70%, and the memory occupancy rate N1 = 60%. Set the weight coefficients of data traffic, processor usage rate, and memory occupancy rate w1 = 0.4, w2 = 0.4, w3 = 0.2. Then the load index L1 = 0.4 * 800 + 0.4 * 70 + 0.2 * 60 = 72; Calculate the load indexes of P2 and P3 in the same way; At this time, P1 has a high load and transmits high-real-time real-time critical data R1 = 1. The adjustment coefficient γ is dynamically set according to the network status: γ = 0.5 when the network is normal, and γ = 2 when congested; In this embodiment, the adjustment coefficient γ = 2; Calculate the load indexes of other platforms, L2 = 50, L3 = 30, through the formula R(D1j) = argmin Pj(Lj + γ * (1 - R1) = argmin Pj (Lj), filter out P3 as the alternative routing target, forward the data originally directly transmitted from P1 through the optimized path via P3, and avoid the high load of P1; at the same time, due to the high load of P1, appropriately reduce the allocation of its non-critical task resources and give priority to ensuring the processing of real-time critical data; Step S400: After determining the scheduling and transmission routes, set monitoring indicators for each type of data based on data classification; For real-time critical data class (Class 1): Set the normal range [0, 1000] of the sensor data value in the data format as the monitoring indicator M11; Monitor in buffer B1. The sensor data value in the data D11 uploaded by P1 to B1 is 1200, which exceeds the range of M11, and it is determined as abnormal data; Summarize all abnormal data features. 100 abnormal data are collected within a period of time. It is statistically found that the sensor data value abnormality appears 60 times; Set the screening threshold μ = 0.4. After sorting by frequency, the frequency of the sensor data abnormality feature Y11 = 60 / 100 = 0.6 > μ, and it is screened as the key abnormality feature K1 = {K11}; The anomaly monitoring model uses K1 as the input feature, combines the marked abnormal and normal real-time critical data as training samples, extracts feature values and normalizes them, and divides the training set and test set according to 8:2; Use the random forest algorithm to build the model. When training, adjust the number of decision trees to 100 and the maximum depth to 8 layers to complete the training; Receive new data in real time, detect anomalies, and output the anomaly platform: P1 and the anomaly type: sensor data value anomaly; According to this information, return to step S102 to update the attribute set A1 of the corresponding node v1 and the corresponding edge weight data in the knowledge graph.
[0020] Finally, it should be noted that the above are only the preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A data management method based on a heterogeneous platform, characterized in that: The method includes: Step S100: Collect the protocol features and task metadata of multi-source heterogeneous platforms in the big data service platform, and construct a dynamic knowledge graph according to the interaction relationships between different heterogeneous platforms; Step S200: Analyze the data uploaded by different heterogeneous platforms according to the knowledge graph, classify the data according to the data format and task priority and mark the real-time level, and formulate a priority scheduling strategy for different types of data; During the scheduling process, introduce a dynamic weight adjustment mechanism to dynamically adjust the weights of task priority and real-time in the scheduling algorithm, so that high-real-time and high-priority data are processed first; Step S300: Real-time monitor the data traffic, processor utilization rate and memory occupancy rate of different heterogeneous platforms to calculate the load index of each platform, and dynamically adjust the data transmission path in combination with the data real-time level and priority scheduling strategy, and at the same time adjust the resource allocation according to the load status; When the load of the platform is too high and the transmitted data is high-real-time data, select the platform with the lowest load as the alternative routing target by calculating the load index; Step S400: Based on the adjustment of the transmission path and resource allocation, real-time monitor and identify abnormal data that appears in different heterogeneous platforms during the transmission process. Based on the protocol features and task metadata, set an abnormal monitoring threshold, and determine the data as abnormal data when the data index exceeds the threshold; Extract various abnormal features from the identified abnormal data, construct an abnormal data monitoring model, and dynamically update the knowledge graph according to the abnormal information fed back by the model, so as to continuously optimize the data classification and scheduling strategy.
2. The data management method based on a heterogeneous platform according to claim 1, wherein: The said step S100 includes: Step S101: Set the heterogeneous platform set as P={P1, P2,..., Pi}, where Pi represents the i-th heterogeneous platform, and each heterogeneous platform includes a data interface, a processing module and a communication module; For the i-th heterogeneous platform Pi, collect the corresponding protocol feature Fc_i and task metadata Fr_i; The protocol features include data format Ff_i, transmission rate Fv_i and communication protocol type Ft_i; The task metadata includes sampling frequency Fl_i, task priority Fp_i and signal strength Fs_i; The data format Ff_i represents the organizational structure of the data of the i-th heterogeneous platform during transmission; The transmission rate Fv_i represents the data transmission speed of the i-th heterogeneous platform under different loads; The communication protocol type Ft_i represents the communication protocol used by the i-th heterogeneous platform and the corresponding communication mechanism; The sampling frequency Fl_i represents the number of times each sensor in the i-th heterogeneous platform collects data per second; The task priority Fp_i represents the priority level of the tasks in the i-th heterogeneous platform, including high priority, medium priority and low priority; The signal strength Fs_i is for the communication module of the i-th heterogeneous platform and is used to quantitatively represent the signal strength in the communication link; Step S102: Select the Neo4j graph database to construct the knowledge graph G(V, E), where the nodes V represent platform types, protocol features, and task metadata, and the edges E represent the association relationships between the nodes; for the i-th heterogeneous platform Pi, create a corresponding node vi in the knowledge graph G, and store the corresponding protocol features and task metadata as the attributes of the node. The attribute set of the node vi is represented as Ai={Fc_i, Fr_i, Ff_i, Fv_i, Ft_i, Fl_i, Fp_i, Fs_i}; according to the data transmission between platforms, establish an edge eij to connect the nodes vi and vj, where vj represents the node corresponding to the j-th heterogeneous platform Pj, and set the weight Wij of the edge to represent the data interaction frequency between the platforms Pi and Pj. The data interaction frequency is calculated by counting the amount of data transmitted between the two platforms within a certain time window: set the time window T, and within the time window T, count the amount of data Dij transmitted from the platform Pi to the platform Pj, and the amount of data Dji transmitted from the platform Pj to the platform Pi. The weight Wij of the edge eij is represented as: Wij=(Dij+Dji) / T; monitor the sampling frequency Fl_i of the platform Pi in real time. When the sampling frequency changes, update the new sampling frequency value to the attribute Ai of the node vi.
3. A data management method based on a heterogeneous platform according to claim 1, characterized in that: The step S200 includes: Step S201: From the dynamic knowledge graph G(V, E), identify the same type of data through the protocol features and task metadata represented by the nodes V; for the node vi corresponding to each heterogeneous platform Pi and its attribute set Ai, determine the data category and mark the real-time level according to the data format and task priority; set the category mapping function C. For the data uploaded by the heterogeneous platform Pi, denoted as Di, obtain the data category C(Di) through the mapping function C; by traversing the data uploaded by all platforms, group the data with the same value into the same category; collect the identified same-category data into a unified data buffer B. For the m-th category of data, the data uploaded by the platform Pi is denoted as Dim, and all Dim are collected into the corresponding m-th sub-buffer Bm in the buffer B; at the same time, mark the real-time level Ri for the data according to the task priority, where the high-real-time data is marked as Ri=1, and the low-real-time data is marked as Ri=0; Step S202: According to the real-time level Ri of the data, set a priority scheduling strategy for each category of data m; in the data buffer B, perform priority sorting on the data in each sub-buffer Bm. The high-real-time data is processed and transmitted first, and the low-real-time data is processed in order; adopt the priority scheduling algorithm: Q(Dim)=α*Ri+β*Fp_i; where Q(Dim) represents the priority value of the data Dim, and α and β are the set real-time and priority weight coefficients respectively, and α+β=1.
4. A data management method based on a heterogeneous platform according to claim 1, characterized in that: The step S300 includes: Step S301: Monitor the load conditions of each platform in real time, including data traffic, processor utilization rate, and memory occupancy rate; calculate the load index Li of each platform: Li = w1*Fi + w2*Ui + W3*Ni; where Fi represents the data traffic of the i-th heterogeneous platform, Ui represents the processor utilization rate of the i-th heterogeneous platform, and Ni represents the memory occupancy rate of the i-th platform; w1, w2, and w3 respectively represent the weight coefficients of data traffic, processor utilization rate, and memory occupancy rate, and w1 + w2 + w3 = 1; Step S302 dynamically adjusts the data transmission path according to the load index Li and the data real-time level. When the load of platform Pi is too high and the transmitted data is high-real-time data, the platform with the lowest load is selected as the alternative routing target by calculating the load index Li: R(Dim)=argmin Pj (Lj + γ * (1 - Ri)); where R(Dim) represents the routing target selected for the data Dim transmitted from platform Pi to platform Pj, Lj represents the load index of the j-th heterogeneous platform Pj, γ is an adjustment coefficient, and by adjusting the value of γ, the platform load and data real-time are comprehensively considered to find the next routing target platform.
5. The data management method based on a heterogeneous platform according to claim 1, wherein: The step S400 includes: Step S401: After determining the scheduling policy and transmission route, monitor the data transmitted by each heterogeneous platform in real time. Based on the data classification in step S200, set monitoring indicators Mm = {Mm1, Mm2,..., Mmn} for each category of data based on protocol characteristics and task metadata, where Mmn represents the n-th monitoring indicator of the m-th category of data; in the data buffer B, monitor the data in each sub-buffer in real time. For the data Dim uploaded by the platform Pi to Bm, extract its corresponding protocol characteristics and task metadata values, and compare them with the monitoring indicators Mm; if a certain characteristic value of the data Dim exceeds the corresponding monitoring indicator range, then determine that the data is abnormal data; Step S402: Extract abnormal characteristics from all the data marked as abnormal. For each category of data m, summarize the abnormal characteristics of all abnormal data to form an abnormal characteristic set Em = {Em1, Em2,..., Ems}, where s is the number of abnormal data, and Ems represents the s-th abnormal data characteristic of the m-th category of abnormal data; perform statistical analysis on the abnormal characteristic set Em and calculate the frequency of occurrence of each characteristic; let the number of times the characteristic k appears in the abnormal characteristic set Em be Nmk, then the frequency of occurrence Ymk of the characteristic k is Ymk = Nmk / s; set a screening threshold μ, sort all the characteristics in the abnormal characteristic set Em in descending order of frequency of occurrence, and screen out the characteristics with Ymk > μ as the key abnormal characteristics, denoted as Km = {Km1, Km2,..., Kmt}, where t is the number of key abnormal characteristics, and Kmt represents the t-th characteristic in the key characteristic set; Step S403: Use the set of key anomaly features Km obtained in step S402 as the input features of the model; meanwhile, use the anomaly data and normal data marked in step S401 as training samples; for each data sample, extract its corresponding key anomaly feature value, and perform data normalization operation. Divide the preprocessed data into a training set and a test set according to the ratio of 8:2; construct an anomaly monitoring model through the random forest algorithm and use the training set for training; by performing multiple random samplings and feature selections on the training data, construct multiple decision trees. During the training process, adjust the number of decision trees in the model and the maximum depth of each decision tree, and finally determine the hyperparameter combination to complete the training of the model; receive new transmitted data in real time and input it into the model. When detecting anomaly data, output the corresponding anomaly location and anomaly type; the anomaly location is located to the specific heterogeneous platform Pi; the anomaly type is the data in the protocol features and task metadata; according to the anomaly information output by the model, return to step S102 to update the attribute set Ai of the corresponding node vi and the corresponding edge weight data in the knowledge graph.
6. A data management system based on a heterogeneous platform, characterized in that: The system includes: a heterogeneous platform data acquisition module, a data priority scheduling module, a transmission path dynamic adjustment module, and a data anomaly monitoring module; The heterogeneous platform data acquisition module is used to collect the protocol features and task metadata of multi-source heterogeneous platforms in the big data service platform, and construct a dynamic knowledge graph according to the interaction relationships between different heterogeneous platforms; The data priority scheduling module analyzes the data uploaded by different heterogeneous platforms according to the knowledge graph, classifies and marks the real-time level of the data according to the data format and task priority, and formulates a priority scheduling strategy for different types of data; during the scheduling process, introduce a dynamic weight adjustment mechanism to dynamically adjust the weights of task priority and real-time in the scheduling algorithm, so that high-real-time and high-priority data are processed first; The transmission path dynamic adjustment module calculates the load index of each platform by real-time monitoring the data traffic, processor usage rate, and memory occupancy rate of different heterogeneous platforms, and dynamically adjusts the data transmission path in combination with the data real-time level and priority scheduling strategy, and at the same time adjusts the resource allocation according to the load status; when the load of the platform is too high and the transmitted data is high-real-time data, screen out the platform with the lowest load as the alternative routing target by calculating the load index; The data anomaly monitoring module, based on the adjustment of the transmission path and resource allocation, monitors and identifies the anomaly data that appears in different heterogeneous platforms during the transmission process in real time. Based on the protocol features and task metadata, set the anomaly monitoring threshold, and determine the data as anomaly data when the data index exceeds the threshold; extract various anomaly features from the identified anomaly data, construct an anomaly data monitoring model, and dynamically update the knowledge graph according to the anomaly information fed back by the model, so as to continuously optimize the data classification and scheduling strategy.
7. A data management system based on a heterogeneous platform according to claim 6, characterized in that: The heterogeneous platform data acquisition module includes a platform information acquisition unit and a knowledge graph construction unit; The platform information collection unit is used to set a heterogeneous platform set and collect the protocol features of each heterogeneous platform, including data format, transmission rate, and communication protocol type; and task metadata, including sampling frequency, task priority, and signal strength; The knowledge graph construction unit uses the Neo4j graph database to construct a knowledge graph, stores the collected data as node attributes, and establishes edges and assigns weights according to the data transmission between platforms.
8. A data management system based on a heterogeneous platform according to claim 6, characterized in that: The data priority scheduling module includes a data classification unit and a priority scheduling unit; the data classification unit determines the data category through the data format and task priority, pools the data of the same category into a unified data buffer, and marks the real-time level; The priority scheduling unit sets a priority scheduling policy according to the real-time level and task priority of the data, sorts and processes the data.
9. The data management system based on a heterogeneous platform according to claim 6, wherein: The transmission path dynamic adjustment module includes a load monitoring unit and a path adjustment decision unit; the load monitoring unit is used to monitor and calculate the load index of each platform, including data traffic, processor usage rate, and memory occupancy rate; the path adjustment decision unit filters out the platform with the lowest load as the alternative routing target according to the load index and data real-time level, and adjusts the data transmission path.
10. A data management system based on a heterogeneous platform according to claim 6, characterized in that: The data anomaly monitoring module includes a data monitoring unit, an anomaly analysis unit, and an anomaly monitoring model construction unit; the data monitoring unit sets monitoring indicators for each category of data, extracts the protocol features and task metadata values of the data, and compares them with the monitoring indicators to determine abnormal data; the anomaly analysis unit extracts anomaly features from the abnormal data, conducts statistical analysis, and filters out key anomaly features; The anomaly monitoring model construction unit uses the key anomaly features as input features to construct a random forest anomaly monitoring model, conducts model training and real-time anomaly detection.
Citation Information
Patent Citations
Data processing task scheduling optimization method and system based on knowledge graph
CN115858141A
Multi-source heterogeneous data processing method and system
CN116842099A
Wireless communication network intelligent optimization architecture and method based on knowledge graph and deep learning
CN117479191A
Task processing method and system for comprehensive resource management of intelligent monitoring system
CN117608840A
Link reliability tracking and optimizing method combined with real-time monitoring
CN119325104A
Cited By
Computer data processing system based on machine learning
CN120316587A
A computer data processing system based on machine learning
CN120316587B
Multi-source data automated acquisition and analysis system integrating business intelligence
CN122570017A