A processing method and system for supporting multi-source heterogeneous data
By combining Kafka and Flink to build a real-time pipeline and using Ranger unified permission model, the real-time and security problems of existing systems when processing multi-source heterogeneous data are solved, and efficient and secure data processing is achieved.
Patent Information
- Application Number
- CN202510363080.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-03-26
AI Technical Summary
Existing systems are difficult to achieve real-time processing when processing multi-source heterogeneous data, and inconsistent permission models lead to a high risk of leaking sensitive information.
Combining the open source stream processing platform Kafka and the stream data processing engine Flink, a real-time pipeline is built, and the permission inconsistency is solved through the Ranger unified permission model, and the polling interval is dynamically adjusted to optimize the model mapping delay.
It improves the real-time nature of the system to process multi-source heterogeneous data, ensures the security and permissions of data processing, and reduces the risk of sensitive information leakage.
Smart Images

Figure CN119884229B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to a method and system for processing multi-source heterogeneous data. Background Art
[0002] Heterogeneous data refers to a collection of data that differs in structure, format, type, or storage method. Such data usually comes from different data sources or systems and has diverse organizational forms.
[0003] In the SPD business of the medical system, to generate reports covering core data such as products, procurement, and inventory, it is necessary to integrate data from different systems and different structures through a multi-source heterogeneous data processing system. When existing systems process multi-source heterogeneous data, they usually adopt ETL (Extract-Transform-Load) and data warehouse technologies, combined with a serialization framework and metadata management. However, traditional ETL tools rely on batch processing and are difficult to handle real-time data streams; when integrating data from multiple systems, the permission models are inconsistent, and it is extremely easy to cause leakage of sensitive information. Summary of the Invention
[0004] In the present invention, the open-source stream processing platform Kafka and the stream data processing engine Flink are combined to build a real-time pipeline, and multiple real-time pipelines are formed into a stream data processing system to improve the system's ability to process real-time data.
[0005] The technical solution proposed by the present invention is: a method for processing multi-source heterogeneous data, the method comprising:
[0006] Collect data from different target systems and add an identification tag corresponding to the target system to each piece of data;
[0007] Build multiple real-time pipelines parallel to the ETL path to process the collected data, write the data with parsing failures into the error data area, and generate a warning message; write the collected data after removing the data with parsing failures into the temporary storage area;
[0008] Write the output of the real-time data processing system and / or the output of the ETL path into the data warehouse;
[0009] Use the data in the data warehouse to create a visual dynamic report, analyze and predict the dynamic report data through a pre-trained machine learning model, and write the prediction results into the dynamic report.
[0010] Preferably, the data with successful parsing constitutes a correct data set, and the data with parsing failures constitutes an error data set; according to the data structure type, use a clustering algorithm to divide the data in the correct data set into multiple sub-data sets;
[0011] Before obtaining data from different target systems, it further includes the steps of: defining a global policy to unify the permissions of each target system;
[0012] The defining of the global policy includes:
[0013] Register the target system in the console of the security management framework Ranger, and import the resource hierarchy of the target system, where the resource hierarchy includes tables and column families;
[0014] Select corresponding resources according to the resource hierarchy of each target system; specify the operations corresponding to the resources, including permitted operations and denied operations;
[0015] The unifying of the permissions of each target system includes:
[0016] Map the resources of each target system to a unified identifier of Ranger; convert the native operations of each target system into standard operations of Ranger;
[0017] Map the permission models of each target system to a unified "resource - operation - condition" model;
[0018] Create a global role, associate the global role with the corresponding model to bind the resources of the corresponding target system; the conditions include the IP range and the valid duration of the operation;
[0019] The mapping of the permission models of each target system to a unified "resource - operation - condition" model includes:
[0020] Use the NTP protocol to synchronize the clocks of Ranger and the target system;
[0021] Set an extension plugin for the permission model of each target system, and use the polling of the extension plugin to implement changes to the permission model;
[0022] Dynamically adjust the polling interval of the extension plugin to ensure that the permission model of the target system can be mapped in a timely manner.
[0023] Preferably, the dynamically adjusting the polling interval of the extension plugin to ensure that the permission model of the target system can be mapped in a timely manner includes:
[0024] Record the start time of each model mapping; after the permission model is mapped, return the mapping completion time;
[0025] Dynamically calculate a new polling interval through a sliding window weighted average delay and a proportional - integral control algorithm;
[0026] Monitor the delay, replace the current polling interval with the new polling interval to adjust the polling interval and ensure that the permission model of the target system can be mapped in a timely manner.
[0027] Preferably, the dynamic calculation of the new polling interval by the sliding window weighted average delay and proportional integral control algorithm includes:
[0028] After each permission model mapping is completed, calculate the time difference ; where represents the mapping completion time, represents the start time of the model mapping;
[0029] Calculate the weighted average delay , where, represents the size of the sliding window, represents the weight coefficient; represents the current moment before the th time point of the time delay;
[0030] Calculate the control variable ; where, represents the target delay threshold, represents the integral gain one, represents the proportional gain one; represents the average delay of the th sliding window;
[0031] Calculate the new polling interval ; where, represents the current polling interval; represents the maximum polling interval, maximum polling interval;
[0032] If , replace the current polling interval with , and skip the calculation of the control variable.
[0033] Preferably, the definition of the global policy further includes:
[0034] Create business semantic tags;
[0035] Associate the business semantic tags with the resources of different target systems;
[0036] Identify the global role information, assign different business semantic tags to different global roles, and the global roles access the corresponding resources through the corresponding business semantic tags.
[0037] Preferably, constructing multiple real-time pipelines parallel to the ETL path to process the collected data includes:
[0038] Combine the open-source stream processing platform Kafka and the stream data processing engine Flink to build a real-time pipeline. Multiple real-time pipelines form a stream data processing system to process multiple data streams;
[0039] Obtain different types of data from different subsets of data, and convert the data into a format supported by Kafka; if the target system only supports batch file transfer, set a batch file transfer plugin in Kafka to pull data from the temporary storage area;
[0040] After cleaning the obtained data set, collect data in real time through a sliding time window to form a data stream;
[0041] Write the data stream to a preset target storage area and lock it. After Flink completes the checkpoint detection, unlock the target storage area and output the data;
[0042] Maintain the original ETL path of the target system; build a progressive switching strategy to control the process of switching ETL to a real-time pipeline, and gradually switch the ETL data processing tasks of the target system to the real-time pipeline.
[0043] Preferably, the building of the progressive switching strategy to control the process of switching ETL to a real-time pipeline and gradually switching the ETL data processing tasks of the target system to the real-time pipeline includes:
[0044] Quantify the risk indicators in the switching process; the risk indicators include the data difference rate before and after switching, the real-time pipeline processing delay time, the Flink detection failure rate, and the downstream report generation error rate;
[0045] Write the data to both Kafka and the original ETL path at the same time, and compare the output of the ETL path and the output of the real-time pipeline. If the data difference rate between the two outputs is greater than the preset difference rate threshold, stop the switching;
[0046] Switch the non-core task data to the real-time pipeline, and the core task data is still transmitted through the original ETL path;
[0047] Detect the core task data transmission volume. If the core task data transmission volume is lower than the preset transmission volume threshold within a preset time period, switch the core task data to the real-time pipeline and maintain the original ETL path months, ;
[0048] When the real-time pipeline processing delay time exceeds the preset delay time threshold, cut off the real-time pipeline and automatically select the original ETL path to transmit data.
[0049] Preferably, when constructing the progressive switching strategy to control the process of ETL switching to the real-time pipeline and gradually switch the ETL data processing task of the target system to the real-time pipeline, it further includes:
[0050] Obtain the output of the real-time data processing system and the output of the ETL path;
[0051] Calculate the relative entropy of the output of the real-time data processing system and the output of the ETL path:
[0052] , where represents the probability distribution of key data in the output of the real-time data processing system, represents the probability distribution of the same key data in the ETL output; represents the Kullback-Leibler divergence;
[0053] Dynamically adjust the switching process through a PID controller, including the following sub-steps:
[0054] Set the proportion of data flowing into the real-time data processing system as:
[0055] ;
[0056] where ; represents the proportional gain two, represents the integral gain two, represents the derivative gain;
[0057] Set a sliding time window two, and collect the output data of the real-time data processing system and the ETL path in real time to calculate and ;
[0058] Use and to calculate and obtain the relative entropy; calculate the change amount of the proportion of data flowing into the real-time data processing system after passing through a sliding time window two, where is the length of the sliding time window two;
[0059] Dynamically optimize , , through the reinforcement learning algorithm PPO, including the following sub-steps:
[0060] Set the state space ,
[0061] where the integral term , the differential term ;
[0062] Set the action space ;
[0063] Among them, the adjustment amount of the proportional gain two ;
[0064] The adjustment amount of the integral gain two ,
[0065] The adjustment amount of the derivative gain ;
[0066] Set the reward function ;
[0067] Among them, represents the error weight, represents the control quantity change penalty, represents the overshoot penalty;
[0068] Overshoot , represents the ratio of the data flowing into the real-time data processing system at a time point before the current moment ;
[0069] Use the Actor-Critic framework to construct a reinforcement learning algorithm, and use the elements in the action space to update the elements in the state space to achieve the update of ;
[0070] The use of the elements in the action space to update the elements in the state space is specifically: , , ; Among them, , , represent the updated proportional gain two, integral gain two and derivative gain.
[0071] Preferably, creating a visual dynamic report using the data in the data warehouse, analyzing and predicting the dynamic report data through a pre-trained machine learning model, and writing the prediction result into the dynamic report, including:
[0072] Obtain the data of different target systems from the data warehouse, and use the BI tool Tableau to create a dynamic report;
[0073] Load the machine learning model, input the data of different target systems to generate corresponding prediction results, and write them into the dynamic report; specifically including:
[0074] Obtain data from the data warehouse, and identify the data of different target systems through the identification labels of the target systems;
[0075] Extract the eigenvalue of the data of different target systems and then perform standardization processing to form a target system feature set;
[0076] Use the elements in the target system feature set as input variables and input them into the machine learning model to output the prediction result;
[0077] Write the prediction result into the data warehouse, use the BI tool to read the prediction result from the data warehouse and write it into the dynamic report to update the dynamic report; the machine learning model includes one or more of the time series prediction model LSTM, the binary classification model XGBoost, and the linear regression model.
[0078] A processing system supporting multi-source heterogeneous data includes a server, a processor, a communication module, a data acquisition interface, and a memory that are communicatively connected to the processor. The processor is connected to the server through the communication module, and the system is used to execute the processing method for supporting multi-source heterogeneous data.
[0079] Advantages of the present invention:
[0080] In the present invention, by replacing the traditional ETL with a real-time pipeline, constructing a progressive switching strategy, controlling the process of switching the ETL of the target system to the real-time pipeline, and gradually switching the ETL data processing task of the target system to the real-time pipeline, the real-time performance of the system for processing multi-source heterogeneous data is improved; using Ranger to unify permissions solves the problem of inconsistent permissions of multiple target systems; by dynamically adjusting the polling interval, the latency of model mapping is optimized. Description of the drawings
[0081] Figure 1 It is a flowchart of a processing method for supporting multi-source heterogeneous data according to the present invention. Detailed implementation manners
[0082] The following description is used to disclose the present invention so that those skilled in the art can implement the present invention. The preferred embodiments described below are only examples, and those skilled in the art can think of other obvious deformations. The basic principles defined in the following description can be applied to other implementation manners, deformation schemes, improvement schemes, equivalent schemes, and other technical schemes that do not depart from the spirit and scope of the present invention.
[0083] It can be understood that the term "one" should be understood as "at least one" or "one or more". That is, in one embodiment, the number of an element can be one, and in other embodiments, the number of the element can be multiple. The term "one" cannot be understood as a limitation on the number.
[0084] Embodiment 1:
[0085] Refer to Figure 1, the technical solution provided by the present invention is: a method and system for processing multi-source heterogeneous data, including the following steps:
[0086] Step 1: Collect data from different target systems and add an identification tag corresponding to the target system to each piece of data; before this step, there is also a step: define a global policy to unify the permissions of each target system.
[0087] The purpose of setting this step is that traditional ETL tools cannot process some real-time data (such as sensor data), which easily leads to delays in inventory reports. Moreover, when controlling cross-system permissions, due to the inconsistent permission models of ERP, WMS, and logistics systems, the data access risk is increased.
[0088] Among them, defining the global policy includes:
[0089] Register the target system in the console of the security management framework Ranger, import the resource hierarchy of the target system, and the resource hierarchy includes tables and column families; select corresponding resources according to the resource hierarchy of each target system; specify the operations corresponding to the resources, including permitted operations and denied operations.
[0090] Among them, unifying the permissions of each target system includes:
[0091] Map the resources of each target system to the unified identifier of Ranger; convert the native operations of each target system to the standard operations of Ranger; map the permission models of each target system to a unified "resource - operation - condition" model; create a global role, associate the global role with the corresponding model to bind the resources of the corresponding target system; the conditions include the IP range and the effective duration of the operation.
[0092] Mapping the permission models of each target system to a unified "resource - operation - condition" model includes:
[0093] Use the NTP protocol to synchronize the clocks of Ranger and the target system;
[0094] Reduce the clock synchronization error, which can ensure the accuracy of the calculation.
[0095] Set an extension plugin for the permission model of each target system, and use the polling of the extension plugin to implement changes to the permission model;
[0096] Dynamically adjust the polling interval of the extension plugin to ensure that the permission model of the target system can be mapped in a timely manner.
[0097] For example, during the generation of the SPD inventory turnover report, after the permissions are unified, it is convenient to extract the daily inventory snapshot from the WMS (MySQL) and the outbound records from the logistics system (MongoDB).
[0098] Then, calculate the average daily inventory: (initial inventory + ending inventory) / 2, aggregate the total monthly outbound quantity by drug ID; write the result into the inventory_turnover table in the data warehouse (Snowflake), and configure a line chart in Tableau to display the turnover trend of each drug.
[0099] In this embodiment, the polling interval of the dynamic adjustment extension plugin is adjusted to ensure that the permission model of the target system can be mapped in a timely manner, including the following steps:
[0100] Record the start time of each model mapping; after the permission model mapping is completed, return the mapping completion time; dynamically calculate the new polling interval through the sliding window weighted average delay and proportional integral control algorithm; monitor the delay, and replace the current polling interval with the new polling interval to adjust the polling interval and ensure that the permission model of the target system can be mapped in a timely manner.
[0101] Among them, dynamically calculating the new polling interval through the sliding window weighted average delay and proportional integral control algorithm includes:
[0102] After each permission model mapping is completed, calculate the time difference ; where represents the mapping completion time, represents the start time of the model mapping;
[0103] Calculate the weighted average delay , where, represents the size of the sliding window, represents the weight coefficient; represents the current moment before the th time point of the time delay;
[0104] Calculate the control variable ; where, represents the target delay threshold, represents the integral gain, represents the proportional gain; represents the average delay of the th sliding window;
[0105] Calculate the new polling interval ; where, represents the current polling interval; represents the maximum polling interval, Maximum polling interval;
[0106] If , replace the current polling interval with , and skip the calculation of the control variable.
[0107] For example, when seconds, seconds, seconds, , .
[0108] If the delays for three consecutive switches are 15 seconds, 18 seconds, and 20 seconds, then the weighted average delay seconds, and the control variable seconds; seconds.
[0109] If the adjusted switch delay is up to 8 seconds, then seconds, seconds.
[0110] Step 2: Construct multiple real-time pipelines parallel to the ETL path to process the collected data, write the parsed failure data into the error data area, and generate warning messages; write the collected data into the temporary storage area after removing the parsed failure data, form the correct data set with the successfully parsed data, and form the error data set with the parsed failure data; according to the data structure type, use the clustering algorithm to divide the data in the correct data set into multiple sub-data sets to facilitate understanding the quantity of each structure data.
[0111] Step 2 includes the following sub-steps:
[0112] Combine the open-source stream processing platform Kafka and the stream data processing engine Flink to construct real-time pipelines, and multiple real-time pipelines form a stream data processing system to process multiple data streams;
[0113] Obtain different types of data from different sub-data sets, and convert the data into the format supported by Kafka; if the target system only supports batch file transfer, set the batch file transfer plug-in in Kafka and pull the data from the temporary storage area;
[0114] After cleaning the obtained data set, collect the data in real time through a sliding time window to form a data stream;
[0115] Write the data stream into the preset target storage area and lock it, and unlock the target storage area after Flink completes the state point detection to output the data;
[0116] Keep the original ETL path of the target system;
[0117] Build a progressive switching strategy to control the process of ETL switching to a real-time pipeline, and gradually switch the ETL data processing tasks of the target system to the real-time pipeline. Specifically:
[0118] Quantify the risk indicators during the switching process; the risk indicators include the data difference rate before and after switching, the real-time pipeline processing delay time, the Flink detection failure rate, and the downstream report generation error rate.
[0119] In this embodiment, if the difference rate between the output results of the original system and the real-time pipeline system is greater than 1%, it is judged as a high risk when evaluating data consistency; if the Flink event time (real-time pipeline processing delay time) is greater than 5 minutes, it is considered a high risk when evaluating system timeliness.
[0120] Write the data to Kafka and the original ETL path simultaneously, compare the output of the ETL path with the output of the real-time pipeline. If the data difference rate between the two outputs is greater than the preset difference rate threshold, stop the switching.
[0121] Switch the non-core task data to the real-time pipeline, and the core task data is still transmitted through the original ETL path.
[0122] Detect the transmission volume of the core task data. If the transmission volume of the core task data is lower than the preset transmission volume threshold within the preset time period, switch the core task data to the real-time pipeline and keep the original ETL path months, 。
[0123] When the real-time pipeline processing delay time exceeds the preset delay time threshold, cut off the real-time pipeline and automatically select the original ETL path to transmit data.
[0124] For example, if the real-time stream delay exceeds 10 minutes, automatically cut off the real-time pipeline and switch to ETL until the delay resumes. The above function can be implemented by embedding the fault tolerance library Resilience4j or the service fuse degradation component Hystrix into the Flink job side.
[0125] The above method of building a real-time pipeline through Kafka and Flink can improve the timeliness of the medical SPD service. At the same time, the dual-line transmission and progressive switching methods can reduce the risk of switching.
[0126] Step 3: Write the output of the real-time data processing system and / or the output of the ETL path into the data warehouse;
[0127] Step 4: Create a visual dynamic report using the data in the data warehouse, analyze and predict the dynamic report data through a pre-trained machine learning model, and write the prediction results into the dynamic report. It includes the following sub-steps:
[0128] Obtain data from different target systems in the data warehouse and create a dynamic report using the BI tool Tableau;
[0129] Load the machine learning model, input the data of different target systems to generate corresponding prediction results, and write them into the dynamic report; the machine learning model includes one or more of the time series prediction model LSTM, the binary classification model XGBoost, and the linear regression model; specifically including:
[0130] Obtain data from the data warehouse and identify the data of different target systems through the identification tags of the target systems;
[0131] Extract the eigenvalue of the data of different target systems and then perform standardization processing to form a target system feature set; for example, the average drug purchase volume in the past 7 days and the standard deviation of the inventory can form a time series feature set; the historical on-time delivery rate of suppliers can form a risk classification feature set of the procurement and supply system, etc.
[0132] Take the elements in the target system feature set as input variables and input them into the machine learning model to output the prediction results; for example, through the LSTM model, the inventory demand within a certain period in the future can be output, which can be used to generate an inventory demand report; through the binary classification model, the default probability of the supplier can be output, and default warning data can be added to the SPD report.
[0133] Write the prediction results into the data warehouse, use the BI tool to read the prediction results from the data warehouse and write them into the dynamic report to update the dynamic report.
[0134] Traditional reports rely on static rules and cannot adapt to dynamic business changes (such as fluctuations in drug demand caused by emergencies). In this embodiment, the LSTM / XGBoost model is used to predict inventory demand and supplier risks, reducing the dependence on manual experience and improving the prediction accuracy.
[0135] Embodiment 2:
[0136] The difference between this embodiment and Embodiment 1 is that in Embodiment 1, defining the global policy through Ranger is an explicit authorization method based on resources, and in this embodiment, an implicit authorization based on tags is given.
[0137] Specifically, it includes the following steps:
[0138] Create business semantic tags; associate the business semantic tags with the resources of different target systems;
[0139] Identify global role information, assign different business semantic tags to different global roles, and the global roles access corresponding resources through the corresponding business semantic tags.
[0140] In this embodiment, by setting the business semantic tag as the "query service tag", the tag permission operation is set to the "read" operation, the global role is the "statistician", and the business semantic tag is associated with the system (such as HDFS, Hive, Kafka, etc.), and the system resources can be accessed through the business tag.
[0141] Embodiment Three:
[0142] The difference between this embodiment and Embodiment One is that in Embodiment One, during the switching process between ETL and the real-time pipeline, after comparing the data difference rate between the two outputs with the difference rate threshold, the switching control is performed. This difference rate threshold is fixed and needs to be manually set, and it cannot adapt to the dynamically changing data processing process. Therefore, in this embodiment, a dynamic control switching strategy is given, that is, by constructing a dynamic feedback system, the migration traffic ratio is adaptively adjusted in real time according to the behavior differences between the ETL and the real-time pipeline systems. The specific steps are as follows:
[0143] Obtain the output of the real-time pipeline and the output of the ETL path;
[0144] Calculate the relative entropy of the output of the real-time pipeline and the output of the ETL path:
[0145] , where represents the probability distribution of the key data in the output of the real-time pipeline, represents the probability distribution of the same key data in the ETL output; represents the Kullback-Leibler divergence.
[0146] Dynamically adjust the switching process through a PID controller, including the following sub-steps:
[0147] Set the proportion of the data flowing into the real-time data processing system as:
[0148] ;
[0149] where ; represents the proportional gain two, represents the integral gain two, represents the derivative gain; represents the sensitivity of the PID controller to the current difference, represents the correction ability of the PID controller to the long-term deviation, Indicates the judgment of the PID controller on the change trend of the difference. When the difference suddenly increases, reduce the value of
[0150] Set the sliding time window two, and collect the output data of the real-time pipeline and the ETL path in real time to calculate and ;
[0151] Use and to calculate and obtain the relative entropy; calculate the change amount of the proportion of the data flowing into the real-time pipeline after passing through a sliding time window two , where is the length of the sliding time window two;
[0152] Dynamically optimize , , through the Proximal Policy Optimization (PPO) reinforcement learning algorithm, including the following sub-steps:
[0153] Set the state space ,
[0154] where the integral term , and the differential term ;
[0155] Set the action space ;
[0156] where the adjustment amount of the proportional gain two ;
[0157] The adjustment amount of the integral gain two ,
[0158] The adjustment amount of the differential gain ;
[0159] Set the reward function ;
[0160] where represents the error weight, represents the control variable change penalty, represents the overshoot penalty;
[0161] The overshoot amount ; represents the proportion of the data flowing into the real-time data processing system at a time point before the current time .
[0162] Use the Actor-Critic framework to construct a reinforcement learning algorithm, and update the elements in the state space by using the elements in the action space to achieve the update of ;
[0163] Updating the elements in the state space by using the elements in the action space specifically includes: , , ; where , , represent the updated proportional gain two, integral gain two, and derivative gain.
[0164] For example, an implementation process of this embodiment may be:
[0165] Set the initial flow ratio to Set , , ;
[0166] Start the PID controller, calculate once every , and update once;
[0167] If the fluctuation amplitude of within a period is greater than the preset fluctuation amplitude threshold, or the fluctuation frequency exceeds the preset fluctuation frequency threshold, then start the reinforcement learning model to optimize , , ;
[0168] When and , and it lasts for more than 24 hours, and it is determined that the migration is completed, then terminate the handover control task.
[0169] The solution of this embodiment can adaptively shorten the policy effective delay through PID, reduce the average delay, can flexibly handle sudden traffic or policy changes, optimize the delay of the global policy to take effect, and improve the handover efficiency.
[0170] The present invention also provides a processing system supporting multi-source heterogeneous data, including a server, a processor, a communication module, a data acquisition interface, and a memory communicatively connected to the processor. The processor is connected to the server through the communication module, and the system is used to execute the described processing method for supporting multi-source heterogeneous data.
[0171] Embodiments disclosed in the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. Embodiments disclosed in the present invention include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from a network through a communication section, and / or installed from a removable medium. When the computer program is executed by a central processing unit (CPU), the above functions defined in the methods of the present invention are performed. It should be noted that the computer-readable medium in the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium can, for example but not limited to, be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include but are not limited to: an electrical connection with one or more wire segments, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program codes. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program for use by or combined with an instruction execution system, apparatus, or device. The program codes contained on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless segments, wire segments, optical cables, RF, etc., or any suitable combination of the above.
[0172] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a part of code, which contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system that performs the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0173] Those skilled in the art should understand that the embodiments of the present invention described above and shown in the accompanying drawings are only examples and do not limit the present invention. The objectives of the present invention have been fully and effectively achieved. The functions and structural principles of the present invention have been demonstrated and illustrated in the embodiments. Without departing from the said principles, any changes or modifications may be made to the embodiments of the present invention.
Claims
1. A processing method supporting multi-source heterogeneous data, characterized in that: The method comprises: Collect data from different target systems and add identification tags of the corresponding target systems to each piece of data; Build multiple real-time data processing systems in parallel with the ETL path to process the collected data, write the parsing failure data into the error data area, and generate warning information; remove the parsing failure data from the collected data and write it into the temporary storage area; Write the output of the real-time data processing system and / or the output of the ETL path into the data warehouse; Use the data in the data warehouse to create visual dynamic reports, analyze and predict the dynamic report data through pre-trained machine learning models, and write the prediction results into the dynamic report; The construction of multiple real-time data processing systems in parallel with the ETL path to process the collected data includes: Combine the open source stream processing platform Kafka and the stream data processing engine Flink to build a real-time data processing system. Multiple real-time data processing systems constitute a stream data processing system to process multiple data streams. Get different types of data from different sub-datasets and convert the data into a format supported by Kafka. If the target system only supports batch file transfer, set up a batch file transfer plug-in in Kafka to pull data from the temporary storage area. After cleaning the acquired data set, data is collected in real time through a sliding time window to form a data stream; Write the data stream into the preset target storage area and lock it. After Flink completes the state point detection, unlock the target storage area and output the data. Maintain the original ETL path of the target system; Build a gradual switching strategy to control the process of switching from ETL to real-time data processing system, and gradually switch the ETL data processing tasks of the target system to the real-time data processing system.
2. A processing method supporting multi-source heterogeneous data according to claim 1, characterized in that: The successfully parsed data constitutes a correct data set, and the unparsed data constitutes an incorrect data set; according to the data structure type, the clustering algorithm is used to divide the data in the correct data set into multiple sub-data sets; Before acquiring data from different target systems, the steps also include: defining a global strategy to unify the permissions of each target system; Defining a global strategy includes: Register the target system in the console of the security management framework Ranger, and import the resource hierarchy of the target system, wherein the resource hierarchy includes tables and column families; Select appropriate resources based on the resource level of each target system; Specify the operations corresponding to the resource, including allow operations and deny operations; The unified permissions of each target system include: Map the resources of each target system to the unified identifier of Ranger; Convert the native operations of each target system into standard operations of Ranger; Map the permission models of each target system into a unified "resource-operation-condition" model; Create a global role and associate the global role with the corresponding model to bind the resources of the corresponding target system; The conditions include the IP range and the validity period of the operation; The mapping of the permission models of each target system into a unified "resource-operation-condition" model includes: Use NTP protocol to synchronize the clock of Ranger and the target system; Set up extension plug-ins for the permission model of each target system, and use the polling of the extension plug-ins to implement the change of the permission model; Dynamically adjust the polling interval of the extension plug-in to ensure that the permission model of the target system can be mapped in time.
3. A processing method supporting multi-source heterogeneous data according to claim 2, characterized in that: The dynamically adjusting the polling interval of the extension plug-in to ensure that the permission model of the target system can be mapped in time includes: Record the start time of each model mapping; After the permission model completes the mapping, the mapping completion time is returned; The new polling interval is dynamically calculated through the sliding window weighted average delay and proportional integral control algorithm; Monitor the delay and replace the current polling interval with the new polling interval to adjust the polling interval to ensure that the permission model of the target system can be mapped in time.
4. A processing method supporting multi-source heterogeneous data according to claim 3, characterized in that: The new polling interval is dynamically calculated by using a sliding window weighted average delay and a proportional integral control algorithm, including: After each permission model mapping is completed, calculate the time difference ;in Indicates the mapping completion time, Indicates the start time of model mapping; Calculate weighted average delay ,in, represents the size of the sliding window, represents the weight coefficient; Indicates the current time The previous The time delay of a time point; Calculate control variables ;in, represents the target latency threshold, represents the integral gain of one, represents proportional gain of one; Indicates The average delay of a sliding window; Calculate new polling interval ;in, Indicates the current polling interval; Indicates the maximum polling interval. Maximum polling interval; if , replace the current polling interval with , and skip the calculation of the control variables.
5. A processing method supporting multi-source heterogeneous data according to claim 4, characterized in that: Defining a global strategy also includes: Create business semantic tags; Associating business semantic tags with resources in different target systems; Identify global role information and assign different business semantic tags to different global roles. Global roles access corresponding resources through corresponding business semantic tags.
6. A processing method supporting multi-source heterogeneous data according to claim 5, characterized in that: The step-by-step switching strategy is constructed to control the process of switching from ETL to the real-time data processing system, and gradually switch the ETL data processing tasks of the target system to the real-time data processing system, including: Quantify the risk indicators in the switching process; the risk indicators include the data difference rate before and after the switching, the processing delay time of the real-time data processing system, the Flink detection failure rate and the downstream report generation error rate; Write data to Kafka and the original ETL path at the same time, compare the output of the ETL path with the output of the real-time data processing system, and stop switching if the data difference rate between the two outputs is greater than the preset difference rate threshold; Switch non-core task data to the real-time data processing system, while core task data is still transmitted through the original ETL path; Detect the core task data transmission volume. If the core task data transmission volume is lower than the preset transmission volume threshold within the preset time period, the core task data will be switched to the real-time data processing system and the original ETL path will be maintained. Month, ; When the processing delay time of the real-time data processing system exceeds the preset delay time threshold, the real-time data processing system is disconnected and the original ETL path is automatically selected to transmit data.
7. A processing method supporting multi-source heterogeneous data according to claim 6, characterized in that: The step-by-step switching strategy is constructed to control the process of switching from ETL to the real-time data processing system, and gradually switches the ETL data processing tasks of the target system to the real-time data processing system, and further includes: Get real-time data processing system output and ETL path output; Calculate the relative entropy of the real-time data processing system output and the ETL path output: ,in, Represents the probability distribution of key data in the output of real-time data processing system, Indicates the probability distribution of the same key data in the ETL output; represents the Kullback-Leibler divergence; The switching process is dynamically adjusted by the PID controller, including the following sub-steps: Assume that the proportion of data flowing into the real-time data processing system is: ; in, ; represents proportional gain two, represents the integral gain of two, represents the differential gain; Set sliding time window 2 to collect output data from real-time data processing system and ETL path in real time, and calculate and ; use and Calculate and obtain relative entropy; Calculate the change in the proportion of data flowing into the real-time data processing system after a sliding time window ,in, is the length of sliding time window 2; Dynamic optimization through reinforcement learning algorithm PPO , , , including the following sub-steps: Setting up the state space , where the integral term , the differential term ; Setting up the action space ; Among them, the adjustment amount of proportional gain 2 ; Adjustment amount of integral gain 2 , Differential gain adjustment ; Setting up the reward function ; in, represents the error weight, represents the penalty for control quantity change, represents overshoot penalty; Overshoot , Indicates at the current moment The proportion of data that flowed into the real-time data processing system at a previous point in time; The Actor-Critic framework is used to build a reinforcement learning algorithm, which uses the elements in the action space to update the elements in the state space. Updates; The use of elements in the action space to update elements in the state space is specifically: , , ;in, , , Represents the updated proportional gain 2, integral gain 2, and differential gain.
8. A processing method supporting multi-source heterogeneous data according to claim 7, characterized in that: The method of creating a visual dynamic report using the data in the data warehouse, analyzing and predicting the dynamic report data through a pre-trained machine learning model, and writing the prediction results into the dynamic report includes: Obtain data from different target systems from the data warehouse and create dynamic reports using the BI tool Tableau; Load the machine learning model, input the data of different target systems to generate corresponding prediction results, and write them into the dynamic report; the machine learning model includes one or more of the time series prediction model LSTM, the binary classification model XGBoost and the linear regression model; specifically includes: Obtain data from the data warehouse and identify data of different target systems through identification tags of target systems; The data of different target systems are subjected to feature value extraction and then standardized to form a feature set of the target system; Input the elements in the target system feature set as input variables into the machine learning model and output the prediction results; Write the forecast results into the data warehouse, use BI tools to read the forecast results from the data warehouse and write them into dynamic reports, and update the dynamic reports.
9. A processing system supporting multi-source heterogeneous data, comprising a server, a processor, a plurality of data acquisition interfaces connected to the processor, a communication module and a memory, wherein the processor is connected to the server via the communication module, characterized in that: The system is used to execute a processing method supporting multi-source heterogeneous data as described in any one of claims 1 to 8.
Citation Information
Patent Citations
Data fusion method for information security multi-source heterogeneous data
CN115001793A
Cited By
Real-time streaming analysis method and system based on multi-source heterogeneous data
CN121744304A