Cross-platform data integration optimization method and device, equipment and storage medium
By acquiring and processing multiple data source information, format conversion and task scheduling, the problem of insufficient data format diversity and real-time in cross-platform data integration is solved, and efficient and real-time data integration and optimization effects are achieved.
Patent Information
- Application Number
- CN202411986125.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-27
AI Technical Summary
In big data and multi-cloud environments, cross-platform data integration faces problems such as data format diversity, system protocol incompatibility and poor real-time requirements, resulting in low integration efficiency, insufficient real-time and waste of resources.
By obtaining multiple data source information, determining their type, acquiring access integration data, format conversion, generating target format data, and scheduling these data to achieve optimization processing.
It realizes the efficiency and real-time performance of cross-platform data integration, and is suitable for a variety of data sources, data formats and business scenarios, optimizing and improving overall performance.
Smart Images

Figure CN120045611A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of data optimization, and particularly to a cross-platform data integration optimization method, device, equipment, and storage medium. Background Art
[0002] In the current big data and multi-cloud environment, enterprises usually use multiple different platforms and systems to store and process data. Therefore, it is necessary to integrate the data of multiple platforms, but the data of different platforms often have problems such as diverse data formats, incompatible system protocols, and poor real-time requirements. These problems pose great challenges to cross-platform data integration.
[0003] In traditional implementation methods, cross-platform data integration mainly relies on static configuration and rule-driven data integration methods. In terms of data source connection and format conversion, users usually need to manually configure connection parameters and conversion rules, and most are oriented to specific data source types and fixed formats. Therefore, they lack automatic adaptability and flexibility. Especially when facing a large number of heterogeneous data sources, dynamic recognition and seamless docking cannot be achieved; at the same time, overall data optimization cannot be realized.
[0004] Moreover, even if the integration operation is achieved through the above manual method, there are still some problems. For example, the integration efficiency is low. Due to the need for a large amount of manual configuration, it is time-consuming and error-prone; the real-time performance is insufficient. The above method has poor support for real-time data synchronization and cannot meet the real-time requirements of some services; there is a lack of intelligent optimization, and the data flow cannot be dynamically optimized, resulting in resource waste and delays. Summary of the Invention
[0005] Based on this, the present application provides a cross-platform data integration optimization method, device, equipment, and storage medium, which can integrate multiple data formats and different data sources, has strong cross-platform compatibility, realizes the high efficiency and real-time performance of cross-platform data integration, is applicable to various data sources, data formats, and business scenarios. At the same time, it can realize the task scheduling of multiple target format data, achieving the effect of optimizing and improving the overall performance.
[0006] In the first aspect, a cross-platform data integration optimization method is provided, and the method includes:
[0007] Obtain data source data, where the data source data includes multiple data source information, and the multiple data source information comes from different data platforms;
[0008] Determine the type of each data source information according to each data source information;
[0009] Obtain multiple access integration data according to the type of each data source information;
[0010] Convert the format of each access integrated data to obtain multiple pieces of data in the target format;
[0011] Perform task scheduling on the multiple pieces of data in the target format to obtain corresponding optimized processed data.
[0012] According to an implementable manner in the embodiments of the present application, each data source information includes a metadata structure and an access protocol; according to each data source information, determine the type of each data source information, including:
[0013] Extract the basic information features of each data source information according to the metadata structure of each data source information;
[0014] Obtain the type of each data source information according to the preset matching rules and the basic information features of each data source information; and / or,
[0015] Obtain the type of each data source information according to the access protocol of each data source information.
[0016] According to an implementable manner in the embodiments of the present application, obtain multiple access integrated data according to the type of each data source information, including:
[0017] Extract the connection parameters of each data source information according to the type of each data source information;
[0018] Generate corresponding multiple connection configuration files according to the connection parameters of each data source information;
[0019] Obtain corresponding multiple access integrated data according to each connection configuration file.
[0020] According to an implementable manner in the embodiments of the present application, convert the format of each access integrated data to obtain multiple pieces of data in the target format, including:
[0021] Identify the data format of each access integrated data, convert each data format to the corresponding standard data format to obtain multiple pieces of data in the standard format;
[0022] Obtain the target format information of each access integrated data;
[0023] Convert each piece of data in the standard format to multiple pieces of data in the target format according to the preset conversion rule table and the target format information of each.
[0024] According to an implementable manner in the embodiments of the present application, perform task scheduling on the multiple pieces of data in the target format to obtain corresponding optimized processed data, including:
[0025] Obtain the priority information and load resource information of the multiple pieces of data in the target format;
[0026] Task scheduling is performed on multiple target format data according to each priority information, load resource information, and a preset task scheduling algorithm to obtain corresponding optimized processed data.
[0027] According to an implementable manner in an embodiment of the present application, after the step of obtaining multiple access integration data, the method further includes:
[0028] Clean and verify each access integration data to obtain corresponding multiple access integration corrected data;
[0029] Convert the format of each access integration corrected data to obtain multiple standard format data; and / or,
[0030] Perform security detection on each access integration data, configure access permissions for each access integration data, and obtain access integration permission data;
[0031] Convert the format of each access integration permission data to obtain multiple standard format data.
[0032] According to an implementable manner in an embodiment of the present application, the method further includes:
[0033] Obtain performance indicators for each optimized processed data;
[0034] Perform performance analysis on each performance indicator to obtain corresponding multiple performance analysis results;
[0035] Optimize and adjust each optimized processed data according to each performance analysis result.
[0036] In a second aspect, a cross-platform data integration optimization device is provided, and the device includes:
[0037] A data acquisition module, configured to acquire data source data, where the data source data includes multiple data source information, and the multiple data source information comes from different data platforms;
[0038] A type determination module, configured to determine the type of each data source information according to each data source information;
[0039] An access data module, configured to obtain multiple access integration data according to the type of each data source information;
[0040] A format conversion module, configured to convert the format of each access integration data to obtain multiple target format data;
[0041] A task scheduling module, configured to perform task scheduling on multiple target format data to obtain corresponding optimized processed data.
[0042] In a third aspect, a computer device is provided, including:
[0043] At least one processor; and
[0044] a memory communicatively connected to the at least one processor; wherein,
[0045] the memory stores computer instructions executable by the at least one processor, and the computer instructions are executed by the at least one processor to enable the at least one processor to execute the method involved in the above first aspect.
[0046] In a fourth aspect, there is provided a computer-readable storage medium having computer instructions stored thereon, characterized in that the computer instructions are used to cause a computer to execute the method involved in the above first aspect.
[0047] According to the technical content provided by the embodiments of the present application, data source data is obtained, wherein the data source data includes a plurality of data source information, and the plurality of data source information comes from different data platforms; according to each data source information, the type of each data source information is determined; according to the type of each data source information, a plurality of access integration data are obtained; the format of each access integration data is converted to obtain a plurality of target format data; task scheduling is performed on the plurality of target format data to obtain corresponding optimized processing data. The above operations can integrate multiple data formats and different data sources, have strong cross-platform compatibility, realize the efficiency and real-time nature of cross-platform data integration, are applicable to multiple data sources, data formats and business scenarios. At the same time, task scheduling of multiple target format data can be realized to achieve the effect of optimizing and improving the overall performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] Figure 1 is an application environment diagram of a cross-platform data integration optimization method in an embodiment;
[0049] Figure 2 is a flowchart of a cross-platform data integration optimization method in an embodiment;
[0050] Figure 3 is a preferred flowchart of a cross-platform data integration optimization method in an embodiment;
[0051] Figure 4 is a structural block diagram of a cross-platform data integration optimization device in an embodiment;
[0052] Figure 5 is a schematic structural diagram of a computer device in an embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0054] A cross-platform data integration optimization method provided by this application can be applied to an application environment as shown in Figure 1 In the application environment shown. Among them, the terminal 102 communicates with the server 104 through a network. Specifically, the server 104 obtains the data source data of the terminal 102. Among them, the data source data includes multiple data source information, and the multiple data source information comes from different data platforms; according to each data source information, determine the type of each data source information; according to the type of each data source information, obtain multiple access integration data; perform format conversion on each access integration data to obtain multiple target format data; perform task scheduling on the multiple target format data to obtain corresponding optimized processing data. Among them, the terminal 102 can be but is not limited to various personal computers, laptop computers, smart phones, tablet computers, etc., and the server 104 can be implemented by an independent server or a server cluster composed of multiple servers.
[0055] In one embodiment, as shown in Figure 2 A cross-platform data integration optimization method is provided. Taking the server in Figure 1 as an example for illustration, the method includes the following steps:
[0056] Step S201: Obtain data source data.
[0057] Among them, the data source data includes multiple data source information, and the multiple data source information comes from different data platforms.
[0058] Here, since the data source data includes multiple data source information, different data source information can come from different systems and platforms. For example: the multiple data source information can come from relational databases, NoSQL (Not Only SQL) databases, and cloud services, etc. Specifically, the server can automatically detect the multiple data source information of different data platforms, that is, the data source data.
[0059] Step S203: Determine the type of each data source information according to each data source information.
[0060] Among them, the data source types include but are not limited to relational databases, NoSQL databases, file systems, cloud storage, and API (Application Programming Interface) interfaces, etc.
[0061] Here, different data source information corresponds to different types. Therefore, based on each data source information, the type of each data source information can be determined through an adaptive connector method.
[0062] Step S205: Obtain multiple access integration data according to the type of each data source information.
[0063] Here, according to the types of information of each data source, connection configurations can be automatically generated without manually inputting any complex parameters. Furthermore, it is possible to access the information of multiple data sources to the current server, implement integrated operations on the information of multiple data sources, and obtain the integrated access data.
[0064] Step S207: Convert the format of each integrated access data to obtain multiple data in target formats.
[0065] Here, since each integrated access data comes from different systems and platforms, the formats of each integrated access data may be completely different. Therefore, the formats of each integrated access data can be converted to obtain data in target formats.
[0066] Step S209: Perform task scheduling on multiple data in target formats to obtain corresponding optimized processed data.
[0067] Here, each data in target format is a running task, and multiple data in target formats form multiple running tasks. By performing task scheduling on multiple running tasks, since task scheduling itself can achieve resource optimization, corresponding optimized processed data can be obtained.
[0068] It can be seen that in the embodiment of the present application, by obtaining data source data, where the data source data includes multiple data source information, and the multiple data source information comes from different data platforms; determining the types of each data source information according to each data source information; obtaining multiple integrated access data according to the types of each data source information; converting the format of each integrated access data to obtain multiple data in target formats; performing task scheduling on multiple data in target formats to obtain corresponding optimized processed data. The above operations can integrate multiple data formats and different data sources, have strong cross-platform compatibility, achieve the efficiency and real-time performance of cross-platform data integration, are applicable to multiple data sources, data formats, and business scenarios. At the same time, task scheduling of multiple data in target formats can be realized to achieve the effect of optimizing and improving the overall performance.
[0069] The following will describe the different steps in the above method flow in detail. First, the above step 203, that is, "determining the types of each data source information according to each data source information", will be described in detail in combination with an embodiment.
[0070] Clean and verify each integrated access data to obtain corresponding multiple corrected integrated access data; convert the format of each corrected integrated access data to obtain multiple data in target formats; and / or perform security detection on each integrated access data, configure access permissions for each integrated access data, and obtain integrated access permission data;
[0071] Convert the format of each access integration permission data to obtain multiple target format data.
[0072] In an implementable manner, extract the basic information features of each data source information according to the metadata structure of each data source information; obtain the types of each data source information according to the preset matching rules and the basic information features of each data source information.
[0073] Among them, each data source information includes a metadata structure and an access protocol. The metadata structure includes, but is not limited to, JDBC (Java Database Connectivity), URL (Uniform Resource Locator), API (Application Programming Interface) files, and file extensions, etc.
[0074] Here, since the data source information includes a metadata structure, the server can extract the basic information features of the corresponding data source information by accessing the metadata structure of the data source information; according to the preset matching rules, map the basic features of each data source information to the known predefined data source types to obtain the types of each data source information. Among them, the preset matching rules include, but are not limited to, preset pattern matching and rule engines, etc. For example, the metadata structure with a json (JavaScript Object Notation) file extension can be recognized as a JSON type file.
[0075] In another implementable manner, obtain the types of each data source information according to the access protocol of each data source information.
[0076] Here, the current server can support multiple access protocols, such as SQL (Structured Query Language), REST (Representational State Transfer), SOAP (Simple Object Access Protocol), FTP (File Transfer Protocol), S3, and file system protocols, etc. For different access protocols, there are corresponding different abstract interfaces, such as query, write, and synchronization interfaces, etc. The functions of different access protocols are different. For example, SQL can execute queries and result set parsing; REST can construct HTTP requests and parse JSON responses, etc.
[0077] Specifically, since the data source information also includes an access protocol, based on the access protocols of the respective data source information, the types of the respective data source information can be obtained. For example, the access protocol jdbc:mysql: / / host:port can be recognized as the MySQL database type.
[0078] Through the above operations, by obtaining the metadata structure and access protocol of the data source information, the corresponding source data types can be obtained, realizing the automatic recognition of diverse data sources, facilitating the subsequent integration of all data source information, achieving support for multiple data formats and different data sources, and having a strong cross-platform compatibility effect.
[0079] Secondly, the above step S205, that is, "obtaining multiple access integration data according to the types of the respective data source information", will be described in detail in combination with the embodiments.
[0080] According to the types of the respective data source information, extract the connection parameters of the respective data source information; according to the connection parameters of the respective data source information, generate corresponding multiple connection configuration files; according to the respective connection configuration files, obtain corresponding multiple access integration data.
[0081] Here, according to the types of the respective data source information, the connection parameters of the corresponding data source information can be extracted, including but not limited to the host address, port, and authentication method, etc. According to the connection parameters of the respective data source information, multiple connection configuration files or connection URLs (Uniform Resource Locator) can be dynamically generated to adapt to the data source information, realizing a one-to-one correspondence between each data source information and each connection configuration file. For example, for database source information, JDBC, URL, and authentication information can be automatically generated; for API data source information, an access configuration file containing an API Key or OAuth Token can be automatically generated. According to the respective connection configuration files, the corresponding connectors are automatically loaded to obtain corresponding multiple access integration data. It should be noted that there may be some unrecognized data source information, and extensible plugins can be preset in advance so that users can customize connectors to achieve automatic connection.
[0082] Through the above operations, according to the types of the respective data source information, appropriate connectors can be dynamically loaded and initialized, realizing cross-platform data access and processing and completing the access integration of data sources, providing a unified entry for subsequent format conversion, achieving the effect of reducing manual configuration work and improving data integration efficiency.
[0083] Next, the above step S207, that is, "performing format conversion on each access integration data to obtain multiple target format data", will be described in detail in combination with the embodiments.
[0084] Identify the data formats of each access integrated data, convert each data format into the corresponding standard data format to obtain multiple standard format data; obtain the target format information of each access integrated data; and convert each standard format data into multiple target format data according to the preset conversion rule table and each target format information.
[0085] Among them, the data formats include but are not limited to JSON (JavaScript Object Notation), XML (eXtensible Markup Language), CSV (Comma-Separated Values), and Parquet (columnar storage) and other formats.
[0086] Here, automatically identify the data format of each access integrated data. The specific identification methods include but are not limited to file format identification, that is, quickly identify based on the file extension, such as.json,.xml, and.csv, etc.; for files without extensions, identify through the magic number or header characteristics of the data content; use regular expressions or pattern matching methods to identify the structural characteristics of the access integrated data, such as XML tags and JSON object levels, etc., and apply predefined rules or machine learning classification models for format type identification, etc. After identifying the corresponding data format, it can be parsed and placed in the memory of the current server for subsequent format conversion.
[0087] For different data formats, corresponding parsing logics can be provided. For example, open-source parsing libraries, namely Jackson, FastXML, Apache Commons, and CSV, etc., can be used for parsing; nested formats, namely JSON and XML, etc., can also be used to recursively generate in-memory data objects; or through dynamic field mapping, the hierarchical structure of the data can be flattened or reorganized into the corresponding standard data format. After the explanation is completed, it can be stored in the memory of the server. It should be noted that for parsing errors, that is, situations such as format inconsistency and field missing, the problem data can be marked through a fault tolerance mechanism and log records for subsequent operators to process.
[0088] Convert each data format into the corresponding standard data format to obtain multiple standard format data. Among them, the standard data format can adopt a general structure, such as JSON or Parquet, etc., as the intermediate data model, that is, the standard data format, to unify the field naming and type representation. Here, to convert the corresponding data format into the standard data format, the methods adopted include but are not limited to automatically matching the source data fields and the standard data format fields to complete the field name mapping; standardizing different format data types, such as dates, numbers, and booleans, into a unified type; expanding or nesting multi-level nested structures to meet the requirements of a unified format.
[0089] Obtain the target format information of each access integrated data; according to the preset conversion rule table and each target format information, convert each standard format data into multiple target format data. Here, obtain the target format information of each access integrated data input by the user and perform conversion based on the preset conversion rules to obtain multiple target format data. Among them, the preset conversion rules include but are not limited to rule engine drive, that is, adopt the preset format conversion rule table to adjust the field format, order, data type, etc. according to the target format information; templated conversion, that is, for a fixed output format, use a templated method to quickly generate data that meets the target requirements; data compression and serialization, that is, for a large amount of data, adopt an efficient serialization method to achieve conversion.
[0090] The above method converts the standard data format into the target format data, achieving the effect of supporting multiple output requirements, ensuring the integrity and consistency of the data content, providing a unified format requirement for subsequent data synchronization and analysis, and eliminating format compatibility obstacles.
[0091] Finally, in combination with the embodiments, the above step S209, that is, "perform task scheduling on multiple target format data to obtain corresponding optimized processed data", is described in detail.
[0092] Obtain the priority information and load resource information of multiple target format data; according to each priority information, load resource information, and the preset task scheduling algorithm, perform task scheduling on multiple target format data to obtain corresponding optimized processed data.
[0093] Among them, the preset task scheduling algorithm includes but is not limited to the preset task priority scheduling algorithm and the preset load balancing algorithm.
[0094] Here, each target format data corresponds to a task, and each task has its corresponding priority information, that is, the task allocation is dynamically adjusted according to the task importance or time sensitivity. Specifically, a preset task priority scheduling algorithm can be adopted based on the priority information for task scheduling. The specific scheduling process is as follows: Since each task is assigned a priority, that is, it is allocated according to the business urgency of the task or user-defined rules. For example, it can be divided into high priority, medium priority, and low priority, etc. High-priority tasks are executed first, and low-priority tasks are postponed or compressed appropriately. At the same time, Deadline-aware Scheduling can be combined to ensure that time-sensitive tasks are completed before the deadline. A priority queue can be used to store tasks to be processed, dynamically adjust the task order, and reorder according to the priority weights updated in real time to ensure the priority execution of critical tasks.
[0095] Since multiple tasks may be assigned to different task platforms for execution, therefore, the load resource information of multiple tasks, that is, each task platform, can be obtained. Among them, the load resource information includes the CPU, memory, and network bandwidth of each task platform. Specifically, a preset load balancing algorithm can be adopted based on the load resource information for task scheduling. This method is applied to multi-platform and multi-node scenarios to achieve the effect of balancing resource utilization. Its specific scheduling process is as follows: The hash consistency algorithm is adopted, and tasks are efficiently allocated through Round Robin or Least Load methods to ensure relatively balanced loads on each resource node. At the same time, combined with historical task load data, by monitoring the CPU, memory, etc. of each node, the future resource consumption of tasks is predicted, and the task allocation strategy is dynamically adjusted.
[0096] It should be noted that the preset task scheduling algorithms also include a preset genetic scheduling algorithm, a preset reinforcement learning scheduling algorithm, and a preset dynamic resource prediction scheduling algorithm, etc. Among them, the preset genetic scheduling algorithm can be applied to scenarios with a large number of tasks and complex resource constraints that require global optimization scheduling. Its specific scheduling method is as follows: By simulating the biological evolution process, through the crossover and mutation of populations, an optimal task allocation scheme can be found to solve the multi-objective optimization problem, and a balance can be achieved between minimizing resource occupancy and maximizing task throughput. The preset reinforcement learning scheduling algorithm can be applied to scenarios with strong dynamics, such as real-time data stream processing, and its scheduling rules need to be automatically adjusted according to the server system state. The specific scheduling method of this algorithm is as follows: Using preset models such as Q-Learning or Deep Reinforcement Learning, train an intelligent agent to dynamically select a scheduling strategy according to the current server system state, such as resource load, task queuing, etc., to achieve the effect of improving the accuracy of decision-making. The preset dynamic resource prediction scheduling algorithm can be applied to scenarios where the server system load fluctuates greatly and resource planning is required in advance. Its specific scheduling method is as follows: Based on historical data, use time series prediction algorithms such as ARIMA or LSTM to predict future task volumes and resource requirements, and allocate or release resources in advance according to the prediction results to avoid resource waste or shortage and dynamically adjust the resource allocation strategy.
[0097] It should be emphasized that the above-mentioned preset scheduling algorithms can be used alone or simultaneously in this application document, depending on the specific needs of the user.
[0098] Through the above operations, the preset task scheduling algorithm is used to schedule multiple tasks to achieve the allocation of task resources, thereby improving the utilization rate of system resources and data processing efficiency.
[0099] In one embodiment, after the step of obtaining multiple access integration data, the method further includes:
[0100] Cleaning and verifying each access integration data to obtain corresponding multiple access integration corrected data; converting the format of each access integration corrected data to obtain multiple standard format data; and / or, performing security detection on each access integration data and configuring access permissions for each access integration data to obtain multiple standard format data.
[0101] In a feasible manner, clean and verify each access integration data to obtain corresponding multiple access integration corrected data; convert the format of each access integration corrected data to obtain multiple standard format data.
[0102] Here, cleaning and verification are performed on each access integrated data. The specific cleaning operations include removing duplicate, incomplete, and invalid access integrated data, etc. Based on the automated verification rules, the data quality of the access integrated data is verified. The specific verification rules include, but are not limited to: integrity rules, that is, whether there are missing values in each access integrated data, ensuring that key fields are not empty. For example, whether there are relevant records in the main table referenced by the foreign key; accuracy rules, that is, verifying whether each access integrated data is within the predefined range. For example, the age is between 0 and 120, and for the time field, check its format and logical correctness. For example, the end time should be later than the start time; consistency rules, that is, whether each access integrated data is consistent among different systems. For example, a unified time format, currency unit, etc., and check whether the field values meet the constraint conditions. For example, the gender field can only be "male" or "female"; uniqueness rules, that is, ensuring that certain fields of each access integrated data have unique values in the dataset. For example, the ID number and order number; normalization rules, that is, determining whether the format of each access integrated data conforms to the standard. For example, the email and phone number formats, and check whether there are extra spaces or illegal characters in the string fields; anomaly detection rules, that is, based on the historical data distribution, identify possible abnormal data points in the current access integrated data, and use machine learning models to mark the abnormal data. For example, the deviation between the predicted value and the actual value exceeds the threshold; timeliness rules, that is, whether each access integrated data is updated within the specified time range. For real-time data, ensure that the latency is within the allowable range. Based on the foregoing verification rules, the verification is completed, and automatic correction or prompting the operator to make corrections, etc. are realized, and then multiple access integrated corrected data are obtained. The format of each access integrated corrected data is converted to obtain multiple standard format data.
[0103] Through the above operations, by correcting the access integrated data, the accuracy, integrity, and consistency of the access integrated data are ensured, achieving the effect of improving data quality and avoiding performance deviation caused by data problems.
[0104] In another realizable way, security detection is performed on each access integrated data, access permissions are configured for each access integrated data to obtain access integrated permission data; the format of each access integrated permission data is converted to obtain multiple standard format data.
[0105] Here, the security detection of the transmission process of each access integrated data is carried out through the TLS (Transport Layer Security) or SSL (Secure Sockets Layer) encryption technology, and access permissions are configured for each access integrated data, that is, the operation scope is restricted according to the user role, so as to obtain the access integrated permission data, and the format of each access integrated permission data is converted to obtain multiple standard format data. At the same time, the log audit function can also be enabled to record data operation behaviors for security inspection and compliance audit, etc.
[0106] Through the above operations, by configuring access permissions, the transmission security of cross-platform data can be ensured, providing security guarantees for the whole process.
[0107] In one embodiment, the method further includes: obtaining the performance indicators of each optimized processing data; performing performance analysis on each performance indicator to obtain corresponding multiple performance analysis results; and performing optimization adjustment on each optimized processing data according to each performance analysis result.
[0108] Among them, the performance indicators include but are not limited to processing speed, error rate, and resource usage, etc.
[0109] Regularly obtain the performance indicators of each optimized processing data, perform performance analysis on each performance indicator to discover bottlenecks or abnormal links, and then obtain corresponding multiple performance analysis results. According to each performance analysis result, perform optimization adjustment on each optimized processing data.
[0110] The specific optimization adjustment process includes but is not limited to: real-time resource allocation optimization, that is, dynamically adjusting the allocation of CPU, memory, and network bandwidth to avoid resource waste or bottlenecks, and automatically increasing resource instances for peak tasks; task priority optimization adjustment, that is, adjusting the task execution order according to the real-time task load, pausing low-priority tasks when resources are insufficient, and giving priority to key tasks; load balancing optimization adjustment, dynamically adjusting the task allocation nodes to avoid single-point overload, using historical load data to predict future resource requirements, and scheduling in advance; data flow optimization adjustment, that is, compressing transmitted data according to the data flow situation to reduce network load, optimizing the data synchronization frequency, and reducing unnecessary real-time update costs; abnormal self-healing optimization adjustment, when detecting abnormalities in the server system, such as task failure or performance degradation, automatically restarting tasks or switching to standby nodes to reduce the probability of future errors; scheduling strategy learning and optimization adjustment, using machine learning algorithms to automatically optimize the scheduling algorithm according to historical operation data, gradually improving resource allocation and task scheduling rules, and improving overall efficiency.
[0111] Based on the performance analysis results, the above operations dynamically adjust the optimization strategy to form a closed-loop feedback, which acts on steps such as data source identification, synchronization, and scheduling, enhancing the stability and efficiency of the overall server system.
[0112] It should be noted that in an implementable manner, the method can also listen for change events of each data source information through an event-driven architecture. When there is a change in the data source information, the event handling mechanism is triggered to push the changed information to a message queue, such as an Apache Kafka or RabbitMQ queue, etc. The server reads the changed data from the message queue and synchronizes it to the current target format data in real time. This realizes real-time data update between multiple platforms, meets the business requirements for data timeliness, ensures data consistency, and enables the quick reflection of the updated data in the next resource scheduling.
[0113] In another implementable manner, the method further includes returning the process information of the entire task creation and execution to a visualization monitoring and operation platform. Specifically, this platform can display the following information: the running status of the server system, that is, the core performance metrics of the system are displayed in real time, such as the CPU, memory, and disk I / O utilization rates, etc., as well as the current task queue length and task completion rate, etc.; the data flow situation, that is, the real-time graph of the data flow from the data source to the target platform, the data processing speed and latency of each platform, etc.; the task execution details, that is, the execution progress and status of each task, such as tasks being executed, completed, or failed, as well as the estimated completion time and used resources of the tasks, etc.; exception and alarm information, that is, data quality exceptions, such as the real-time statistics of data duplication rate and missing rate, server system resource alarms, and network transmission exceptions, such as high latency or packet loss; historical data statistics, that is, the task completion situation, data transmission volume, and resource utilization trend graph over a past period of time, etc.; dynamic optimization feedback, that is, the implementation situation of the server system optimization strategy, such as the adjusted task priority distribution, and the comparison before and after resource load balancing optimization; security and permission management, that is, data access logs and user operation records, as well as the permission distribution and actual operation statistics of different roles. Based on the information displayed above, operators can obtain a comprehensive view of the running status of the server system for adjustment to achieve the effect of improving operation and maintenance efficiency. At the same time, an automatic alarm can be triggered when the server system is abnormal, and the operation and maintenance personnel can be notified by means such as email and text message.
[0114] Combined with the implementation manners in the above embodiments, the following Figure 3 gives an example description of a preferred method flow provided by the embodiments of the present application. As Figure 3 shown, the method may include the following steps:
[0115] Step S301: Obtain data from data sources. The data from data sources includes multiple data source information, and the multiple data source information comes from different data platforms; each data source information includes a metadata structure and an access protocol.
[0116] Step S302: Extract the basic information features of each data source information according to the metadata structure of each data source information.
[0117] Step S303: Obtain the type of each data source information according to the preset matching rules and the basic information features of each data source information, and obtain the type of each data source information according to the access protocol of each data source information.
[0118] Step S304: Extract the connection parameters of each data source information according to the type of each data source information.
[0119] Step S305: Generate corresponding multiple connection configuration files according to the connection parameters of each data source information.
[0120] Step S306: Obtain corresponding multiple access integrated data according to each connection configuration file.
[0121] Step S307: Identify the data formats of each access integrated data, convert each data format into the corresponding standard data format, and obtain multiple standard format data.
[0122] Step S308: Obtain the target format information of each access integrated data.
[0123] Step S309: Convert each standard format data into multiple target format data according to the preset conversion rule table and each target format information.
[0124] Step S310: Obtain the priority information and load resource information of multiple target format data.
[0125] Step S311: Perform task scheduling on multiple target format data according to each priority information, load resource information and the preset task scheduling algorithm, and obtain the corresponding optimized processed data.
[0126] It should be understood that although Figure 1 、 Figure 3 each step in the flowchart of Figure 1 、 Figure 3At least a part of the steps may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0127] Figure 4 FIG. 4 is a schematic structural diagram of a cross-platform data integration optimization device provided by an embodiment of the present application. The device can be set in Figure 1 the server in the application environment shown, and is used to execute the method flow shown in Figure 1 , Figure 3 shown in. As shown in Figure 4 shown, the device may include: a data acquisition module 401, a type determination module 403, an access data module 405, a format conversion module 407, and a task scheduling module 409. The main functions of each component module are as follows:
[0128] The data acquisition module 401 is used to acquire data source data. Among them, the data source data includes multiple data source information, and the multiple data source information comes from different data platforms;
[0129] The type determination module 403 is used to determine the type of each data source information according to each data source information;
[0130] The access data module 405 is used to obtain multiple access integration data according to the type of each data source information;
[0131] The format conversion module 407 is used to perform format conversion on each access integration data to obtain multiple target format data;
[0132] The task scheduling module 409 is used to perform task scheduling on multiple target format data to obtain corresponding optimized processing data.
[0133] In one embodiment, each data source information includes a metadata structure and an access protocol. The type determination module 403 is further used to:
[0134] Extract the basic information features of each data source information according to the metadata structure of each data source information;
[0135] Obtain the type of each data source information according to the preset matching rule and the basic information features of each data source information; and / or,
[0136] Obtain the type of each data source information according to the access protocol of each data source information.
[0137] In one embodiment, the access data module 405 is further used to:
[0138] Extract the connection parameters of each data source information according to the type of each data source information;
[0139] Generate corresponding multiple connection configuration files according to the connection parameters of each data source information;
[0140] Obtain corresponding multiple access integrated data according to each connection configuration file.
[0141] In one embodiment, the format conversion module 407 is further configured to:
[0142] Identify the data formats of each access integrated data, convert each data format into a corresponding standard data format, and obtain multiple standard format data;
[0143] Obtain the target format information of each access integrated data;
[0144] Convert each standard format data into multiple target format data according to the preset conversion rule table and each target format information.
[0145] In one embodiment, the task scheduling module 409 is further configured to:
[0146] Obtain the priority information and load resource information of multiple target format data;
[0147] Perform task scheduling on multiple target format data according to each priority information, load resource information, and the preset task scheduling algorithm, and obtain corresponding optimized processed data.
[0148] In one embodiment, the device is further configured to:
[0149] Clean and verify each access integrated data to obtain corresponding multiple access integrated corrected data;
[0150] Convert the format of each access integrated corrected data to obtain multiple standard format data; and / or,
[0151] Perform security detection on each access integrated data, configure access permissions for each access integrated data, and obtain access integrated permission data;
[0152] Convert the format of each access integrated permission data to obtain multiple standard format data.
[0153] In one embodiment, the device is further configured to:
[0154] Obtain the performance metrics for each optimized processed data;
[0155] Perform performance analysis on each performance metric to obtain corresponding multiple performance analysis results;
[0156] Optimize and adjust each piece of optimized processing data according to the results of each performance analysis.
[0157] For the same or similar parts among the above embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the apparatus embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and for the relevant parts, reference can be made to the partial description of the method embodiments.
[0158] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, within the scope permitted by the applicable laws and regulations of the country where it is located (such as when the user clearly consents, is effectively notified to the user, and the user clearly authorizes, etc.), user-specific personal data can be used in the solutions described herein within the scope permitted by the applicable laws and regulations.
[0159] According to the embodiments of the present application, the present application also provides a computer device and a computer-readable storage medium.
[0160] As Figure 5 shown, it is a block diagram of a computer device according to an embodiment of the present application. The computer device is intended to represent various forms of digital computers or mobile devices. Among them, the digital computer may include a desktop computer, a portable computer, a workbench, a personal digital assistant, a server, a mainframe computer, and other suitable computers. The mobile device may include a tablet computer, a smart phone, a wearable device, etc.
[0161] As Figure 5 shown, the device 500 includes a computing unit 501, a ROM 502, a RAM 503, a bus 504, and an input / output (I / O) interface 505. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other through the bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.
[0162] The computing unit 501 can execute various processes in the method embodiments of this application according to computer instructions stored in the read-only memory (ROM) 502 or computer instructions loaded from the storage unit 508 into the random access memory (RAM) 503. The computing unit 501 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. The computing unit 501 can include, but is not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. In some embodiments, the method provided in the embodiments of this application can be implemented as a computer software program, which is tangibly contained in a computer-readable storage medium, such as the storage unit 508.
[0163] The RAM 503 can also store various programs and data required for the operation of the device 500. Part or all of the computer programs can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509.
[0164] The input unit 506, output unit 507, storage unit 508, and communication unit 509 in the device 500 can be connected to the I / O interface 505. Among them, the input unit 506 can be, such as, a keyboard, a mouse, a touch screen, a microphone, etc.; the output unit 507 can be, such as, a display, a speaker, an indicator light, etc. The device 500 can exchange information, data, etc. with other devices through the communication unit 509.
[0165] It should be noted that this device may also include other components necessary for normal operation. It may also only include the components necessary to implement the solution of this application, and does not necessarily include all the components shown in the figure.
[0166] Various embodiments of the systems and technologies described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGA), application-specific integrated circuits (ASIC), application-specific standard products (ASSP), systems-on-a-chip (SOC), complex programmable logic devices (CPLD), computer hardware, firmware, software, and / or combinations thereof.
[0167] The computer instructions for implementing the method of this application can be written in any combination of one or more programming languages. These computer instructions can be provided to the computing unit 501, such that when the computer instructions are executed by a computing unit 501 such as a processor, the steps involved in the method embodiments of this application are executed.
[0168] The computer-readable storage medium provided by the present application can be a tangible medium that can contain or store computer instructions for executing the various steps involved in the method embodiments of the present application. The computer-readable storage medium can include, but is not limited to, storage media in the forms of electronic, magnetic, optical, electromagnetic, etc.
[0169] The above specific implementation manners do not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A cross-platform data integration optimization method, characterized in that: The method comprises: Acquire data source data, wherein the data source data includes multiple data source information, and the multiple data source information comes from different data platforms; Determine the type of each piece of data source information according to each piece of data source information; According to the type of each data source information, a plurality of access integrated data are obtained; Performing format conversion on each of the access integrated data to obtain multiple target format data; Task scheduling is performed on the plurality of target format data to obtain corresponding optimized processing data.
2. The method according to claim 1, characterized in that Each of the data source information includes a metadata structure and an access protocol; and determining the type of each of the data source information according to each of the data source information includes: Extracting basic information features of each of the data source information according to the metadata structure of each of the data source information; According to the preset matching rules and the basic information characteristics of each of the data source information, the type of each of the data source information is obtained; and / or, The type of each data source information is obtained according to the access protocol of each data source information.
3. The method according to claim 2, characterized in that The obtaining of a plurality of access integration data according to the type of each of the data source information includes: Extracting connection parameters of each piece of data source information according to the type of each piece of data source information; Generate corresponding multiple connection configuration files according to the connection parameters of each data source information; According to each of the connection configuration files, a corresponding plurality of the access integration data are obtained.
4. The method according to claim 1, characterized in that: The format conversion is performed on each of the access integrated data to obtain a plurality of target format data, including: Identify the data format of each access integrated data, convert each data format into a corresponding standard data format, and obtain a plurality of standard format data; Acquiring target format information of each access integrated data; According to the preset conversion rule table and the target format information, each standard format data is converted into a plurality of target format data.
5. The method according to claim 4, characterized in that The step of performing task scheduling on the plurality of target format data to obtain corresponding optimized processing data includes: Acquire priority information and load resource information of a plurality of target format data; According to the priority information, load resource information and a preset task scheduling algorithm, the plurality of target format data are task scheduled to obtain the corresponding optimized processing data.
6. The method according to any one of claims 1 to 5, characterized in that: After the step of obtaining a plurality of access integrated data, the method further comprises: Cleaning and verifying each access integration data to obtain corresponding multiple access integration correction data; Performing format conversion on each of the access integrated correction data to obtain a plurality of the standard format data; and / or, Perform security detection on each access integration data, configure access permission for each access integration data, and obtain access integration permission data; The format of each access integration authority data is converted to obtain a plurality of standard format data.
7. The method according to any one of claims 1 to 5, characterized in that: The method further comprises: Obtaining performance indicators of each of the optimized processing data; Performing performance analysis on each of the performance indicators to obtain corresponding multiple performance analysis results; According to each of the performance analysis results, each of the optimization processing data is optimized and adjusted.
8. A cross-platform data integration optimization device, characterized in that: The device comprises: A data acquisition module is used to acquire data source data, wherein the data source data includes a plurality of data source information, and the plurality of data source information comes from different data platforms; A type determination module, used to determine the type of each piece of data source information according to each piece of data source information; An access data module, used to obtain multiple access integrated data according to the type of each data source information; A format conversion module, used for performing format conversion on each of the access integrated data to obtain a plurality of target format data; The task scheduling module is used to perform task scheduling on the plurality of target format data to obtain corresponding optimized processing data.
9. A computer device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores computer instructions that can be executed by the at least one processor, and the computer instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having computer instructions stored thereon, characterized in that: The computer instructions are used to make a computer execute the method according to any one of claims 1 to 7.