Cross-platform big data sharing and collaborative intelligent processing method
Through cross-platform big data sharing and collaborative intelligent processing methods, the technical problems of data interoperability and integration between different data platforms are solved, efficient data sharing and collaborative processing are achieved, and data consistency and reliability are ensured.
Patent Information
- Application Number
- CN202411898291.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-23
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The data format, storage structure and access protocols between different data platforms vary greatly, resulting in technical difficulties in cross-platform data interoperability and integration. When dealing with large-scale heterogeneous data, problems such as system consistency, resource scheduling and fault handling are complex.
Cross-platform big data sharing and collaborative intelligent processing methods are adopted, including data structure standardization, data integration and virtualization, access control-based data sharing mechanism, collaborative processing task scheduling and resource allocation, cross-platform collaborative analysis and automatic optimization, data consistency control and fault tolerance mechanism and other technical means.
It realizes seamless integration and interoperability of data between different platforms, improves system compatibility and scalability, reduces the complexity of data integration, improves the collaboration efficiency between platforms, and ensures the consistency and reliability of data in a distributed environment.
Smart Images

Figure CN119988468A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data processing technology, and in particular to a cross-platform big data sharing and collaborative intelligent processing method. Background Art
[0002] With the rapid development of information technology and the Internet, various data processing platforms and systems are constantly emerging, and data sharing and collaborative processing between different platforms are becoming increasingly important. Especially in application scenarios such as big data, the Internet of Things, and cloud computing, data is distributed on multiple heterogeneous platforms, including databases, file systems, and cloud services. The storage structure and access protocols of data vary greatly. How to achieve efficient collaboration, data sharing, and intelligent processing between these platforms has become a key issue facing current technology.
[0003] However, due to differences in data formats, storage structures, access protocols, etc. on different platforms, cross-platform data interoperability and integration present significant technical challenges. In addition, as the amount of data continues to increase, it becomes particularly complex to ensure that the system maintains consistency when processing large-scale heterogeneous data, optimize resource scheduling, and effectively respond to system failures. Existing technologies usually rely on a single platform or data format, lack a unified standardized processing framework, and are unable to cope with the challenges brought by heterogeneous data sources. Summary of the invention
[0004] Based on the above objectives, the present invention provides a cross-platform big data sharing and collaborative intelligent processing method.
[0005] The cross-platform big data sharing and collaborative intelligent processing method includes the following steps:
[0006] S1, cross-platform data structure standardization: perform unified structural standardization on raw data from different platforms, and use data mapping and conversion technology to ensure consistency in storage format and access protocol of raw data on each platform;
[0007] S2, data integration and virtualization: The data sources of different platforms are uniformly abstracted and virtualized through the data integration framework to form a cross-platform data virtualization layer. The virtualization layer dynamically adapts the data models of each platform to achieve data interoperability between platforms without relying on the technical details of the specific platform.
[0008] S3, Data sharing mechanism design based on access control model: Design a data sharing mechanism based on role-based access control or attribute-based access control model, and formulate a secure access control strategy according to the permission requirements of different platforms;
[0009] S4, collaborative processing task scheduling and resource allocation: Using intelligent scheduling algorithms, dynamically adjust resource allocation based on the computing power, task type and data volume of each platform to ensure efficient collaborative execution of data processing tasks across multiple platforms, avoiding resource waste and processing bottlenecks;
[0010] S5, cross-platform collaborative analysis and automatic optimization: Combine big data mining and machine learning technologies to collaboratively analyze data from different platforms, automatically identify potential patterns and rules, and optimize processing procedures, data sharing strategies, and system performance;
[0011] S6, data consistency control and fault-tolerance mechanism: By designing a data consistency control mechanism and fault-tolerance mechanism based on a distributed transaction protocol, cross-platform data between platforms can always remain consistent in a distributed environment, and provide data recovery and error handling capabilities in the event of a system failure.
[0012] Optionally, the S1 includes:
[0013] S11, data source identification and classification: Identify data sources from different platforms, identify the types of data sources (such as relational databases, non-relational databases, file systems, etc.), and classify the data storage formats and access protocols (such as SQL, NoSQL, RESTful API, etc.) of each platform to form a metadata model of the data source;
[0014] S12, data format analysis and attribute extraction: Detailed analysis of the data formats of each platform, including the data structure type (such as table type, document type, key-value pair type, etc.), and extraction of key attributes in the data (such as field name, data type, data range, etc.);
[0015] S13, data mapping rule formulation: formulate data mapping rules based on the key attributes extracted in S12;
[0016] S14, data conversion and standardization: According to the data mapping rules formulated in S13, the original data from different platforms are structurally converted, and data conversion algorithms (such as data extraction, conversion, and loading technology in the ETL process) are used to convert data from the source format to the target standard format;
[0017] S15, data consistency verification: after the data conversion is completed, the converted data is verified for consistency;
[0018] S16. Standardized data output: Output the standardized data into a unified format.
[0019] Optionally, S2 includes:
[0020] S21, Data Integration Framework Design: Design and build a cross-platform data integration framework that receives data sources from multiple platforms and abstracts them uniformly;
[0021] S22, construction of data virtualization layer: based on the data integration framework, build a cross-platform data virtualization layer;
[0022] S23, data conversion and standardization: In the data virtualization layer, data conversion algorithms are applied to standardize data on different platforms;
[0023] S24, dynamic data adaptation and model mapping: through the data virtualization layer, dynamic adaptation of data models on each platform is achieved;
[0024] S25, cross-platform data interoperability and data synchronization: With the support of the data virtualization layer, cross-platform data can be interoperable and synchronized in real time. Through the data synchronization mechanism, the data in the virtualization layer can be updated in time between platforms.
[0025] Optionally, the S3 includes:
[0026] S31, role definition and access rights division: define the role model for cross-platform data access, including data owners, data administrators, and data users, and divide the permissions of different roles based on business needs. Each role has a different permission set, which includes data reading, writing, modifying, and deleting.
[0027] S32, formulation of role and permission mapping rules: According to the data security requirements of different platforms, a set of specific permissions is assigned to each role. Through the mapping rules between roles and permissions, each role can access the data set that matches its responsibilities;
[0028] S33, Dynamic allocation and control of access rights: Dynamically adjust access rights in real time based on user operations and role changes in the cross-platform system;
[0029] S34, Data sharing strategy formulation and implementation: Design data sharing strategies based on the roles and permissions.
[0030] Optionally, S3 further includes:
[0031] S35, Attribute definition and classification: Define a multi-dimensional attribute model for users and data, including user attributes (such as department, rank, geographic location, etc.) and data attributes (such as data sensitivity, data source, data type, etc.). By describing users and data in an attributed manner, this provides a basis for access control.
[0032] S36, Attribute-based access control strategy design: Develop access control strategies based on user attributes and data attributes to ensure that data is only accessible to qualified users;
[0033] S37, attribute evaluation and real-time authorization: Evaluate the attributes of users and data in real time, dynamically match the attribute values of users (such as rank, department, access frequency, etc.) with the attribute values of data (such as data sensitivity, data type, etc.), and authorize users to access data;
[0034] S38, multi-level access control implementation: Implement multi-level access control strategies, including layer-by-layer authorization for data access (such as access control at the data level, field level, and record level), and set different security thresholds for different levels of control;
[0035] S39, data sharing strategy implementation and compliance assurance: Ensure that the entire data sharing process complies with data privacy protection and compliance requirements.
[0036] Optionally, the S4 includes:
[0037] S41, Task type and resource requirement analysis: Analyze the data processing tasks on each platform, identify the computing requirements, data access requirements and execution time of each task, and evaluate the requirements for computing resources, storage resources and bandwidth according to the characteristics of the task;
[0038] S42, computing capability evaluation and resource pool management: Evaluate the computing capability of each platform, including the platform's CPU / GPU performance, memory size, and network bandwidth. Based on the evaluation results, aggregate the computing resources, storage resources, and bandwidth resources of each platform into a dynamic resource pool, which is dynamically updated based on the real-time status of the platform.
[0039] S43, Intelligent Scheduling Algorithm Design: Adopt intelligent scheduling algorithm to dynamically adjust the distribution of tasks among platforms according to the resource requirements of tasks, computing power of platforms and system load;
[0040] S44, dynamic task scheduling and load balancing: During the task execution process, the scheduling system dynamically adjusts the task allocation according to the real-time monitoring data and performs load balancing;
[0041] S45, resource release and recycling mechanism: when the task is completed, the occupied computing resources and storage resources are automatically released.
[0042] Optionally, the S5 includes:
[0043] S51, Design of cross-platform data collaborative analysis framework: Design a cross-platform collaborative analysis framework that abstracts data from different platforms into a multidimensional data set;
[0044] S52, Data Pattern and Rule Mining: Based on data mining algorithms, data from different platforms are analyzed to automatically identify potential patterns, correlations and trend changes in the data;
[0045] S53, Pattern Recognition and Data Fusion: Through machine learning algorithms, the analysis results are further processed to fuse global models from heterogeneous data on different platforms;
[0046] S54, Feedback mechanism for cross-platform collaborative analysis results: Establish a cross-platform analysis result feedback mechanism to feed back collaborative analysis results to each platform to achieve data-driven decision support. This feedback mechanism uses real-time data streams (such as Kafka, Apache Pulsar, etc.) to ensure that analysis results can be fed back in a timely manner and promote cross-platform optimization.
[0047] Optionally, S5 further includes:
[0048] S55, automatic optimization objective function design: design the automatic optimization objective function based on the results of cross-platform collaborative analysis;
[0049] S56, optimization algorithm design and implementation: Based on Q-learning reinforcement learning algorithm, the processing flow, data sharing strategy and system performance are automatically optimized according to the objective function;
[0050] S57, Feedback mechanism and dynamic adjustment: By real-time monitoring of the system's operating status, the resource allocation strategy is automatically adjusted according to the optimization results;
[0051] S58, optimization effect evaluation and verification: Evaluate the effect of the optimized system and analyze the optimized resource utilization efficiency, task execution time and data sharing efficiency.
[0052] Optionally, the S6 includes:
[0053] S61, Distributed Transaction Protocol Design: Design a data consistency control mechanism based on the distributed transaction protocol to ensure data consistency across multiple platforms;
[0054] S62, data consistency protocol implementation: adopt a distributed consistency algorithm based on the Paxos protocol or the Raft protocol to implement a distributed consistency protocol;
[0055] S63, Fault-Tolerant Mechanism Design: Design a fault-tolerant mechanism based on redundant backup to prevent data loss or inconsistency caused by node failure or network interruption between platforms;
[0056] S64, Data recovery and error handling mechanism: Design a data recovery mechanism to deal with data loss or errors in the event of system failure;
[0057] S65, Failover and automatic repair: Design an automatic failover mechanism to automatically switch to the backup node to continue data processing tasks when a platform node failure is detected.
[0058] Beneficial effects of the present invention:
[0059] The present invention can effectively eliminate the differences in data formats and protocols between different platforms and achieve seamless data integration and intercommunication by introducing a unified data structure standardization and data virtualization layer in the process of cross-platform data sharing and collaborative intelligent processing. By designing data mapping and conversion technology, it is ensured that the original data of each platform is highly consistent in storage format and access protocol, providing a solid foundation for subsequent data integration, virtualization and dynamic adaptation. This not only greatly improves the compatibility and scalability of the system, but also effectively reduces the complexity of data integration and improves the collaborative efficiency between platforms.
[0060] The present invention adopts an intelligent scheduling and resource allocation mechanism based on the reinforcement learning algorithm of Q-learning, which can dynamically adjust resource allocation according to the computing power, task type and data volume of the platform, thereby avoiding resource waste and processing bottlenecks. By analyzing the resource requirements of tasks and the resource status of each platform in real time, the intelligent scheduling algorithm can achieve efficient resource scheduling and task coordination between multiple platforms to ensure the maximization of the overall performance of the system. This solution enables the computing resources between platforms to be optimally configured, improves the efficiency and real-time performance of data processing, and avoids system delays and resource idleness caused by improper resource configuration.
[0061] The present invention designs a powerful data consistency control and fault-tolerant mechanism based on distributed transaction protocols (such as 2PC or 3PC) and distributed consistency algorithms (such as Paxos and Raft protocols). In a multi-platform environment, the consistency of cross-platform data is ensured, so that data can be updated synchronously in real time in a distributed system, and data recovery and error handling capabilities can be provided through a fault-tolerant mechanism even in the event of node failure or network interruption. This design ensures that the system can still operate stably in the face of partial failures or network fluctuations, greatly enhancing the reliability of data and the stability of the system, and is particularly suitable for cross-platform big data environments with high requirements for data consistency and reliability. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0063] Figure 1 A schematic diagram of a method flow of an embodiment of the present invention;
[0064] Figure 2 Schematic diagram of S2 process of an embodiment of the present invention. DETAILED DESCRIPTION
[0065] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. At the same time, it is explained here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments, and those skilled in the art may also adopt other alternatives to implement some known technologies; and the accompanying drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.
[0066] It should be noted that the references to "one embodiment", "an embodiment", "an exemplary embodiment", "some embodiments" and the like in the specification indicate that the embodiments described may include specific features, structures or characteristics, but not every embodiment may include the specific features, structures or characteristics. In addition, when a specific feature, structure or characteristic is described in conjunction with an embodiment, it should be within the knowledge of a person skilled in the art to implement such feature, structure or characteristic in conjunction with other embodiments (whether or not explicitly described).
[0067] In general, a term can be understood, at least in part, from its use in context. For example, depending, at least in part, on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in the singular sense, or can be used to describe a combination of features, structures, or characteristics in the plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but can instead, depending, at least in part, on the context, allow for the presence of other factors that are not necessarily explicitly described.
[0068] like Figure 1-Figure 2 As shown, the cross-platform big data sharing and collaborative intelligent processing method includes the following steps:
[0069] S1, cross-platform data structure standardization: Unify the structure of raw data from different platforms and use data mapping and conversion technology to ensure the consistency of raw data on each platform in storage format and access protocol, providing basic support for subsequent data integration and virtualization;
[0070] S2, data integration and virtualization: The data sources of different platforms are uniformly abstracted and virtualized through the data integration framework to form a cross-platform data virtualization layer. The virtualization layer dynamically adapts the data models of each platform to achieve data interoperability between platforms without relying on the technical details of the specific platform.
[0071] S3, Data sharing mechanism design based on access control model: Design a data sharing mechanism based on role-based access control or attribute-based access control model, formulate a secure access control strategy according to the permission requirements of different platforms, and ensure efficient data sharing and real-time data synchronization under the premise of meeting privacy protection and data security requirements;
[0072] S4, collaborative processing task scheduling and resource allocation: Using intelligent scheduling algorithms, dynamically adjust resource allocation based on the computing power, task type and data volume of each platform to ensure efficient collaborative execution of data processing tasks across multiple platforms, avoiding resource waste and processing bottlenecks;
[0073] S5, cross-platform collaborative analysis and automatic optimization: Combine big data mining and machine learning technologies to collaboratively analyze data from different platforms, automatically identify potential patterns and rules, optimize processing procedures, data sharing strategies and system performance, and enhance overall decision support capabilities;
[0074] S6, data consistency control and fault-tolerance mechanism: By designing a data consistency control mechanism and fault-tolerance mechanism based on a distributed transaction protocol, cross-platform data between platforms can always remain consistent in a distributed environment, and provide data recovery and error handling capabilities in the event of a system failure.
[0075] S1 includes:
[0076] S11, data source identification and classification: Identify data sources from different platforms, identify the types of data sources (such as relational databases, non-relational databases, file systems, etc.), and classify the data storage formats and access protocols (such as SQL, NoSQL, RESTful API, etc.) of each platform to form a metadata model of the data source, providing basic information for subsequent data standardization processing;
[0077] S12, data format analysis and attribute extraction: Detailed analysis of the data formats of each platform, including the data structure type (such as table type, document type, key-value pair type, etc.), and extraction of key attributes in the data (such as field name, data type, data range, etc.). During the analysis process, data mining techniques (such as cluster analysis) are used to identify the common attributes and differences in the data of each platform, and determine the target format for cross-platform standardization;
[0078] S13, data mapping rule formulation: According to the key attributes extracted in S12, formulate data mapping rules. The mapping rules include mapping relationships between data fields and attributes of different platforms to ensure that the data on different platforms can achieve consistency in format. The mapping rules are automatically generated by the rule engine. The specific rules are expressed as follows:
[0079] F target =M(Fsource );
[0080] Among them, F source is the source data field, M is the mapping rule function, and F target This rule ensures that the source data field F source Can be accurately converted into the target data field F according to the mapping rule M target ;
[0081] S14, data conversion and standardization: According to the data mapping rules formulated in S13, the original data from different platforms are structurally converted, and data conversion algorithms (such as data extraction, conversion, and loading technology in the ETL process) are used to convert data from the source format to the target standard format. The conversion process includes the following steps:
[0082] Data extraction: extracting raw data from data sources on each platform;
[0083] Data cleaning: Clean the extracted data to remove missing values, duplicate values and abnormal data to ensure data quality;
[0084] Data conversion: converting data from the original format to the standard format according to the mapping rules formulated in S13;
[0085] Data loading: Load the converted data into a unified standard data warehouse or database;
[0086] S15, data consistency verification: After the data conversion is completed, the converted data is verified for consistency to ensure that the data on different platforms are completely consistent in syntax, logic, and semantics. For example, a verification algorithm is used to compare the differences between the source data and the target data to ensure that the converted data is consistent with the original data in terms of content. The verification formula is as follows:
[0087]
[0088] Among them, D source and D target are the source data and target data respectively, δ represents the data consistency verification function, n is the total number of data fields, and the verification result is the consistency score;
[0089] S16. Standardized data output: The standardized data is output to a unified format and stored in the cross-platform data virtualization layer for subsequent data integration and virtualization processing. The output data format includes but is not limited to structured formats such as JSON, XML, and Parquet to ensure that cross-platform data can be efficiently shared and exchanged between different systems;
[0090] Through the above steps S11 to S16, the consistency of data storage formats and access protocols on different platforms can be ensured, providing a solid foundation support for subsequent data integration and virtualization, and improving data sharing and processing efficiency.
[0091] S2 includes:
[0092] S21, Data Integration Framework Design: Design and build a cross-platform data integration framework. The data integration framework receives data sources from multiple platforms and abstracts them uniformly. The framework is based on middleware technology and data access layer (such as ETL tools, API gateways, etc.) to ensure the stability and efficiency of data flow. The framework accesses the data sources of each platform through a unified interface and protocol (such as RESTAPI, SOAP, etc.) to avoid technical differences between platforms;
[0093] S22, construction of data virtualization layer: Based on the data integration framework, a cross-platform data virtualization layer is constructed. The virtualization layer abstracts and unifies the data models of different platforms and provides data access to the outside through virtual views. The virtualization layer can dynamically adapt to the data models of each platform, including table structure, document structure, and key-value pair structure, and dynamically generate a cross-platform unified data view through the metadata management system. The structure of the virtualization layer is as follows:
[0094] V=f(D 1 , D 2 , ..., D n );
[0095] Where V is the virtualized data view, D 1 , D 2 , ..., D n is the original data source from different platforms, f is the data abstraction function, which means integrating multiple data sources through a unified view;
[0096] S23, Data conversion and standardization: In the data virtualization layer, data conversion algorithms are applied to standardize data on different platforms. Specific methods include unified conversion of data types, formats, and units to ensure data consistency. The data conversion rules used in the data conversion process can be dynamically generated by the rule engine to cope with data differences between different platforms. After conversion, all data will conform to a unified standard format to facilitate cross-platform data interoperability;
[0097] S24, Dynamic Data Adaptation and Model Mapping: Through the data virtualization layer, dynamic adaptation of data models on various platforms is achieved. When data models on different platforms change, the virtualization layer can automatically identify changes in data structure and make adjustments. This process is completed through an automated model mapping algorithm, which maps the source data model to the target model and ensures the consistency of the data structure. The mapping rules are as follows:
[0098] M target =g(M source );
[0099] Among them, M source is the source data model, M target is the target data model, g is the mapping function, which represents the automatic conversion from the source model to the target model;
[0100] S25, cross-platform data intercommunication and data synchronization: With the support of the data virtualization layer, cross-platform data can be intercommunication and synchronized in real time. The data synchronization mechanism ensures that the data in the virtualization layer is updated in time between platforms and ensures the consistency of data between different platforms. The synchronization mechanism adopts an incremental synchronization algorithm and is carried out through the following steps:
[0101] D sync =Δ(D platform1 , D platform2 );
[0102] Among them, D sync is the synchronized data, Δ is the incremental synchronization function, which represents the synchronization process of data differences between platforms;
[0103] The implementation of steps S21 to S25 provides a unified framework and virtualization layer, which solves the problem of structural differences in cross-platform data sources through dynamic adaptation and automatic mapping technology. The data integration framework accesses the data sources of each platform through a unified interface, and the virtualization layer provides dynamic adaptation to different data models and ensures data consistency through automated mapping algorithms. Ultimately, cross-platform data can achieve efficient interoperability and real-time synchronization, thereby greatly improving the efficiency of data sharing and collaborative processing.
[0104] S3 includes:
[0105] S31, role definition and access rights division: define the role model for cross-platform data access, including data owners, data administrators, and data users, and divide the permissions of different roles based on business needs. Each role has a different permission set, which includes data reading, writing, modifying, and deleting.
[0106] S32, formulation of role and permission mapping rules: According to the data security requirements of different platforms, a set of specific permissions is assigned to each role. Through the mapping rules between roles and permissions, each role can access the data set consistent with its responsibilities. The mapping rules are implemented through the Access Control Matrix (ACM), where rows represent roles and columns represent data items. Each cell in the matrix represents the role's access rights to the data, expressed as:
[0107] ACM ij ={Read,Write,Modify,Delete};
[0108] Among them, ACM ij Indicates the access rights of role i to data item j. The permission values include read, write, modify, and delete.
[0109] S33, Dynamic allocation and control of access rights: According to the user's operation behavior and role changes in the cross-platform system, access rights are adjusted dynamically in real time. Through intelligent permission scheduling algorithms (such as dynamic adjustment algorithms based on access time and frequency), accurate permission allocation to user roles is achieved while ensuring privacy protection and data security.
[0110] S34, Data Sharing Strategy Development and Implementation: Based on the roles and permissions, design data sharing strategies to ensure that data is only shared between authorized roles to prevent unauthorized access. The strategy includes safeguards such as data encryption and data transmission security protocols (such as HTTPS, TLS) to ensure that data is always protected during transmission and storage;
[0111] A cross-platform data sharing mechanism is designed through the role-based access control model (RBAC), ensuring that each role can obtain corresponding permissions on different platforms, thereby achieving secure data sharing and real-time synchronization. The solution ensures the security and compliance of data access through the mapping rules between roles and permissions, dynamic permission adjustment, and the implementation of data sharing strategies.
[0112] S3 also includes:
[0113] S35, Attribute definition and classification: Define a multi-dimensional attribute model for users and data, including user attributes (such as department, rank, geographic location, etc.) and data attributes (such as data sensitivity, data source, data type, etc.). By describing users and data in an attributed manner, this provides a basis for access control.
[0114] S36, Attribute-based access control strategy design: Develop access control strategies based on user attributes and data attributes to ensure that data is only accessible to qualified users. The specific strategy is expressed as:
[0115] AC i ={A j |User attributes meet the conditions and data attributes meet the requirements};
[0116] Among them, AC i A represents the access control set of user i to data, j is a set of data attributes, and the access control conditions are determined by the matching of user attributes and data attributes;
[0117] S37, attribute evaluation and real-time authorization: Real-time evaluation of user and data attributes, dynamic matching based on user attribute values (such as rank, department, access frequency, etc.) and data attribute values (such as data sensitivity, data type, etc.), authorize users to access data, and dynamically adjust user access rights based on attribute weights and sensitivity through intelligent authorization algorithms;
[0118] S38, multi-level access control implementation: Implement multi-level access control strategies, including layer-by-layer authorization for data access (such as access control at the data level, field level, and record level). At the same time, set different security thresholds for different levels of control to improve the accuracy and security of data access;
[0119] S39, Data Sharing Strategy Implementation and Compliance Assurance: Ensure that the entire data sharing process complies with data privacy protection and compliance requirements, especially in a cross-platform environment, to achieve refined control over user and data attributes to avoid data leakage or unauthorized data access;
[0120] Another data sharing mechanism was designed through the attribute-based access control model (ABAC), which uses fine-grained user and data attribute matching for dynamic access control, achieving more flexible and accurate access management. This solution can ensure data security and privacy protection during cross-platform data sharing, and improve sharing efficiency and security.
[0121] S4 includes:
[0122] S41, task type and resource requirement analysis: perform type analysis on the data processing tasks on each platform, identify the computing requirements, data access requirements and execution time of each task, and evaluate its requirements for computing resources, storage resources and bandwidth according to the characteristics of the task (such as computing-intensive tasks, I / O-intensive tasks, mixed tasks, etc.). The resource requirements of each task are represented by a task descriptor, which includes information such as task type, data volume, required computing power, and estimated execution time;
[0123] S42, computing power evaluation and resource pool management: Evaluate the computing power of each platform, including the platform's CPU / GPU performance, memory size, and network bandwidth. Based on the evaluation results, aggregate the computing resources, storage resources, and bandwidth resources of each platform into a dynamic resource pool. The resource pool is dynamically updated based on the real-time status of the platform (such as load, response time, etc.) so that the intelligent scheduling algorithm can allocate tasks based on the latest resource conditions.
[0124] S43, Intelligent Scheduling Algorithm Design: Adopt intelligent scheduling algorithm to dynamically adjust the distribution of tasks among platforms according to the resource requirements of tasks, computing power of platforms and system load. Specifically, the scheduling algorithm is based on multi-objective optimization strategy to balance the computing load, network bandwidth and data storage requirements of tasks to ensure efficient execution of tasks. The objective function of the algorithm is:
[0125]
[0126] Among them, F is the scheduling optimization target, n is the number of platforms, C i is the computational requirement of task i, R i is the computing power of platform i, D i is the data volume of task i, B i is the bandwidth of platform i;
[0127] The optimization goal of this objective function is to minimize the sum of computational requirements and bandwidth consumption, ensuring that the task is executed on the platform with the most sufficient resources to improve resource utilization and avoid bottlenecks;
[0128] S44, dynamic task scheduling and load balancing: During the task execution process, the scheduling system dynamically adjusts the task allocation and performs load balancing based on real-time monitoring data (such as platform load, network delay, task progress, etc.). When a platform encounters resource overload or processing bottleneck, the system automatically migrates some tasks to other platforms with lower load to ensure balanced execution of tasks and avoid performance bottlenecks caused by overload of a single platform.
[0129] S45, resource release and recycling mechanism: When a task is completed, the occupied computing resources and storage resources are automatically released to ensure that the resource pool can provide sufficient resources for subsequent tasks. The system uses a resource recycling mechanism to ensure dynamic allocation and efficient utilization of resources between platforms;
[0130] Through the above steps S41 to S45, the coordinated processing of task scheduling and resource allocation ensures the efficient execution of cross-platform data processing tasks. Through task type and resource demand analysis, computing power evaluation and resource pool management, the intelligent scheduling algorithm can dynamically adjust task allocation and optimize resource utilization during task execution. Real-time monitoring and load balancing mechanisms further improve the processing efficiency of the system, avoid resource waste and processing bottlenecks, and thus improve the performance and responsiveness of the entire cross-platform data processing system.
[0131] S5 includes:
[0132] S51, Design of cross-platform data collaborative analysis framework: Design a cross-platform collaborative analysis framework. The collaborative analysis framework abstracts data from different platforms into a unified multidimensional data set. After data integration, the collaborative computing framework is used to achieve unified data analysis. Big data mining techniques, such as cluster analysis and association rule mining, are used to discover potential patterns and rules from heterogeneous data on different platforms. The collaborative analysis framework uses a distributed computing architecture (such as Hadoop, Spark, etc.) for parallel processing to ensure that the analysis process is efficient and scalable.
[0133] S52, Data pattern and rule mining: Based on data mining algorithms, analyze data from different platforms to automatically identify potential patterns, associations, and trend changes in the data. For example, use clustering algorithms (such as K-means, DBSCAN, etc.) to group data, or use association rule mining algorithms (such as Apriori, FP-growth, etc.) to discover the associations between data from different platforms;
[0134] S53, pattern recognition and data fusion: Through machine learning algorithms (such as deep learning, reinforcement learning, etc.), the analysis results are further processed to fuse the global model from the heterogeneous data of different platforms. This model can perform real-time pattern recognition and data fusion according to the data changes of different platforms, thereby providing a basis for subsequent optimization and decision-making. The fused data model is expressed as:
[0135] M global =f(M 1 , M 2 , ..., M n );
[0136] Among them, M 1 , M 2, ..., M n Indicates the data mode of different platforms, M global is the data mode after global fusion, and f is the fusion function;
[0137] S54, Feedback mechanism for cross-platform collaborative analysis results: Establish a cross-platform analysis result feedback mechanism to feed back collaborative analysis results to each platform to achieve data-driven decision support. This feedback mechanism uses real-time data streams (such as Kafka, Apache Pulsar, etc.) to ensure that analysis results can be fed back in a timely manner and promote cross-platform optimization;
[0138] The collaborative analysis process from step S51 to step S54 can effectively integrate data from different platforms and identify potential patterns and rules in the data through big data mining and machine learning algorithms. This analysis framework can not only improve data utilization efficiency, but also provide a reliable basis for subsequent automatic optimization.
[0139] The S5 also includes:
[0140] S55, automatic optimization objective function design: Based on the results of cross-platform collaborative analysis, the automatic optimization objective function is designed. The objective function comprehensively considers multiple factors such as processing flow, data sharing strategy and system performance, aiming to maximize resource utilization efficiency and reduce system latency and processing bottlenecks. The objective function is expressed as:
[0141] Objective=α·Efficiency-β·Latency+γ·Cost;
[0142] Among them, α, β, γ are weighted coefficients, Efficiency represents the efficiency of the processing flow, Latency represents the delay of the system, and Cost represents the system operation cost;
[0143] S56, Optimization Algorithm Design and Implementation: Based on the Q-learning reinforcement learning algorithm, the processing flow, data sharing strategy and system performance are automatically optimized according to the objective function, including:
[0144] S561, Q-learning algorithm framework construction: First, build the framework of the Q-learning algorithm, which contains the following main elements:
[0145] State space: defined as the current state of the system, including real-time information such as platform load, resource utilization, task queue length, data transmission volume, etc.
[0146] Action space: defined as the optimization actions that can be performed, such as adjusting task scheduling strategies, changing data sharing strategies, choosing different data processing paths, etc.
[0147] Reward function: Set up a reward mechanism and give corresponding rewards based on system performance and resource utilization efficiency. For example, if the adjustment of task scheduling reduces waiting time and data transmission delay, a positive reward will be given; if the system wastes resources or tasks are blocked, a negative reward will be given.
[0148] Through the Q-learning algorithm, the intelligent agent evaluates the value of the action according to the reward function after each action is performed, thereby dynamically learning the optimal action strategy, which is expressed as:
[0149]
[0150] Where Q(s,a) is the value of executing action a in state s, R(s,a) is the immediate reward after executing action a in state s, γ is the discount factor used to evaluate the impact of future rewards, and α is the learning rate, which controls the update ratio of new experience to old value.
[0151] S562, real-time update of state and action: Based on real-time system feedback, the Q-learning algorithm dynamically adjusts the state and action of the system. For example, when the system's resource utilization efficiency is low, Q-learning will select the optimal action through exploration and exploitation strategies, thereby gradually improving the efficiency of task scheduling and data sharing strategies;
[0152] S563, Dynamically optimize the processing flow: The Q-learning algorithm is used to dynamically optimize the processing flow, including task allocation, data transmission scheduling, and processing order. For example, when multiple platforms are co-processing, if the computing load of a platform is too high, the Q-learning algorithm will automatically choose to transfer some tasks to other platforms with lighter loads to avoid computing bottlenecks and resource waste;
[0153] S564, Data Sharing Strategy Adjustment: The Q-learning algorithm dynamically adjusts the data sharing strategy based on real-time monitoring of platform load, task execution, network bandwidth and other information. By continuously evaluating and optimizing data flow paths and data transmission methods (such as choosing direct transmission or cache transfer), the algorithm can effectively reduce data transmission delays and improve overall system performance;
[0154] S565, system performance improvement and real-time feedback: Through the iterative learning process, the Q-learning algorithm can continuously adjust system parameters in actual operation and optimize the system's operating efficiency. For example, after multiple optimizations, the system can autonomously identify the best task scheduling solution, data sharing strategy, and optimal computing resource allocation method, thereby achieving global performance improvement;
[0155] S566, optimization result evaluation and verification: After each optimization process, the optimization results are evaluated by monitoring key performance indicators such as the system's response time, processing efficiency, and resource utilization. If the optimization effect does not meet expectations, the system will further adjust the learning strategy or explore new optimization solutions to ensure the effectiveness and sustainability of the optimization process;
[0156] S57, Feedback mechanism and dynamic adjustment: By real-time monitoring of the system's operating status, the resource allocation strategy is automatically adjusted according to the optimization results. During the operation of the optimization algorithm, the parameters can be dynamically adjusted according to the system feedback, so that the optimization strategy can be adjusted in real time as the platform status changes. The optimization results are transmitted to each platform through the feedback mechanism to ensure that cross-platform data processing is always in the best state;
[0157] S58, optimization effect evaluation and verification: Evaluate the effect of the optimized system, analyze the resource utilization efficiency, task execution time and data sharing efficiency after optimization, and verify the effectiveness of the optimization algorithm by comparing with the system before optimization. The evaluation results will be used to further adjust the optimization strategy and improve the overall performance of the system;
[0158] By setting a reasonable objective function and combining it with a machine learning optimization algorithm, automatic optimization of cross-platform data processing processes, data sharing strategies, and system performance is achieved. This optimization process not only improves resource utilization, but also effectively reduces system latency and operating costs, further enhancing the system's decision-making support capabilities.
[0159] S6 includes:
[0160] S61, Distributed transaction protocol design: Design a data consistency control mechanism based on the distributed transaction protocol to ensure data consistency between multiple platforms. This mechanism is based on the two-phase commit protocol (2PC) or the three-phase commit protocol (3PC). By coordinating the transaction operations of each platform, it ensures that all platforms can reach a consistent transaction state when updating data to avoid data inconsistency. The protocol transmits confirmation information of transaction operations between platforms through message transmission to ensure the atomicity of transactions in a distributed environment;
[0161] S62, data consistency protocol implementation: Use a distributed consistency algorithm based on the Paxos protocol or the Raft protocol to implement a distributed consistency protocol to ensure that data copies between multiple platforms remain consistent during synchronization, and that the final consistency of data can be guaranteed even in the event of network partitions or node failures.
[0162] The specific algorithm is:
[0163] Paxos protocol: ensures that in a distributed environment, even if some platforms fail, a consistent leader node can still be elected and the node coordinates data updates;
[0164] Raft protocol: Synchronizes data updates between the primary node and the backup node through the log replication mechanism to ensure data consistency;
[0165] S63, Fault-tolerant mechanism design: Design a fault-tolerant mechanism based on redundant backup to prevent data loss or inconsistency caused by node failure or network interruption between platforms. The mechanism includes regular data backup and real-time data replication. Every time the data is updated, the data update is copied to multiple backup nodes in real time through a distributed storage system (such as HDFS, Ceph, etc.), and the data of the backup nodes is ensured to be consistent with the primary node. In the event of a node failure, the system can quickly switch to the backup node to ensure the continuity and integrity of the data processing task;
[0166] S64, Data recovery and error handling mechanism: Design a data recovery mechanism to deal with data loss or errors in the event of system failure, regularly log data operations, and implement data recovery through log playback technology (such as the WAL protocol). Error detection and automatic repair mechanisms ensure that when data errors or inconsistencies occur, they can be repaired through the consistency protocol. For example, when a transaction fails, the system will roll back to a consistent state and recover data through compensating transactions to ensure high availability and data integrity of the system;
[0167] S65, Failover and automatic repair: Design an automatic failover mechanism. When a platform node failure is detected, it automatically switches to the backup node to continue the data processing task, while ensuring data consistency and availability. The fault-tolerant mechanism uses heartbeat detection and timeout mechanisms to detect the availability of nodes in a timely manner, and automatically repair or reschedule tasks when a failure occurs.
[0168] By designing a data consistency control mechanism and fault tolerance mechanism based on a distributed transaction protocol, data consistency and reliability are achieved between multiple platforms, ensuring high availability and consistency of cross-platform data in a distributed environment. At the same time, redundant backup, data recovery, and automatic failover mechanisms are adopted to improve the system's fault tolerance, ensure reliable data recovery in the event of a failure, avoid data loss or inconsistency, and improve the stability and robustness of the overall system.
[0169] The present invention covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present invention. In order to make the public have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, but those skilled in the art can fully understand the present invention without the description of these details. In addition, in order to avoid unnecessary confusion about the essence of the present invention, well-known methods, processes, procedures, components and circuits are not described in detail.
[0170] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A cross-platform big data sharing and collaborative intelligent processing method, characterized in that: The following steps are involved: S1, cross-platform data structure standardization: perform unified structural standardization on raw data from different platforms, and use data mapping and conversion technology to ensure consistency in storage format and access protocol of raw data on each platform; S2, data integration and virtualization: The data sources of different platforms are uniformly abstracted and virtualized through the data integration framework to form a cross-platform data virtualization layer. The virtualization layer dynamically adapts the data models of each platform to achieve data interoperability between platforms without relying on the technical details of the specific platform. S3, Data sharing mechanism design based on access control model: Design a data sharing mechanism based on role-based access control or attribute-based access control model, and formulate a secure access control strategy according to the permission requirements of different platforms; S4, collaborative processing task scheduling and resource allocation: using intelligent scheduling algorithms, dynamically adjust resource allocation based on the computing power, task type and data volume of each platform to avoid resource waste and processing bottlenecks; S5, cross-platform collaborative analysis and automatic optimization: Combine big data mining and machine learning technologies to collaboratively analyze data from different platforms, automatically identify potential patterns and rules, and optimize processing procedures, data sharing strategies, and system performance; S6, data consistency control and fault-tolerance mechanism: By designing a data consistency control mechanism and fault-tolerance mechanism based on a distributed transaction protocol, cross-platform data between platforms can always remain consistent in a distributed environment, and provide data recovery and error handling capabilities in the event of a system failure.
2. The cross-platform big data sharing and collaborative intelligent processing method according to claim 1 is characterized in that: The S1 includes: S11, data source identification and classification: Identify data sources from different platforms, identify the types of data sources, and classify the data storage formats and access protocols of each platform to form a metadata model of the data source; S12, data format analysis and attribute extraction: Analyze the data format of each platform, including the structure type of the data, and extract the key attributes of the data; S13, data mapping rule formulation: formulate data mapping rules based on the key attributes extracted in S12; S14, data conversion and standardization processing: According to the data mapping rules formulated in S13, the raw data from different platforms are structurally converted, and the data is converted from the source format to the target standard format using the data conversion algorithm; S15, data consistency verification: after the data conversion is completed, the converted data is verified for consistency; S16. Standardized data output: Output the standardized data into a unified format.
3. The cross-platform big data sharing and collaborative intelligent processing method according to claim 2 is characterized in that: The S2 includes: S21, Data Integration Framework Design: Design and build a cross-platform data integration framework that receives data sources from multiple platforms and abstracts them uniformly; S22, construction of data virtualization layer: based on the data integration framework, build a cross-platform data virtualization layer; S23, data conversion and standardization: In the data virtualization layer, data conversion algorithms are applied to standardize data on different platforms; S24, dynamic data adaptation and model mapping: through the data virtualization layer, dynamic adaptation of data models on each platform is achieved; S25, cross-platform data interoperability and data synchronization: With the support of the data virtualization layer, cross-platform data can be interoperable and synchronized in real time. Through the data synchronization mechanism, the data in the virtualization layer can be updated in time between platforms.
4. The cross-platform big data sharing and collaborative intelligent processing method according to claim 3 is characterized in that: The S3 includes: S31, role definition and access rights division: define the role model for cross-platform data access, including data owners, data administrators, and data users, and divide the permissions of different roles based on business needs. Each role has a different permission set, which includes data reading, writing, modifying, and deleting. S32, formulation of role and permission mapping rules: According to the data security requirements of different platforms, a set of permissions is assigned to each role. Through the mapping rules between roles and permissions, each role can access the data set that matches its responsibilities; S33, Dynamic allocation and control of access rights: Dynamically adjust access rights in real time based on user operations and role changes in the cross-platform system; S34, Data sharing strategy formulation and implementation: Design data sharing strategies based on the roles and permissions.
5. The cross-platform big data sharing and collaborative intelligent processing method according to claim 4 is characterized in that: The S3 further includes: S35, Attribute definition and classification: Define the multi-dimensional attribute model of users and data, including user attributes and data attributes, and provide a basis for access control by describing users and data in an attributed manner; S36, Attribute-based access control strategy design: Develop access control strategies based on user attributes and data attributes to ensure that data is only accessible to qualified users; S37, attribute evaluation and real-time authorization: Evaluate the attributes of users and data in real time, dynamically match the attribute values of users with the attribute values of data, and authorize users to access data; S38, Multi-level access control implementation: Implement multi-level access control strategies, including layer-by-layer authorization for data access, and set different security thresholds for different levels of control; S39, data sharing strategy implementation and compliance assurance: Ensure that the entire data sharing process complies with data privacy protection and compliance requirements.
6. The cross-platform big data sharing and collaborative intelligent processing method according to claim 5 is characterized in that: The S4 includes: S41, Task type and resource requirement analysis: Analyze the data processing tasks on each platform, identify the computing requirements, data access requirements and execution time of each task, and evaluate the requirements for computing resources, storage resources and bandwidth according to the characteristics of the task; S42, computing capability evaluation and resource pool management: Evaluate the computing capability of each platform, including the platform's CPU / GPU performance, memory size, and network bandwidth. Based on the evaluation results, aggregate the computing resources, storage resources, and bandwidth resources of each platform into a dynamic resource pool, which is dynamically updated based on the real-time status of the platform. S43, Intelligent Scheduling Algorithm Design: Adopt intelligent scheduling algorithm to dynamically adjust the distribution of tasks among platforms according to the resource requirements of tasks, computing power of platforms and system load; S44, dynamic task scheduling and load balancing: During the task execution process, the scheduling system dynamically adjusts the task allocation according to the real-time monitoring data and performs load balancing; S45, resource release and recycling mechanism: when the task is completed, the occupied computing resources and storage resources are automatically released.
7. The cross-platform big data sharing and collaborative intelligent processing method according to claim 6 is characterized in that: The S5 includes: S51, Design of cross-platform data collaborative analysis framework: Design a cross-platform collaborative analysis framework that abstracts data from different platforms into a multidimensional data set; S52, Data Pattern and Rule Mining: Based on data mining algorithms, data from different platforms are analyzed to automatically identify potential patterns, correlations and trend changes in the data; S53, Pattern Recognition and Data Fusion: The analysis results are further processed through machine learning algorithms to fuse global models from heterogeneous data on different platforms; S54, feedback mechanism for cross-platform collaborative analysis results: Establish a cross-platform analysis result feedback mechanism to feed back the collaborative analysis results to each platform.
8. The cross-platform big data sharing and collaborative intelligent processing method according to claim 7 is characterized in that: The S5 further includes: S55, automatic optimization objective function design: design the automatic optimization objective function based on the results of cross-platform collaborative analysis; S56, optimization algorithm design and implementation: Based on Q-learning reinforcement learning algorithm, the processing flow, data sharing strategy and system performance are automatically optimized according to the objective function; S57, Feedback mechanism and dynamic adjustment: By real-time monitoring of the system's operating status, the resource allocation strategy is automatically adjusted according to the optimization results; S58, optimization effect evaluation and verification: Evaluate the effect of the optimized system and analyze the optimized resource utilization efficiency, task execution time and data sharing efficiency.
9. The cross-platform big data sharing and collaborative intelligent processing method according to claim 8 is characterized in that: The S6 includes: S61, Distributed Transaction Protocol Design: Design a data consistency control mechanism based on the distributed transaction protocol to ensure data consistency across multiple platforms; S62, data consistency protocol implementation: adopt a distributed consistency algorithm based on the Paxos protocol or the Raft protocol to implement a distributed consistency protocol; S63, Fault-Tolerant Mechanism Design: Design a fault-tolerant mechanism based on redundant backup to prevent data loss or inconsistency caused by node failure or network interruption between platforms; S64, Data recovery and error handling mechanism: Design a data recovery mechanism to deal with data loss or errors in the event of system failure; S65, Failover and automatic repair: Design an automatic failover mechanism to automatically switch to the backup node to continue data processing tasks when a platform node failure is detected.