A heterogeneous storage method, system, device and medium for data

By using the Apache Flink engine and built-in policy adapters to achieve heterogeneous storage of financial business data, this technology solves the problems of complex storage procedures and high difficulty in customized development in existing technologies, and realizes standardized, process-oriented and automated data storage.

CN118296066BActive Publication Date: 2025-09-09BEIJING HAIKE RONGTONG PAYMENT SERVICE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410357738.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-27
Publication Date
2025-09-09
Estimated Expiration
2044-03-27

AI Technical Summary

Technical Problem

Existing technologies for storing heterogeneous financial business data are complex, difficult to customize, and inflexible in supporting business changes.

Method used

By acquiring the target policy task and the data stream to be stored, and using the Apache Flink engine and built-in policy adapter, the data stream is heterogeneously stored into the corresponding heterogeneous data storage according to the context policy logic, thus realizing a standardized, streamlined, and automated storage process.

Benefits of technology

It reduces redundant development work, simplifies heterogeneous storage processes, and improves adaptability to business needs and data storage security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118296066B_ABST
    Figure CN118296066B_ABST
Patent Text Reader

Abstract

This application relates to a method, system, device, and medium for heterogeneous data storage. The method includes: obtaining a target policy task and a data stream to be stored; determining a target policy adapter based on the target policy task, wherein the target policy adapter includes contextual policy logic; and heterogeneously storing the data stream to be stored based on the contextual policy logic. The heterogeneous storage method stores different types of data streams to be stored in corresponding adapted heterogeneous data storage devices. This method solves the problem in existing technologies of complex business processes and the difficulty of customized development for processing and storing large amounts of business data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to a method, system, device and medium for heterogeneous storage of data. Background Art

[0002] With the advent of the big data era, data collection in financial business systems faces enormous challenges. In particular, as data regulation increasingly demands higher standards for data quality and security, financial enterprises are continuously strengthening their data asset management efforts. Consequently, in real-time data synchronization scenarios for financial businesses, the process of exporting data to heterogeneous storage requires encryption and other intermediate operations in accordance with the financial industry's level protection requirements.

[0003] However, in a large number of complex business scenarios, it is necessary to embed matching logic and other requirements according to business needs. Although there are a large number of data synchronization tools that can simplify the complexity of the process, business environments with a large number of customized requirements still face problems such as complex processes, repeated development, and inflexible support for business changes. At the same time, the data construction process on different heterogeneous storage terminals inevitably needs to meet the corresponding storage requirements, which also increases the complexity of the large amount of business adaptation and development work required in the process of realizing business data output. Therefore, the business process for processing and storing large amounts of business data in the existing technology is relatively complex, and customized development is relatively difficult. Summary of the Invention

[0004] In order to overcome the problems in the prior art of complex business processes for processing and storing large amounts of business data and the difficulty of customized development, the present application provides a heterogeneous storage method, system, device and medium for data.

[0005] In a first aspect, in order to solve the above technical problems, the present application provides a method for heterogeneous storage of data, comprising:

[0006] Obtain the target policy tasks and data streams to be stored;

[0007] According to the target policy task, a target policy adapter is determined, and the target policy adapter includes context policy logic;

[0008] According to the contextual policy logic, the data stream to be stored is stored heterogeneously; wherein the heterogeneous storage represents storing different types of data streams to be stored in corresponding adapted heterogeneous data storage devices.

[0009] In a second aspect, the present application further provides a heterogeneous storage system for data, comprising:

[0010] The acquisition module is used to obtain the target policy tasks and the data stream to be stored;

[0011] A determination module, configured to determine a target policy adapter according to a target policy task, wherein the target policy adapter includes a contextual policy logic;

[0012] The storage module is used to perform heterogeneous storage on the data stream to be stored according to the context policy logic; wherein the heterogeneous storage represents storing different types of data streams to be stored in corresponding adapted heterogeneous data storage devices.

[0013] In a third aspect, the present application also provides a computing device, including a memory, a processor, and a program stored in the memory and running on the processor. When the processor executes the program, the steps of a heterogeneous storage method for data as described above are implemented.

[0014] In a fourth aspect, the present application also provides a computer-readable storage medium, in which instructions are stored. When the instructions are executed on a terminal device, the terminal device executes the steps of a heterogeneous storage method for data.

[0015] The beneficial effects of this application are: by obtaining pre-set target policy tasks, it is possible to standardize storage tasks in advance according to actual business needs, which can reduce the repetitive development of policy tasks. The heterogeneous storage process for the data stream to be stored is formulated as follows: determining the target policy task according to the target policy task, and storing the data stream to be stored according to the context policy logic in the target policy task, which can standardize, streamline and automate the heterogeneous storage of the data stream to be stored, making the heterogeneous storage of the data stream to be stored simple, while reducing the repetitive development of the heterogeneous storage process of the data stream to be stored, thereby reducing the difficulty of customized development of the heterogeneous storage process for the data stream to be stored. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] Figure 1 A flowchart of a heterogeneous data storage method for this application;

[0017] Figure 2 This is another flowchart of a heterogeneous data storage method according to the present application;

[0018] Figure 3 for Figure 2 Specific examples of the illustrated embodiments;

[0019] Figure 4 This is an architectural diagram of a heterogeneous data storage method for this application;

[0020] Figure 5 This is a schematic diagram of the structure of a heterogeneous storage system for data in this application. DETAILED DESCRIPTION

[0021] The following examples are provided to further explain and supplement the present application and do not constitute any limitation to the present application.

[0022] The following describes a heterogeneous data storage method, system, device, and medium according to an embodiment of the present application with reference to the accompanying drawings.

[0023] An embodiment of the present application provides a heterogeneous storage method for data, which is applied to a terminal device. In the present application scheme, the terminal device is used as the execution subject to illustrate the present application scheme, and the terminal device is used to execute the steps of a heterogeneous storage method for data.

[0024] like Figure 1 As shown, the present application provides a heterogeneous storage method for data, including:

[0025] Step S1, obtaining the target policy task and the data flow to be stored;

[0026] Step S2: determining a target policy adapter according to the target policy task, where the target policy adapter includes contextual policy logic;

[0027] Step S3: heterogeneously storing the data stream to be stored according to the contextual policy logic; wherein the heterogeneous storage means storing different types of data streams to be stored in corresponding adapted heterogeneous data storage devices.

[0028] The heterogeneous storage method for data of this embodiment obtains pre-set target policy tasks, so that storage tasks can be standardized in advance according to actual business needs, which can reduce the repeated development of policy tasks. The heterogeneous storage process for the data stream to be stored is formulated as follows: determining the target policy task according to the target policy task, and storing the data stream to be stored according to the context policy logic in the target policy task. This can standardize, streamline and automate the heterogeneous storage of the data stream to be stored, making the heterogeneous storage of the data stream to be stored simple, while reducing the repeated development of the heterogeneous storage process of the data stream to be stored, thereby reducing the difficulty of customized development of the heterogeneous storage process for the data stream to be stored.

[0029] The method of this embodiment is implemented using the Apache Flink engine. Apache Flink is a framework and distributed processing engine for performing stateful computations on both unbounded and bounded data streams. The Flink engine runs in all common cluster environments and can perform computations at high memory speeds and at any scale.

[0030] In some embodiments, as Figure 2As shown, after obtaining the policy task (target policy task), all current data streams that need to be heterogeneously stored are merged, that is, N streams are merged into one data stream (data stream to be stored). According to the contextual policy logic in the policy task, the data stream to be stored is subjected to heterogeneous storage operations, and the output is stored in the corresponding adapted heterogeneous data storage device to realize heterogeneous storage of the data stream to be stored.

[0031] Optionally, the context policy logic includes data storage requirements. According to the context policy logic, the data stream to be stored is heterogeneously stored, including:

[0032] Determine the type of data stream to be stored according to data storage requirements. The data storage requirements represent the requirements for storing different types of data streams to be stored in corresponding adapted heterogeneous data storage devices.

[0033] According to the type of data stream to be stored, the corresponding target heterogeneous data storage is determined from the preset storage medium library;

[0034] According to the data storage requirements, the data stream to be stored is stored in the target heterogeneous data storage.

[0035] In this embodiment, since the contextual policy logic includes a data storage requirement for storing different types of to-be-stored data streams in corresponding, adapted heterogeneous data storage devices, the data storage requirement includes the data stream type of the to-be-stored data stream. This data storage requirement represents the heterogeneous data storage device's storage requirements for storable data streams. Thus, based on the data stream type in the data storage requirement, the corresponding target heterogeneous data storage device can be directly determined from a pre-set storage media library. The to-be-stored data stream can then be stored in the corresponding, adapted target heterogeneous data storage device in accordance with the data storage requirement, thereby achieving standardized, streamlined, and automated heterogeneous storage of the to-be-stored data streams.

[0036] Utilizing the real-time synchronization module in the Apache Flink engine, the data stream to be stored is written to the target heterogeneous data storage in real time according to the data storage requirements configured in the policy context. Figure 3 for Figure 2 The specific example of the embodiment shown shows the business process description corresponding to the target strategy task in the real-time synchronization module in the Apache Flink engine. Figure 3As shown, the scenario for policy configuration task 1 (target policy task 1) is configured. Policy configuration task 1 is: merge N data streams into one data stream to be stored and store it in the target heterogeneous data storage "Mongo database". The data stream type of the data stream in the N data streams is a fact table. Merge the order table stream and the order details table stream into the order details table (merged table) stream. Since the data stream type of the order table stream and the order details table stream is both a fact table, according to policy configuration task 1, the order details table (merged table) stream is stored in the Mongo database. The scenario for policy configuration task 2 (target policy task 2) is configured. Policy configuration task 2 is: merge N data streams into one data stream to be stored and store it in the target heterogeneous data storage "HBase database". The data stream types of the data streams in the N data streams include fact tables and dimension tables. The order table stream, order details table stream, agent table stream, and region table stream are merged into an order summary table (merged table) stream. Since the data stream types of the order table stream and order details table stream are both fact tables, and the data stream types of the agent table stream and region table stream are both dimension tables, according to the policy configuration task 2, the order summary table (merged table) stream is stored in the HBase database. In this way, based on the data stream type in the data storage requirement, the corresponding target heterogeneous data storage can be directly determined from the preset storage media library. In accordance with the data storage requirement, the data stream to be stored is stored in the corresponding adapted target heterogeneous data storage, thereby achieving standardization, process-based, and automated heterogeneous storage of the data stream to be stored.

[0037] Optionally, obtain the target policy task, including:

[0038] Generate policy templates using the preset built-in policy adapter;

[0039] Formulate policy scheduling corresponding to policy templates. Policy scheduling includes policy rules and scheduling resources.

[0040] Generate corresponding policy tasks to be executed according to policy rules;

[0041] Verify policy rules and scheduling resources to obtain verification results;

[0042] When the verification result indicates that the verification has passed, the policy task to be executed is used as the target policy task.

[0043] In this embodiment, a policy template is generated through a preset built-in policy adapter. This policy template is universal and can template the storage requirements of data streams that require heterogeneous storage, thereby reducing repetitive development work. A policy schedule corresponding to the policy template is formulated, and a policy task to be executed corresponding to the policy schedule is generated, so that the policy task to be executed can meet the requirements for heterogeneous storage of the data stream to be stored. At the same time, the policy task to be executed is executed as the target policy task only when the verification results corresponding to the policy rules and scheduling resources in the policy schedule are verified to pass. This can improve the security of the target policy task, thereby improving the security of the heterogeneous storage of the data stream to be stored.

[0044] In this embodiment, the built-in policy adapter is located in the policy center, which manages policies and uses the built-in policy adapter to generate policy templates for heterogeneous storage of data streams to be stored. This allows for template-based storage requirements for data streams requiring heterogeneous storage, thereby generating the target policy tasks required for heterogeneous storage of data streams to be stored. The policy rules in the policy templates generated by the built-in policy adapter's mapping primarily include rule structures such as a data source end, a data storage end, a data source logical segment, a data storage logical segment, and a feature definition segment.

[0045] Due to the varying scenarios and types of data streams to be stored, the business requirements for heterogeneous storage of these data streams also vary. Therefore, a policy template must be developed based on the business requirements of the data streams to be stored. This policy template defines the Kafka source (consumer) configuration rule section and data storage section within the policy schedule. These sections are known as policy rules. These sections include producer, consumer, and server configuration rule sections. Furthermore, resource configuration requests must be submitted based on the resource queues of the allocated Apache Flink engine's real-time cluster to configure the scheduling resources within the policy schedule. These resource configuration requests include memory, core count, and parallelism. In policy scheduling, both policy rule configuration and scheduling resource configuration are automatically generated by the system based on the data source and target storage selections. The default scheduling resource configuration is automatically generated by the system based on the resource availability of all data streams. Customization of the default scheduling resource configuration is required only for specific scenarios. Policy scheduling supports scheduling policies in the cluster environment of the Apache Flink engine, including YARN's FIFO scheduling, capacity scheduling, and fair scheduling.

[0046] Optionally, verify policy rules and scheduling resources, including:

[0047] Determine whether the rule constraints corresponding to the policy rules and the corresponding built-in configurators are legal;

[0048] Determine whether the scheduled resources fall within the preset resource range;

[0049] Determine whether the scheduling resources are legal.

[0050] In this embodiment, since the policy rules are configured and output by the corresponding built-in configurator, and the policy rules need to be constrained by the cluster environment when executed, the policy rules can only be safely executed in the cluster environment when the built-in configurator corresponding to the policy rules is legal, that is, the constraint requirements of the built-in configurator corresponding to the policy rules meet the configuration constraint requirements of the cluster environment, and the rule constraints corresponding to the policy rules are legal, that is, the rule constraints corresponding to the policy rules meet the rule constraints of the cluster environment, thereby improving the legitimacy of the policy rules. At the same time, the preset resource scope is the resource situation of all user data that can exist in the cluster environment of the Apache Flink engine. Only when the resources in the scheduling resources fall within the resource scope of the data in the resource situation and the scheduling resources are legal resources can the corresponding data source resources to be stored be scheduled and stored according to the scheduling resources, thereby improving the legitimacy of the scheduling resources. In this way, verifying the policy rules and scheduling resources can improve the legitimacy of policy scheduling.

[0051] Optionally, when the verification result indicates a verification failure, a new scheduling policy corresponding to the policy template is formulated, where the new policy scheduling includes new policy rules and new scheduling resources;

[0052] The new policy rules and the new scheduling resources are verified, and the verification result is determined until the verification result indicates that the verification has passed.

[0053] In this embodiment, when the verification result indicates that the verification has failed, it means that the current scheduling policy is illegal. If it is enforced, the security of the data stream to be stored during heterogeneous storage will be reduced. Therefore, it is necessary to re-formulate a new scheduling policy corresponding to the policy template and continue to verify the new scheduling resources until the verification passes. Only then can the scheduling policy corresponding to the verification be used to generate and execute the corresponding policy task, thereby improving the security of the data stream to be stored during heterogeneous storage.

[0054] Optionally, the data stream to be stored includes at least one piece of changed data, and obtaining the data stream to be stored includes:

[0055] Collect multiple user data in real time;

[0056] Perform a search operation on each piece of user data in a preset database to determine a matching result corresponding to each piece of user data;

[0057] For each piece of user data, if the matching result indicates that no matching data is found, the user data is treated as the changed data;

[0058] All changed data is treated as the data stream to be stored.

[0059] In this embodiment, since some of the user data collected in real time may have been stored before, only the data in the user data collected in real time that is not stored in the preset database of historical collection is treated as changed data, and all the changed data is stored heterogeneously as the data stream to be stored. This can reduce the invalid storage of duplicate data in the user data collected in real time and speed up the storage speed of valid data in the user data collected in real time.

[0060] Optionally, according to the target policy task, a target policy adapter is determined, including:

[0061] Analyze the target strategy task and obtain the analysis result;

[0062] The parsing results are mapped and matched in the preset policy adapter library to determine the corresponding target policy adapter.

[0063] In this embodiment, the target policy task includes execution adaptation requirements for heterogeneous storage of the data stream to be stored. After parsing the target policy task, the parsing result obtained can understand the corresponding execution adaptation requirements, so that the target policy adapter with the same execution adaptation requirements can be matched from the preset policy adapter library according to the execution adaptation requirements. The context policy logic in the target policy adapter includes the execution steps for heterogeneous storage of the data stream to be stored in the target policy task, so that the data stream to be stored can be heterogeneously stored according to the context policy logic in the target policy adapter, thereby making the heterogeneous storage of the data stream to be stored standardized, streamlined and automated.

[0064] In some embodiments, the target policy task submits the policy schedule to the Apache Flink cluster via a gateway interface. Apache Flink's real-time scheduling task module then performs management operations such as startup, offline, editing, and deletion based on the target policy task. If an exception occurs during the startup, submission, or running of a scheduled task startup or online task, the scheduler's built-in retry mechanism will initiate a retry. Within the retry policy scope, an exception alert will be issued, indicating that the abnormal task has successfully retried. If the retry policy scope is outside the retry policy scope, an active alert mechanism will be initiated to notify the configured alert target message channel.

[0065] The preset policy adapter library is stored in the cluster of the Apache Flink engine. The preset policy adapter library contains multiple preset policy adapters, as shown in Table 1.

[0066] Table 1

[0067]

[0068] Preset policy adapters include data storage adapters, type adapters, security adapters, and data format adapters. Data storage adapters primarily provide adaptation functions for heterogeneous data stores, including but not limited to MongoDB, HBase, ElasticSearch (ES), MySQL, and Oracle databases. These data storage adapters include MongoDB adapters, HBase adapters, ES adapters, and MySQL adapters. Type adapters primarily implement adaptation functions such as type-related mapping and conversion between heterogeneous data sources and data storage sources. These type adapters include MySQL-Mongo adapters, MySQL-HBase adapters, MySQL-ES adapters, and Oracle-Mongo adapters. Security adapters primarily implement adaptation functions for general security, business-defined security, and other security requirements. These security adapters include encryption adapters, policy validity adapters, blacklist and whitelist adapters, and custom security adapters. Data format adapters primarily implement adaptation functions for time and specific format validation. These data format adapters include time format adapters, JSON format adapters, business-specific format adapters, and custom format adapters. The data related to the source end is obtained through the message queue, and the real-time data synchronization task logic processing is performed on the data stream according to the context policy logic of the policy adapter, that is, the storage data stream is heterogeneously stored according to the context policy logic of the policy adapter.

[0069] This application provides a heterogeneous storage method for data, which can solve the problems of complex processes, repeated development and inflexible support for business changes caused by a large number of customized requirements in the current real-time data synchronization process of financial services. Figure 4 This application is an architectural diagram of a heterogeneous data storage method, such as Figure 4 As shown, it includes a policy center for generating target policy tasks, a message system for collecting data streams to be stored, an Apache Flink real-time synchronization module for determining target policy adapters and performing heterogeneous storage on data streams to be stored, and a heterogeneous storage module for storing data streams to be stored.

[0070] A real-time data collection system collects actor (user) data in real time and matches it against a database (which contains all historical and currently collected data). This data is then identified as changes, forming a data stream to be stored and an incremental change log for the database. This data stream is then sent to a messaging system, which in turn sends it to the Apache Flink real-time synchronization module.

[0071] The policy center includes a built-in policy adapter and a policy manager. The built-in policy adapter generates multiple policy templates, which form a template list. The policy manager selects the corresponding policy template based on the storage task requirements of the data stream to be stored and formulates a policy schedule corresponding to the policy template. The policy schedule includes policy rules and scheduling resources, and generates the target policy task corresponding to the policy rule. The target policy task is then sent to the Apache Flink real-time synchronization module.

[0072] The Apache Flink real-time synchronization module performs synchronization tasks: it performs policy parsing and mapping for policy scheduling, and determines the corresponding target policy adapter, which includes contextual policy logic. Based on this contextual policy logic, it heterogeneously stores the data stream to be stored in the corresponding target heterogeneous data storage. Target heterogeneous data storage includes databases such as Mongo, HBase, and ElasticSearch.

[0073] This application uses the strategy center to effectively abstract and automate the large amount of complex business development that requires customization during the real-time synchronization of financial data, avoiding a large amount of repetitive development work. At the same time, it greatly improves the flexibility of responding to changes in business needs through flexible policy rule configuration.

[0074] Furthermore, the Policy Center manages and creates policy schedules through the Policy Management Module. Policy schedules support policy templates in various formats, multiple built-in adapters for various business policies, and error checking based on built-in policy definitions. Once a policy schedule is created and the scheduled task is started, a real-time synchronization task is automatically initiated.

[0075] This application improves the adaptability to a large number of customized scenarios in the development of financial data synchronization through the functional integration of a data collection cluster, a policy center, a message system, and an Apache Flink real-time synchronization module composed of multiple servers. The policy template generated by the built-in policy adapter center is loaded through the policy center page, and the policy rule item configuration is completed according to the synchronization business requirements. The policy scheduler selects the generated policy rule data to create a target policy task, and sends the target policy task to the Apache Flink real-time synchronization cluster for synchronization. The Apache Flink real-time synchronization module parses the target policy task and finds the corresponding target policy adapter through the policy mapping module. The policy adapter adapts the policy rule data to process the database change data (data stream to be stored) in the message system and synchronizes the data to the heterogeneous target database in real time. Through the above process, this application simplifies complex business change requirements with configurable policy rule operations through the configuration center, and leaves it to the built-in policy adapter to automatically schedule tasks to perform synchronization functions, effectively solving the current problems of complex processes, repeated development, and inflexible support for business changes in business environments with a large number of customized requirements.

[0076] like Figure 5 As shown, the present application provides a heterogeneous storage system for data, including:

[0077] The acquisition module is used to obtain the target policy tasks and the data stream to be stored;

[0078] A determination module, configured to determine a target policy adapter according to a target policy task, wherein the target policy adapter includes a contextual policy logic;

[0079] The storage module is used to perform heterogeneous storage on the data stream to be stored according to the context policy logic; wherein the heterogeneous storage represents storing different types of data streams to be stored in corresponding adapted heterogeneous data storage devices.

[0080] Optionally, the storage module is specifically configured to:

[0081] Determine the type of data stream to be stored based on data storage requirements;

[0082] According to the type of data stream to be stored, the corresponding target heterogeneous data storage is determined from the preset storage medium library;

[0083] According to the data storage requirements, the data stream to be stored is stored in the target heterogeneous data storage.

[0084] Optionally, obtain a module, specifically for:

[0085] Generate policy templates using the preset built-in policy adapter;

[0086] Formulate policy scheduling corresponding to policy templates. Policy scheduling includes policy rules and scheduling resources.

[0087] Generate corresponding policy tasks to be executed according to policy rules;

[0088] Verify policy rules and scheduling resources to obtain verification results;

[0089] When the verification result indicates that the verification has passed, the policy task to be executed is used as the target policy task.

[0090] Optionally, obtain a module, specifically for:

[0091] Determine whether the rule constraints corresponding to the policy rules and the corresponding built-in configurators are legal;

[0092] Determine whether the scheduled resources fall within the preset resource range;

[0093] Determine whether the scheduling resources are legal.

[0094] Optionally, obtain a module, specifically for:

[0095] When the verification result indicates a verification failure, a new scheduling policy corresponding to the policy template is formulated. The new policy scheduling includes new policy rules and new scheduling resources.

[0096] The new policy rules and the new scheduling resources are verified, and the verification result is determined until the verification result indicates that the verification has passed.

[0097] Optionally, obtain a module, specifically for:

[0098] Collect multiple user data in real time;

[0099] Perform a search operation on each piece of user data in a preset database to determine a matching result corresponding to each piece of user data;

[0100] For each piece of user data, if the matching result indicates that no matching data is found, the user data is treated as the changed data;

[0101] All changed data is treated as the data stream to be stored.

[0102] Optionally, a module is determined, specifically for:

[0103] Analyze the target strategy task and obtain the analysis result;

[0104] The parsing results are mapped and matched in the preset policy adapter library to determine the corresponding target policy adapter.

[0105] A computing device according to an embodiment of the present application includes a memory, a processor, and a program stored in the memory and running on the processor. When the processor executes the program, some or all steps of the above-mentioned heterogeneous storage method for data are implemented.

[0106] Among them, the computing device can be a computer, and correspondingly, its program is computer software. The above-mentioned parameters and steps in a computing device of the present application can refer to the parameters and steps in an embodiment of a heterogeneous storage method for data above, and will not be repeated here.

[0107] In an embodiment of the present application, a computer-readable storage medium is provided, in which instructions are stored. When the instructions are executed, the steps of the above-mentioned method for heterogeneous storage of data are executed.

[0108] The computer-readable storage medium may be a transient computer-readable storage medium or a non-transitory computer-readable storage medium.

[0109] The technical solution of the embodiments of the present disclosure can be embodied in the form of a software product, which is stored in a storage medium and includes one or more instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method of the embodiments of the present disclosure. The aforementioned computer-readable storage medium can be a non-transitory computer-readable storage medium, including: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program code, or a transient computer-readable storage medium.

[0110] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Among them, each box in the flowchart or block diagram can represent a module, program segment, or part of the code, and the above-mentioned module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0111] Those skilled in the art will appreciate that the present application may be implemented as a system, method, or computer program product. Therefore, the present disclosure may be specifically implemented in the following forms, namely, in the form of complete hardware, complete software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, the present application may also be implemented in the form of a computer program product in one or more computer-readable media, the computer-readable medium containing a computer-readable program code. Computer-readable storage media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or components, or any combination thereof.

[0112] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0113] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and cannot be understood as limitations on the present application. Ordinary technicians in this field can change, modify, replace and modify the above embodiments within the scope of the present application.

Claims

1. A method for heterogeneous storage of data, characterized in that: include: Obtain the target policy tasks and data streams to be stored; Parsing the target policy task to obtain a parsing result, wherein the parsing result corresponds to the execution adaptation requirement; The parsing result is mapped and matched in a preset policy adapter library to determine a target policy adapter with the same execution adaptation requirement; wherein the preset policy adapter library includes multiple preset policy adapters, the preset policy adapters are data storage adapters, type adapters, security adapters or data format adapters, and the target policy adapter includes contextual policy logic, and the contextual policy logic includes data storage requirements for storing different types of data streams to be stored in corresponding adapted heterogeneous data storage devices; According to the contextual policy logic, the data stream to be stored is heterogeneously stored; wherein the heterogeneous storage represents storing different types of data streams to be stored in corresponding adapted heterogeneous data storage devices; The data stream to be stored is heterogeneously stored according to the contextual policy logic, including: determining the type of the data stream to be stored according to the data storage requirements, the type of the data stream to be stored including fact tables and dimension tables; determining the corresponding target heterogeneous data storage from a preset storage media library according to the type of the data stream to be stored, the target heterogeneous data storage including the ElsticSearch database; storing the data stream to be stored in the target heterogeneous data storage according to the data storage requirements.

2. The method according to claim 1, characterized in that The target strategy acquisition task includes: Generate policy templates using the preset built-in policy adapter; Formulate a policy schedule corresponding to the policy template, wherein the policy schedule includes policy rules and scheduling resources; Generate corresponding policy tasks to be executed according to the policy rules; Verifying the policy rules and the scheduling resources to obtain a verification result; When the verification result indicates that the verification is passed, the to-be-executed policy task is used as the target policy task.

3. The method according to claim 2, characterized in that The verifying the policy rule and the scheduling resource includes: Determine whether the rule constraints corresponding to the policy rule and the corresponding built-in configurator are legal; Determining whether the scheduled resources fall within a preset resource range; Determine whether the scheduled resources are legal.

4. The method according to claim 2, characterized in that Also includes: When the verification result indicates a verification failure, formulating a new scheduling policy corresponding to the policy template, wherein the new policy scheduling includes new policy rules and new scheduling resources; The new policy rule and the new scheduling resource are verified, and a verification result is determined until the verification result indicates that the verification has passed.

5. The method according to any one of claims 1 to 4, characterized in that The data stream to be stored includes at least one piece of changed data, and obtaining the data stream to be stored includes: Collect multiple user data in real time; Perform a search operation on each piece of user data in a preset database to determine a matching result corresponding to each piece of user data; For each piece of user data, when the matching result indicates that no matching data is found, the user data is used as the changed data; All the changed data are taken as the data stream to be stored.

6. A heterogeneous storage system for data, characterized in that: include: The acquisition module is used to obtain the target policy tasks and the data stream to be stored; A determination module, configured to parse the target policy task and obtain a parsing result, wherein the parsing result corresponds to the execution adaptation requirement; The parsing result is mapped and matched in a preset policy adapter library to determine a target policy adapter with the same execution adaptation requirement; wherein the preset policy adapter library includes multiple preset policy adapters, the preset policy adapters are data storage adapters, type adapters, security adapters or data format adapters, and the target policy adapter includes contextual policy logic, and the contextual policy logic includes data storage requirements for storing different types of data streams to be stored in corresponding adapted heterogeneous data storage devices; A storage module, configured to perform heterogeneous storage on the data stream to be stored according to the contextual policy logic; wherein the heterogeneous storage represents storing different types of data streams to be stored in corresponding adapted heterogeneous data storage devices; The storage module is specifically used to: determine the type of the data stream to be stored according to the data storage requirements, and the type of the data stream to be stored includes a fact table and a dimension table; determine the corresponding target heterogeneous data storage from a preset storage medium library according to the type of the data stream to be stored, and the target heterogeneous data storage includes an ElsticSearch database; store the data stream to be stored in the target heterogeneous data storage according to the data storage requirements.

7. A computing device comprising a memory, a processor, and a program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the steps of the method for heterogeneous storage of data as described in any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium, characterized in that The computer-readable storage medium stores instructions, and when the instructions are executed on a terminal device, the terminal device executes the steps of a heterogeneous storage method for data according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Heterogeneous storage expansion system and method

    CN108959398A

  • Distributed multi-data source acquisition implementation method

    CN117033952A