Data management method and device, equipment, medium and product
By combining the data management method of Bloom filter and B+ tree file database in streaming computing, the problems of data query efficiency and storage compatibility are solved, and efficient data query and storage are achieved.
Patent Information
- Application Number
- CN202511040864.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-10
AI Technical Summary
Existing data storage management solutions cannot balance data query efficiency and large-scale data storage, especially in streaming computing, they cannot meet the requirements of high concurrency query requests and low latency.
A combination of preset Bloom filters and B+ tree file databases is used to store incremental and existing data of the business system respectively. The Bloom filter is used to quickly determine the existence of query data, and efficient queries are performed in the B+ tree file database.
It realizes efficient data query in streaming computing, supports massive data storage, reduces query latency and system resource consumption, and improves query efficiency.
Smart Images

Figure CN120763173A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of computer technology, and in particular to a data management method, apparatus, device, medium, and product. Background Art
[0002] In streaming data processing, table lookups are often required to retrieve the corresponding data. Lookup tables typically contain large amounts of data that changes infrequently, but are frequently queried. To meet the real-time requirements of streaming computing, lookup tables must support high concurrency and extremely low query latency. Furthermore, the data in the lookup tables must maintain a high degree of accuracy and consistency to ensure the correctness of decisions and operations based on the lookup tables.
[0003] Currently, data storage is mostly done using in-memory data storage, columnar storage formats, traditional relational databases using B+ trees, or LSM tree (Log-Structured Merge Tree) databases to support data lookup.
[0004] However, in-memory data storage cannot handle massive amounts of data, and using a distributed solution is extremely costly. Columnar storage formats have low query efficiency and cannot efficiently support updates. Traditional relational databases using B+ trees cannot efficiently compress data, consume a lot of storage space, and have low write efficiency. Initial writes of massive amounts of data take a long time to complete, and data cannot be loaded locally on compute nodes. LSM tree databases, such as HBase, offer high write efficiency, but because data may exist in many ordered key-value files, query efficiency is low. No single data organization method can meet the data search requirements of streaming computing. Summary of the Invention
[0005] Embodiments of the present invention provide a data management method, apparatus, device, medium, and product that can support efficient data query in streaming computing while achieving large-scale data storage.
[0006] In a first aspect, an embodiment of the present invention provides a data management method, the method comprising:
[0007] In response to the data query request, determining a data field identifier of the target query data;
[0008] Determine whether the target query data corresponding to the data field identifier is in the corresponding preset incremental data set through a preset Bloom filter; wherein the preset incremental data set stores the incremental data of the business data of the target business system based on a B+ tree file database;
[0009] In response to the target query data being in the preset incremental data set, a data query is performed in the preset incremental data set according to the data field identifier to obtain the target query data.
[0010] In a second aspect, an embodiment of the present invention provides a data management device, the device comprising:
[0011] A request response module, configured to determine a data field identifier of target query data in response to a data query request;
[0012] A data pre-query module is used to determine whether the target query data corresponding to the data field identifier is in the corresponding preset incremental data set through a preset Bloom filter; wherein the preset incremental data set stores the incremental data of the business data of the target business system based on a B+ tree file database;
[0013] The data query module is used for performing data query in the preset incremental data set according to the data field identifier to obtain the target query data in response to the target query data being in the preset incremental data set.
[0014] In a third aspect, an embodiment of the present invention further provides a computer device, comprising:
[0015] one or more processors;
[0016] a memory for storing one or more programs;
[0017] When the one or more programs are executed by one or more processors, the one or more processors implement the data management method provided by any embodiment of the present invention.
[0018] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the data management method provided by any embodiment of the present invention.
[0019] In a fifth aspect, an embodiment of the present invention further provides a computer program product, comprising a computer program, which, when executed by a processor, implements the data management method provided by any embodiment of the present invention.
[0020] The embodiments of the above invention have the following advantages or beneficial effects:
[0021] In an embodiment of the present invention, in response to a data query request, a data field identifier of target query data is determined; through a preset Bloom filter, whether the target query data corresponding to the data field identifier is in a corresponding preset incremental data set is determined; wherein the preset incremental data set stores incremental data of the business data of the target business system based on a B+ tree file database; in response to the target query data being in the preset incremental data set, a data query is performed in the preset incremental data set according to the data field identifier to obtain the target query data. The technical solution of the embodiment of the present invention solves the problem that existing data storage management solutions cannot take into account both data query efficiency and large-scale data storage. It can support efficient data query in streaming computing while achieving large-scale data storage. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 is a flow chart of a data management method provided by an embodiment of the present invention;
[0023] Figure 2 is a flow chart of another data management method provided by an embodiment of the present invention;
[0024] Figure 3 This is a structural diagram of a data management device provided by an embodiment of the present invention;
[0025] Figure 4 It is a structural diagram of a computer device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0026] The present invention will be further described in detail below with reference to the accompanying drawings and examples. It will be understood that the specific embodiments described herein are intended only to illustrate the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all structures.
[0027] Figure 1 This is a flowchart of a data management method provided in an embodiment of the present invention. This embodiment is applicable to scenarios involving the management and query analysis of large amounts of data storage, particularly in streaming computing scenarios involving real-time data access and querying of data storage spaces. This method can be executed by a data management device, which can be implemented in software and / or hardware and integrated into a computer device with application development capabilities.
[0028] like Figure 1 As shown, the data management method of this embodiment includes the following steps:
[0029] S110 : In response to a data query request, determine a data field identifier of target query data.
[0030] Data query requests are requests for data queries to achieve business goals within a predefined business scenario. Predefined business scenarios can be of different types. For example, in a real-time monitoring scenario, the core of the query is "filtering real-time data that meets the predefined criteria." In a real-time statistics scenario, the core of the query is "aggregate results," focusing on "statistical values within a specific time window." In a dynamic recommendation scenario, the core of the query is "page browsing feature data," focusing on "time-sequenced page content."
[0031] Take real-time fraud detection in the telecommunications industry as an example. In this scenario, operators need to monitor user call and text message activity in real time to identify potential fraudulent activity. A lookup table might contain a user's credit score, historical behavior patterns, and known fraud characteristics. Whenever a user initiates a communication, the system must quickly query the table, comparing the user's current behavior with the data in the table to determine if there's fraud risk. This requires the table to handle massive amounts of data while ensuring rapid query response and real-time data updates. Therefore, the performance of the lookup table directly impacts the efficiency and effectiveness of the fraud detection system.
[0032] Data query requests contain structured data field definitions to identify the data fields to be queried. Examples include basic fields that describe the core attributes of the data, such as user IDs, order representations, or event times. They can also include derived fields generated through calculations, time fields and data source fields for time window constraints, and fields used for data filtering and selection conditions, such as numerical and logical conditions.
[0033] Furthermore, the data query request may be parsed using the syntax structure corresponding to the data query request to identify at least one of the above-mentioned fields contained in the data query request, thereby locking the data content to be queried and determining the data field identifier of the target query data. The data query request may be a request assigned to the computing node where the data management device is located.
[0034] S120 : Determine, through a preset Bloom filter, whether the target query data corresponding to the data field identifier is in the corresponding preset incremental data set.
[0035] A Bloom filter is used to determine whether an element exists in a set. It consists of a bit array and multiple hash functions. It can be constructed based on the data elements in a preset incremental dataset. It offers high space efficiency, high query efficiency, and is friendly to distributed scenarios. It can quickly determine whether the target query data is in the corresponding preset incremental dataset.
[0036] The preset incremental data center stores the incremental business data of the target business system based on a B+ tree file database. In other words, in this embodiment, the existing and incremental business data are stored and managed separately. Each compute node uses a B+ tree file system database similar to BoltDataBase to store and manage incremental data. The B+ tree structure provides efficient data updates and queries, making it particularly suitable for handling frequent data changes.
[0037] A Bloom filter is established for each B+ tree file system database. Every time data is updated or inserted, the Bloom filter is maintained accordingly to quickly determine whether the data exists in the incremental database. The introduction of the Bloom filter further optimizes the efficiency of incremental data queries by quickly eliminating non-existent keys and reducing unnecessary access to the B+ tree database.
[0038] Incremental data is stored in a single B+ tree file. Compared with the multiple incremental files generated by the LSM tree, it has higher query efficiency, lower maintenance costs, and better system resource utilization.
[0039] S130 : In response to the target query data being in the preset incremental data set, a data query is performed in the preset incremental data set according to the data field identifier to obtain the target query data.
[0040] If the target query data is in the preset incremental data set, the target query data can be quickly queried in the corresponding B+ tree and feedback can be provided. This avoids querying existing data and reduces data query pressure.
[0041] Furthermore, in response to the target query data not being included in the pre-set incremental data set, a data query can be performed based on the data field identifier in the database corresponding to the existing data of the target business system to obtain the target query data; wherein the database corresponding to the existing data stores data in an ordered key-value pair format on pre-set distributed computer nodes. In other words, in this embodiment, the business data is split into two parts: existing data and incremental data files. The existing part is stored using ordered key-value pairs, and the incremental part is stored using a B+ tree.
[0042] Specifically, the processing of stock data includes the following steps:
[0043] Before the data query request is acquired, the inventory data is data-sharded according to a preset data-sharding rule to obtain a plurality of inventory data sub-sets, and each inventory data sub-set is associated with a computing node in a preset distributed computing storage architecture. The storage format of data in each inventory data sub-set is converted into an ordered key-value pair storage format. That is, the inventory data shards are stored on each computing node, and this localized storage strategy reduces data cross-node transmission, reduces network delay, and improves query response speed. Each computing node is only responsible for data within its shard, reducing the load of a single node and improving query efficiency.
[0044] The preset data-sharding rule can be data-sharding according to attributes (such as key value, range, hash value, etc.) of the data itself. The range can be at least one of a numerical range, a time range, or a business range.
[0045] The technical solution of the embodiment responds to the data query request to determine the data field identifier of the target query data, determines whether the target query data corresponding to the data field identifier is in the corresponding preset incremental data set through the preset Bloom filter, wherein the incremental data in the preset incremental data set is stored based on the B+ tree file database of the business data of the target business system, and the target query data is obtained through data query in the preset incremental data set according to the data field identifier in response to the target query data in the preset incremental data set. The technical solution of the embodiment solves the problem that the existing data storage management scheme cannot balance data query efficiency and large data storage, and can support efficient data query in stream computing while realizing large data storage.
[0046] Figure 2 A flowchart of a data management method provided by the embodiment of the application, the data management method in the embodiment and the above-mentioned embodiments belong to the same inventive concept, and further describes the process of incremental data updating. The method can be executed by a data management device, which can be realized by software and / or hardware and integrated in a computer device with application development function.
[0047] As shown in FIG. 1, the data management method of the embodiment includes the following steps: Figure 2
[0048] S210, in response to the data update dynamic, updating the update data associated with the data update dynamic to the preset incremental data set.
[0049] Data update dynamics are sent to the computing node where the data management device resides in the form of a streaming change log. This data update dynamics includes information such as the corresponding operation type, data source table name, data before and after the change, and the event that caused the change. This allows the updated data associated with the data update dynamics to be updated in the pre-set incremental dataset.
[0050] The preset incremental data set is the data set for the corresponding computing node to manage incremental data. In this data set, the B+ tree structure is used for incremental data storage, providing efficient data update and query performance.
[0051] S220: Update the Bloom filter corresponding to the preset incremental data set based on the updated data.
[0052] The data in the streaming change log is updated in real time to the corresponding B+ tree file system database, and the Bloom filter is updated synchronously. This real-time update mechanism ensures data consistency and freshness.
[0053] When a new element is inserted or an existing element is deleted in the preset incremental data set, a Bloom filter update is triggered. The update of the Bloom filter can be a filter reconstruction or a dynamic expansion.
[0054] S230: Merge the data in the updated preset incremental data set with the stock data of the corresponding computing node to obtain a new version of the stock data, and clear the updated preset incremental data set.
[0055] Specifically, when the amount of data in the updated preset incremental data set is greater than the preset data amount threshold, the data in the updated preset incremental data set will be merged with the existing data of the corresponding computing node; or, according to the preset existing data update time condition, the data in the updated preset incremental data set will be merged with the existing data of the corresponding computing node.
[0056] Merge incremental data with existing data to generate a new version of the existing data file containing the latest incremental data, reducing the amount of incremental data and alleviating the pressure on the incremental data. The merged new version of the existing data file replaces the old version and becomes the new query base file. This helps control the growth of incremental data and prevent the incremental data set from becoming too large.
[0057] After the merge operation, the old incremental data is integrated into the existing data, and the new incremental data set starts from zero. This cycle is repeated, ensuring the effective use of system resources and the stability of query performance.
[0058] The merge strategy can be dynamically adjusted based on the actual load and performance feedback of the system to adapt to different data update frequencies and query requirements.
[0059] S240: In response to the data query request, determine the data field identifier of the target query data.
[0060] S250: Determine, through a preset Bloom filter, whether the target query data corresponding to the data field identifier is in the corresponding preset incremental data set.
[0061] The preset incremental data set stores the incremental data of the target business system's business data based on a B+ tree file database. The preset incremental data set is the data set updated in the above steps.
[0062] Incremental data is stored in a single localized B+ tree file. This file system database provides fast data insertion and query capabilities and is particularly suitable for handling frequent data updates.
[0063] The introduction of Bloom filters further optimizes the query efficiency of incremental data. It can quickly exclude non-existent keys and reduce unnecessary access to the B+ tree database.
[0064] Since the incremental dataset contains only the changed data, its data volume is relatively small and the query scope is limited, which makes the query operation faster.
[0065] S260 : In response to the target query data being in the preset incremental data set, perform a data query in the preset incremental data set according to the data field identifier to obtain the target query data.
[0066] The technical solution of this embodiment is to update the updated data associated with the data update dynamics into the preset incremental data set in response to the data update dynamics; wherein the data update dynamics are presented in the form of a streaming change log; update the bloom filter corresponding to the preset incremental data set based on the updated data; merge the data in the updated preset incremental data set with the stock data of the corresponding computing node to obtain the new version of the stock data, and clear the updated preset incremental data set; in response to the data query request, determine the data field identifier of the target query data; determine whether the target query data corresponding to the data field identifier is in the corresponding preset incremental data set through the preset bloom filter; in response to the target query data being in the preset incremental data set, perform data query in the preset incremental data set according to the data field identifier to obtain the target query data. The technical solution of the embodiment of the present invention solves the problem that the existing data storage management solution cannot take into account both data query efficiency and large-scale data storage. It can support efficient data query in streaming computing while realizing large-scale data storage, and can also efficiently associate updated data.
[0067] Figure 3A structural schematic diagram of a data management device provided by an embodiment of the present application is shown in the figure. The embodiment can be applied to the scene of data management storage and query, especially real-time data query in stream computing. The data management device can be realized by software and / or hardware and integrated in a computer terminal device with application development function.
[0068] As shown in the figure, the data management device comprises a request response module 310, a data pre-query module 320 and a data query module 330. Figure 3
[0069] The request response module 310 is configured to determine the data field identifier of the target query data in response to a data query request. The data pre-query module 320 is configured to determine whether the target query data corresponding to the data field identifier is in the corresponding preset incremental data set through a preset Bloom filter. The preset incremental data set stores incremental data of the business data of the target business system based on a B+ tree file database. The data query module 330 is configured to perform data query in the preset incremental data set according to the data field identifier to obtain the target query data in response to the target query data being in the preset incremental data set.
[0070] The technical solution of the embodiment determines the data field identifier of the target query data in response to a data query request, determines whether the target query data corresponding to the data field identifier is in the corresponding preset incremental data set through a preset Bloom filter, stores incremental data of the business data of the target business system based on a B+ tree file database in the preset incremental data set, and performs data query in the preset incremental data set according to the data field identifier to obtain the target query data in response to the target query data being in the preset incremental data set. The technical solution of the embodiment of the present application solves the problem that the existing data storage management scheme cannot balance data query efficiency and large data storage, and can support efficient data query in stream computing while realizing large data storage.
[0071] In an optional implementation, the data query module 330 can also be configured to:
[0072] perform data query in the database corresponding to the stock data of the target business system according to the data field identifier to obtain the target query data in response to the target query data not being in the preset incremental data set.
[0073] The database corresponding to the stock data stores data in a preset distributed computer node in an ordered key-value pair format.
[0074] In an optional implementation, the data management device further comprises a data storage management module configured to:
[0075] Before obtaining the data query request, the stock data is sharded according to a preset data sharding rule to obtain a plurality of stock data subsets, and each of the stock data subsets is associated with a computing node in a preset distributed computing storage architecture;
[0076] The storage format of the data in each of the existing data subsets is converted into an ordered key-value pair storage format.
[0077] In an optional embodiment, the data management device further includes a data storage management module, and can also be used to:
[0078] In response to the data update dynamics, updating the updated data associated with the data update dynamics into a preset incremental data set; wherein the data update dynamics are presented in the form of a streaming change log;
[0079] Update the Bloom filter corresponding to the preset incremental data set based on the updated data.
[0080] In an optional embodiment, the data management device further includes a data storage management module, and can also be used to:
[0081] The data in the updated preset incremental data set is merged with the stock data of the corresponding computing node to obtain a new version of the stock data, and the updated preset incremental data set is cleared.
[0082] In an optional embodiment, the data management device further includes a data storage management module, and can also be used to:
[0083] When the amount of data in the updated preset incremental data set is greater than a preset data amount threshold, merging the data in the updated preset incremental data set with the stock data of the corresponding computing node; or
[0084] According to the preset stock data update time condition, the data in the updated preset incremental data set is merged with the stock data of the corresponding computing node.
[0085] The data management device provided by the embodiment of the present invention can execute the data management method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0086] Figure 4 A schematic structural diagram of a computer device provided in an embodiment of the present invention. Figure 4 A block diagram of an exemplary computer device 12 suitable for use in implementing embodiments of the present invention is shown. Figure 4The computer device 12 shown is only an example and should not limit the functionality and scope of use of the embodiments of the present invention. The computer device 12 can be any terminal device with computing capabilities, such as an intelligent controller, a server, a mobile phone, or other terminal devices.
[0087] like Figure 4 As shown, computer device 12 is implemented as a general-purpose computing device. Components of computer device 12 may include, but are not limited to, one or more processors or processing units 16, system memory 28, and a bus 18 that connects various system components (including system memory 28 and processing unit 16).
[0088] Bus 18 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, an Industry Standard Architecture (ISA) bus, a Micro Channel Architecture (MAC) bus, an Enhanced ISA bus, a Video Electronics Standards Association (VESA) local bus, and a Peripheral Component Interconnect (PCI) bus.
[0089] The computer device 12 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computer device 12, including volatile and non-volatile media, removable and non-removable media.
[0090] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache 32. Computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be configured to read and write non-removable, non-volatile magnetic media ( Figure 4 Not shown, often called a "hard drive"). Although Figure 4 Not shown, a magnetic disk drive for reading and writing to a removable non-volatile magnetic disk (e.g., a "floppy disk"), and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a CD-ROM, DVD-ROM, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or more data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.
[0091] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 42 generally perform the functions and / or methods of the embodiments described herein.
[0092] The computer device 12 may also communicate with one or more external devices 14 (e.g., a keyboard, a pointing device, a display 24, etc.), one or more devices that enable a user to interact with the computer device 12, and / or any device that enables the computer device 12 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). Such communication may be performed via an input / output (I / O) interface 22. Furthermore, the computer device 12 may also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) via a network adapter 20. As shown, the network adapter 20 communicates with the other modules of the computer device 12 via the bus 18. It should be understood that although Figure 4 Not shown, other hardware and / or software modules may be used in conjunction with the computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, artificial intelligence systems, tape drives, and data backup storage systems.
[0093] The processing unit 16 executes various functional applications and data management by running programs stored in the system memory 28, such as implementing the data management method provided in the embodiment of the present invention, which includes:
[0094] In response to the data query request, determining a data field identifier of the target query data;
[0095] Determine, through a preset Bloom filter, whether the target query data corresponding to the data field identifier is in the corresponding preset incremental data set; wherein the preset incremental data set stores the incremental data of the business data of the target business system based on a B+ tree file database;
[0096] In response to the target query data being in the preset incremental data set, a data query is performed in the preset incremental data set according to the data field identifier to obtain the target query data.
[0097] An embodiment of the present invention further provides a computer-readable storage medium having a computer program stored thereon. When the program is executed by a processor, the data management method provided in any embodiment of the present invention is implemented. The method includes:
[0098] In response to the data query request, determining a data field identifier of the target query data;
[0099] Determine, through a preset Bloom filter, whether the target query data corresponding to the data field identifier is in the corresponding preset incremental data set; wherein the preset incremental data set stores the incremental data of the business data of the target business system based on a B+ tree file database;
[0100] In response to the target query data being in the preset incremental data set, a data query is performed in the preset incremental data set according to the data field identifier to obtain the target query data.
[0101] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of computer-readable storage media include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device.
[0102] A computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0103] Program code embodied on a computer-readable medium may be transmitted using any appropriate medium, including but not limited to wireless, wireline, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0104] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, Python, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0105] The embodiments of the present application further provide a computer program product, comprising a computer program which, when executed by a processor, implements the data management method provided by any of the embodiments of the present application.
[0106] Computer program code for carrying out operations of the present application can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, Python, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0107] Those skilled in the art should clearly understand that each module or each step of the present application described above can be implemented by using a general computing device, and can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. Alternatively, each module or each step can be implemented by using computer device executable program code, so that each module or each step can be stored in a storage device and executed by a computing device, or each module or each step can be respectively manufactured as an integrated circuit module, or multiple modules or steps can be manufactured as a single integrated circuit module. Thus, the present application is not limited to any specific combination of hardware and software.
[0108] Note that the above merely describes preferred embodiments of the present application and the principles of the technology applied. Those skilled in the art will understand that the present application is not limited to the specific embodiments described herein, and that various obvious changes, modifications and substitutions can be made without departing from the scope of the present application. Therefore, although the present application has been described in detail through the above embodiments, the present application is not limited to the above embodiments, and can include more other equivalent embodiments without departing from the concept of the present application, and the scope of the present application is determined by the scope of the claims.
Claims
1. A data management method, characterized in that: include: In response to the data query request, determining a data field identifier of the target query data; Determining, by a preset Bloom filter, whether the target query data corresponding to the data field identifier is in a corresponding preset incremental data set; wherein the preset incremental data set stores incremental data of the business data of the target business system based on a B+ tree file database; In response to the target query data being in the preset incremental data set, a data query is performed in the preset incremental data set according to the data field identifier to obtain the target query data.
2. The method according to claim 1, characterized in that The method further comprises: In response to the target query data not being in the preset incremental data set, performing a data query in a database corresponding to the stock data of the target business system according to the data field identifier to obtain the target query data; The database corresponding to the stock data stores data in a preset distributed computer node in an ordered key-value pair format.
3. The method according to claim 2, characterized in that Before obtaining the data query request, the method further includes: Slice the stock data according to a preset data slicing rule to obtain a plurality of stock data subsets, and associate each of the stock data subsets with a computing node in a preset distributed computing storage architecture; The storage format of the data in each of the existing data subsets is converted into an ordered key-value pair storage format.
4. The method according to claim 1, wherein The method further comprises: In response to a data update dynamic, updating the update data associated with the data update dynamic into a preset incremental data set; wherein the data update dynamic is presented in the form of a streaming change log; The Bloom filter corresponding to the preset incremental data set is updated based on the updated data.
5. The method according to claim 4, characterized in that The method further comprises: The data in the updated preset incremental data set is merged with the stock data of the corresponding computing node to obtain a new version of the stock data, and the updated preset incremental data set is cleared.
6. The method according to claim 5, characterized in that Merging the updated data in the preset incremental data set with the stock data of the corresponding computing node includes: When the amount of data in the updated preset incremental data set is greater than a preset data amount threshold, merging the data in the updated preset incremental data set with the stock data of the corresponding computing node; or According to the preset stock data update time condition, the data in the updated preset incremental data set is merged with the stock data of the corresponding computing node.
7. A data management device, characterized in that: include: A request response module, configured to determine a data field identifier of target query data in response to a data query request; A data pre-query module is used to determine whether the target query data corresponding to the data field identifier is in a corresponding preset incremental data set through a preset Bloom filter; wherein the preset incremental data set stores incremental data of the business data of the target business system based on a B+ tree file database; The data query module is configured to query the target query data in the preset incremental data set according to the data field identifier in response to the target query data being in the preset incremental data set to obtain the target query data.
8. A computer device, characterized in that: The computer device comprises: one or more processors; a memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data management method according to any one of claims 1 to 6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the data management method according to any one of claims 1 to 6 is implemented.
10. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the computer program implements the data management method according to any one of claims 1 to 6.