Data completion method and device
By creating a cache pool on the local disk and using CQEngine and RocksDB tools to complete data, the problems of high operation and maintenance costs and inaccurate data in the existing technology are solved, and high-performance data completion is achieved.
Patent Information
- Application Number
- CN202510475820.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-16
- Publication Date
- 2025-07-08
AI Technical Summary
The existing data completion methods have problems such as high operation and maintenance costs, high learning costs and inaccurate associated data.
Use CQEngine tool and RocksDB tool to create a local cache pool on the local disk, build routing data and dictionary data by parsing the structured data to be completed, and use the local cache pool to match and complete data, avoiding relying on big data services and data warehouses.
It achieves no operation and maintenance costs, is out of the box, and ensures high-performance data completion effect, reducing the degree of intrusion of code.
Smart Images

Figure CN120277084A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a data completion method and apparatus. Background Art
[0002] In the process of data processing, it is inevitable to associate a batch of data content, and according to the matching conditions, obtain the corresponding information from this batch of data and supplement it to the content of the current data. For example, to complete the class information of a batch of student data, there is a class ID in the student data. At this time, it is necessary to obtain the class information that meets the class ID from the class data set (i.e., the associated data) and supplement it to the student data. Therefore, a method of associated data completion is introduced.
[0003] In the prior art, one is to use the ready-made interfaces provided by mature big data frameworks such as Flink and Spark to complete the above operations, but it requires the operation and maintenance experience of big data services such as Flink and Spark, and deploy the corresponding services. This method requires the access of third-party services, and the operation and maintenance cost is relatively high. At the same time, certain experience in using these big data frameworks is required, otherwise there will be problems with the accuracy and performance of data processing, and the learning cost is relatively high. The other belongs to secondary processing. Using a storage medium of an intermediate database type, the data is saved to a data warehouse such as ClickHouse or Hive, and then the associated data is used for data completion in a manner similar to SQL or batch processing tasks. This method can only use the batch processing method, and there will be a bottleneck in performance. Moreover, when obtaining data from one data source and importing it into the data warehouse, the data type will be distorted because it needs to be compatible with the type of the data warehouse. Therefore, there will be a problem of inaccurate associated data.
[0004] In summary, the problems of high operation and maintenance cost, high learning cost, and inaccurate associated data existing in the existing data completion methods are urgent problems to be solved at present. Summary of the Invention
[0005] In view of this, the embodiments of the present invention provide a data completion method and apparatus to achieve the purpose of no operation and maintenance cost, out-of-the-box use, and high performance.
[0006] To achieve the above object, the embodiments of the present invention provide the following technical solutions:
[0007] The first aspect of the embodiments of the present invention discloses a data completion method, and the method includes:
[0008] Receiving structured data to be completed;
[0009] Parsing the structured data to be completed to obtain the field names and field contents of one or more target fields;
[0010] Query the matching routing data from the pre-built routing set based on the field name of the target field; each piece of the routing data corresponds to a local cache pool; the local cache pool is pre-created on the local disk using the CQEngine tool and the RocksDB tool; the local cache pool pre-stores local data; the local data is obtained by assigning the first dictionary data containing business data content to a storage object;
[0011] Based on the field name and field content of the target field, construct the second dictionary data that does not contain business data content;
[0012] Query the local data matching the second dictionary data from the local cache pool corresponding to the routing data;
[0013] Use the local data to complete the structured data to be completed.
[0014] Preferably, the process of storing the local data into the local cache pool includes:
[0015] Receive the structured data to be stored;
[0016] Parse the structured data to be stored, and distinguish the condition fields carrying matching conditions and the business fields carrying business data content;
[0017] Construct the first dictionary data based on the condition fields and the business fields;
[0018] Based on the field name of the condition fields and the field name of the business fields, construct routing data and add the routing data to the pre-built routing set;
[0019] Concatenate the field name of the condition field and the field name of the business field to obtain a key value;
[0020] Judge whether there is a local cache pool corresponding to the key value;
[0021] If so, assign the first dictionary data to a storage object to obtain local data, and store the local data into the local cache pool;
[0022] If not, create a local cache pool corresponding to the key value, assign the first dictionary data to a storage object to obtain local data, and store the local data into the local cache pool.
[0023] Preferably, the constructing the first dictionary data based on the condition fields and the business fields includes:
[0024] Distinguish the IP type fields and non-IP type fields from the condition fields;
[0025] Sort the field names of each of the non-IP type fields, and splice the field contents of each of the non-IP type fields according to the sorting to obtain a string;
[0026] Sort the field names of each of the IP type fields, and put the field contents of each of the IP type fields into an array according to the sorting to obtain a target array;
[0027] Save the string, the target array, and the service data content carried by the service field into different fields respectively to obtain first dictionary data.
[0028] Preferably, constructing second dictionary data that does not include service data content based on the field names and field contents of the target fields includes:
[0029] Distinguish IP type fields and non-IP type fields from the conditional fields;
[0030] Sort the field names of each of the non-IP type fields, and splice the field contents of each of the non-IP type fields according to the sorting to obtain a string;
[0031] Sort the field names of each of the IP type fields, and put the field contents of each of the IP type fields into an array according to the sorting to obtain a target array;
[0032] Save the string and the target array into different fields respectively to obtain second dictionary data.
[0033] Preferably, querying the local data matching the second dictionary data from the local cache pool corresponding to the routing data includes:
[0034] Construct a query statement based on the second dictionary data;
[0035] Use the query statement and the CQEngine tool to query the local data matching the second dictionary data from the local cache pool corresponding to the routing data.
[0036] Preferably, using the local data to complete the to-be-completed structured data includes:
[0037] Obtain the index data corresponding to the local data from a pre-constructed index data cache;
[0038] Use the index data to query the target data required for the to-be-completed structured data from the local data;
[0039] Complete the to-be-completed structured data by using the target data.
[0040] The second aspect of the embodiments of the present invention discloses a data completion device, and the device includes:
[0041] A first receiving unit, configured to receive to-be-completed structured data;
[0042] A first parsing unit, configured to parse the to-be-completed structured data to obtain the field names and field contents of one or more target fields;
[0043] A matching unit, configured to query matching routing data from a pre-constructed routing set based on the field names of the target fields; each routing data corresponds to a local cache pool; the local cache pool is pre-created on a local disk by using a CQEngine tool and a RocksDB tool; the local cache pool pre-stores local data; the local data is obtained by assigning first dictionary data containing service data content to a storage object;
[0044] A first constructing unit, configured to construct second dictionary data that does not contain service data content based on the field names and field contents of the target fields;
[0045] A querying unit, configured to query the local data matching the second dictionary data from the local cache pool corresponding to the routing data;
[0046] A completion unit, configured to complete the to-be-completed structured data by using the local data.
[0047] Preferably, the device further includes:
[0048] A second receiving unit, configured to receive to-be-stored structured data;
[0049] A second parsing unit, configured to parse the to-be-stored structured data to distinguish conditional fields carrying matching conditions and service fields carrying service data content;
[0050] A second constructing unit, configured to construct first dictionary data based on the conditional fields and the service fields;
[0051] A third constructing unit, configured to construct routing data based on the field names of the conditional fields and the field names of the service fields, and add the routing data to a pre-constructed routing set;
[0052] A splicing unit, configured to splice the field names of the conditional fields and the field names of the service fields to obtain a key value;
[0053] A storage unit is used to determine whether there is a local cache pool corresponding to the key value; if so, assign the first dictionary data to a storage object to obtain local data, and store the local data in the local cache pool; if not, create a local cache pool corresponding to the key value, assign the first dictionary data to a storage object to obtain local data, and store the local data in the local cache pool.
[0054] Preferably, the second construction unit is specifically used for:
[0055] Distinguish IP type fields and non-IP type fields from the condition fields;
[0056] Sort the field names of each non-IP type field, and splice the field contents of each non-IP type field according to the sorting to obtain a string;
[0057] Sort the field names of each IP type field, and put the field contents of each IP type field into an array according to the sorting to obtain a target array;
[0058] Save the string, the target array, and the service data content carried by the service field to different fields respectively to obtain the first dictionary data.
[0059] Preferably, the first construction unit is specifically used for:
[0060] Distinguish IP type fields and non-IP type fields from the condition fields;
[0061] Sort the field names of each non-IP type field, and splice the field contents of each non-IP type field according to the sorting to obtain a string;
[0062] Sort the field names of each IP type field, and put the field contents of each IP type field into an array according to the sorting to obtain a target array;
[0063] Save the string and the target array to different fields respectively to obtain the second dictionary data.
[0064] Based on the data completion method and device provided by the embodiments of the present invention above, receive the structured data to be completed; parse the structured data to be completed to obtain the field names and field contents of one or more target fields; query the matching routing data from the pre-constructed routing set based on the field names of the target fields; each of the routing data corresponds to a local cache pool; the local cache pool is pre-created on the local disk using the CQEngine tool and the RocksDB tool; the local cache pool pre-stores local data; the local data is obtained by assigning the first dictionary data containing business data content to a storage object; construct the second dictionary data that does not contain business data content based on the field names and field contents of the target fields; query the local data matching the second dictionary data from the local cache pool corresponding to the routing data; use the local data to complete the structured data to be completed. In this solution, the degree of invasive code is low and there is no need to rely on a data warehouse, thus achieving the goals of no operation and maintenance costs, out-of-the-box use, and high performance assurance. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only the embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained according to the provided drawings without creative work.
[0066] Figure 1 It is the overall business architecture diagram of a data completion engine disclosed in the embodiments of the present invention;
[0067] Figure 2 It is the flowchart of a data completion method disclosed in the embodiments of the present invention;
[0068] Figure 3 It is the flowchart of local data storage disclosed in the embodiments of the present invention;
[0069] Figure 4 It is the design schematic diagram of a dictionary data disclosed in the embodiments of the present invention;
[0070] Figure 5 It is the logical schematic diagram of storage and matching disclosed in the embodiments of the present invention;
[0071] Figure 6 It is the structure diagram of a data completion device disclosed in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0073] In this application, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0074] First, the technical terms appearing in this application are explained as follows:
[0075] The CQEngine (Collection Query Engine) tool is a high-performance Java in-memory retrieval engine that allows developers to perform SQL-like queries on Java collections with extremely low latency.
[0076] The RocksDB (Rocks Data Base) tool is a high-performance embedded key-value storage engine that adopts the LSM-Tree (Log-Structured Merge-Tree) architecture, optimized for high-speed storage hardware such as SSDs and Flash, and is suitable for scenarios that require high throughput and low latency.
[0077] Dictionary data is an object used for storage and query, and is a specially designed data object structure in this application, which can be understood as an intermediate data temporary carrier object.
[0078] The local cache pool is a rewritten cache pool that utilizes the ability of RocksDB to read and write local files to cache data sets in the CQEngine memory.
[0079] JAVABEAN is a special java object used to represent structured data and save data content.
[0080] Metadata is a type of data that describes field information and is used to represent the types of structured data fields, such as integers, strings, floating-point types, IP types, etc.
[0081] As can be seen from the background art, there are problems of high operation and maintenance costs, high learning costs, and inaccurate associated data in existing data completion methods.
[0082] Therefore, an embodiment of the present invention discloses a data completion method and device. The structured data to be stored is parsed into first dictionary data in advance, and then assigned to a storage object to obtain local data, and the local data is stored in the corresponding local cache pool. Thus, when receiving the structured data to be completed, it is parsed into second dictionary data, the matching local data is retrieved in the local cache pool, and the completion is completed using the local data. In this solution, the degree of intrusion code is low and there is no need to rely on a data warehouse, thereby achieving the purpose of no operation and maintenance costs, out-of-the-box use, and ensuring high performance.
[0083] As Figure 1 shown, it is an overall business architecture diagram of a data completion engine disclosed by an embodiment of the present invention. The business architecture includes: an interface layer, a storage layer, and an engine layer.
[0084] Among them, the interface layer is responsible for the data interaction of the entire engine, custom functions, etc.
[0085] Specifically, the interface layer includes: an initialization interface for initializing the data completion engine; a data storage interface for receiving the structured data to be stored and passing it to the storage layer; an index data interface for receiving the structured data to be completed and passing it to the storage layer; a custom index interface for providing users with a custom index function.
[0086] The storage layer is used to parse the structured data to be stored passed by the interface layer and convert it into first dictionary data; parse the structured data to be completed passed by the interface layer and convert it into second dictionary data.
[0087] Build index data based on local data and store it in the index data cache; build a local cache pool; manage the local cache pool and its data, including creating and destroying the local cache pool and performing addition, deletion, modification, and query on the data in the local cache pool.
[0088] The engine layer is used to screen out conditional fields and non-conditional fields from the parsed structured data to be stored using a condition extractor, build routing data based on the conditional fields and non-conditional fields and cache it; use a data engine cache to find the local cache pool bound to the routing data. If not found, establish a local cache pool through the storage layer and bind the routing data to the local cache pool; use a data operator and a data storage to assign the first dictionary data to a storage object to obtain local data and store it in the local cache pool corresponding to the routing data.
[0089] Use a condition matcher to find routing data that matches the condition fields in the structured data to be completed; based on the matching routing data, find the bound local cache pool; construct a query statement based on the second dictionary data, and find the matching local data from the local cache pool; obtain the index data corresponding to the local data, extract the target data from the local data based on the index data, and use the target data to complete the structured data to be completed.
[0090] Based on the overall business architecture of a data completion engine disclosed in the embodiments of the present invention above, as Figure 2 shown, is a flowchart of a data completion method disclosed in an embodiment of the present invention, including the following steps:
[0091] Step S201: Receive the structured data to be completed.
[0092] In step S201, the format of the structured data to be completed includes but is not limited to: JSON, CSV, JAVABEAN, and MAP.
[0093] Step S202: Parse the structured data to be completed to obtain the field names and field contents of one or more target fields.
[0094] In step S202, the target fields are used to match the condition fields in the first dictionary data in the local cache pool. After successful matching, the content in the first dictionary data can be used to complete the structured data to be completed.
[0095] Step S203: Query the matching routing data from the pre-constructed routing set based on the field names of the target fields.
[0096] Among them, each routing data corresponds to a local cache pool; the local cache pool is created on the local disk in advance using the CQEngine tool and the RocksDB tool; the local cache pool stores local data in advance; the local data is obtained by assigning the first dictionary data containing business data content to a storage object.
[0097] In the specific implementation process of step S203, according to the number of fields and field names of the target fields, obtain the routing data from the routing set in turn (the routing data is sorted in the routing set), and the one with fewer fields in the routing data is matched first. Obtain the corresponding field content from the target fields according to the field names in the routing data. If the field content cannot be obtained, it means that the routing data does not match, and then obtain the next routing data. If all the field names in the routing data can obtain the corresponding field content from the target fields, it means that the routing data matches.
[0098] As Figure 3As shown in the figure, it is a local data storage flow chart disclosed in an embodiment of the present invention, including the following steps:
[0099] Step S301: Receive the structured data to be stored.
[0100] In step S301, the formats of the structured data to be stored include but are not limited to: JSON, CSV, JAVABEAN, and MAP.
[0101] Step S302: Parse the structured data to be stored, and distinguish the conditional fields carrying matching conditions and the business fields carrying business data content.
[0102] Step S303: Construct the first dictionary data based on the conditional fields and the business fields.
[0103] In the specific implementation process of step S303, first, distinguish the IP type fields and non-IP type fields from the conditional fields.
[0104] It should be noted that the IP fields and non-IP fields can be screened out through the information provided by the metadata. Metadata is field description information, which describes the data type to which the field belongs. It is a part of the storage layer but belongs to another system and is not within the scope of discussion in this application. After docking with any system that can provide field description information, it can be used as long as it can meet the field description. The description information includes the field name and the field type.
[0105] Secondly, sort the field names of each non-IP type field, and splice the field contents of each non-IP type field according to the sorting to obtain a string.
[0106] Preferably, sort the field names of each non-IP type field in alphabetical order.
[0107] For example, there are fields msg1: abc, msg2: bcd, aIP: 127.0.0.1, bIP: 192.168.1.23 / 24. Then the msg1 and msg2 fields are non-ip. The non-ip fields are merged and spliced. In alphabetical order, msg1 is in front, so the non-ip field content is merged into abcbcd.
[0108] Then, sort the field names of each IP type field, and put the field contents of each IP type field into an array according to the sorting to obtain the target array.
[0109] For example, for the IP type field bIP: 192.168.1.23 / 24 and aIP: 127.0.0.1, the field names are sorted as aIP, bIP. Then, 127.0.0.1 is placed in the first element of the array with an index of 0, and 192.168.1.23 is placed in the second element of the array with an index of 1.
[0110] Finally, the business data content carried by the string, target array, and business field is saved to different fields respectively to obtain the first dictionary data.
[0111] Specifically, the first dictionary data consists of three fields. The business data content carried by the string, target array, and business field is assigned to these three fields respectively to obtain the first dictionary data.
[0112] It should be noted that the content of the IP type field and the non-IP type field are assigned separately to leave an expansion space for matching the IP type field. Because the field content of the IP type field has a range attribute. For example, if the IP is 192.168.12.12 / 24 as a field content of a field data, then the IP content of 192.158.12.23 should hit this data. Therefore, the calculation method of IP is different from that of non-IP. If the content of the IP type field is processed in the same way as the content of the non-IP type field, 192.158.12.23 cannot hit 192.168.12.12 / 24.
[0113] Such as Figure 4 shown, it is a design schematic diagram of a dictionary data disclosed in an embodiment of the present invention.
[0114] Among them, since the configuration information is used to generate the data of the open source rule engine framework, it is necessary to convert the data and generate a SimpleNullableAttribute object to describe the field attributes. The ConfigInfo object is responsible for the description matching of the field attributes. The conditionList field indicates which fields are responsible for condition matching, as well as the description of the field name and field content. The valuations are responsible for indicating which fields are assigned corresponding values after the conditions are met.
[0115] The dictionary data operation object must reference the ConfigInfo object collection and perform operations on the data through the marked fields, that is, the function of the data operator. For example, it can be determined through a string flag whether the accessed structured data is added to the local cache pool, deleted from the local cache pool, or updated in the corresponding data in the local cache pool, etc. Similar operations can also be completed through enumeration.
[0116] Step S304: Based on the field names of the condition fields and the business fields, construct routing data, and add the routing data to a pre-constructed routing set.
[0117] Among them, the routing data is used to locate the local cache pool during storage and query. For performance and data volume considerations, the local cache pool is bound to the routing data, and data with the same fields is stored in the same local cache pool.
[0118] For example, assume there are three pieces of structured data to be stored, represented by data a, data b, and data c. The fields in data a are csrcip, cdevip, and level, then the routing data is represented as [csrcip, cdevip, level]. The fields in data b are the same as those in data a, and the routing data is also represented as [csrcip, cdevip, level]. The fields in data c are csrcip and level, then the routing data is represented as [csrcip, level].
[0119] The routing set is a two-dimensional array set, that is, a two-dimensional array structure formed by using a list data nested list data structure, and it will automatically expand according to the size. Based on the above example, if the routing set is created for the first time, the routing set can be represented as [[csrcip, level], [cdevip, csrcip, level]], which only contains two pieces of routing data because the fields of data a and data b are the same, so they are merged into one piece of routing data and saved in the routing set.
[0120] At the same time, the routing data in the routing set is sorted according to the number of fields it contains, and the routing data with fewer fields is ranked in the front. In addition, the content put into the two-dimensional array also needs to be sorted alphabetically. For example, cdevip will be ranked in front of csrcip.
[0121] Step S305: Concatenate the field names of the condition fields and the business fields to obtain the key value.
[0122] For example, for the three pieces of data in the above example, the key value of data a is: cdevipcsrciplevel. Since the fields of data b are the same as those of data a, the key value is also the same. The key value of data c is csrciplevel.
[0123] Step S306: Determine whether there is a local cache pool corresponding to the key value; if so, execute Step S307; if not, execute Step S308.
[0124] In the embodiment of the present invention, a Map data structure is used to save multiple local cache pools. The key value of the Map is represented by a string concatenated from the field names in the routing data, and the value represents the local cache pool.
[0125] Step S307: Assign the first dictionary data to the storage object to obtain local data, and store the local data in the local cache pool.
[0126] Step S308: Create a local cache pool corresponding to the key value, assign the first dictionary data to the storage object to obtain local data, and store the local data in the local cache pool.
[0127] In steps S306 to S308, obtain the local cache pool from the Map through different key values. If not, create a local cache pool based on RocksDB and add it to the Map. At the same time, assign the first dictionary data to the storage object to obtain local data, and store the local data in the local cache pool.
[0128] The local cache pool constructs a memory cache interface based on CQEngine, and uses the storage logic of RocksDB to rewrite this interface. That is, it uses the local disk storage data ability of RocksDB to rewrite the com.googlecode.cqengine.IndexedCollection interface of CQEngine to construct a storage engine with local storage ability, replacing the default storage engine of CQEngine, achieving the ability to store data on the local disk (i.e., the local cache pool). It can persist data to the disk, enabling massive data to be stored on the disk instead of piling up in memory, thus achieving the purpose of saving memory.
[0129] Specifically, use CQEngine to assign the IP type fields and non-IP type fields in the first dictionary data to different storage objects to obtain local data, and store it in the local cache pool.
[0130] Step S204: Based on the field name and field content of the target field, construct a second dictionary data that does not contain business data content.
[0131] It should be noted that the second dictionary data is used to match the first dictionary data stored in the local cache pool. Therefore, the second dictionary data only needs to contain the conditional fields for matching.
[0132] In the specific implementation process of step S204, distinguish the IP type fields and non-IP type fields from the conditional fields; sort the field names of each non-IP type field, and splice the field contents of each non-IP type field according to the sorting to obtain a string; sort the field names of each IP type field, and put the field contents of each IP type field into an array according to the sorting to obtain a target array; save the string and the target array to different fields respectively to obtain the second dictionary data.
[0133] It can be understood that the construction process of the second dictionary data is similar to that of the first dictionary data, except that the fields for storing business data content in the second dictionary data are left blank.
[0134] Step S205: Query and obtain the local data matching the second dictionary data from the local cache pool corresponding to the routing data.
[0135] In the specific implementation process of step S205, a query statement is constructed based on the second dictionary data; using the query statement and the CQEngine tool, the local data matching the second dictionary data is queried and obtained from the local cache pool corresponding to the routing data.
[0136] It can be understood that a structured query statement of CQEngine is constructed according to the second dictionary data, and the local cache pool implemented by RocksDB is retrieved through CQEngine. Since all interfaces are implemented by RocksDB, CQEngine can operate on the local cache pool only with its own logic. Programming facing the interface avoids the implementation logic, so where the data is stored is transparent to CQEngine.
[0137] Step S206: Use the local data to complete the structured data to be completed.
[0138] In the specific implementation process of step S206, the index data corresponding to the local data is obtained from the pre-constructed index data cache; the target data required for the structured data to be completed is queried from the local data using the index data; the structured data to be completed is completed using the target data.
[0139] It should be noted that the index data is constructed correspondingly when the local data is obtained, and is used to achieve fast retrieval and matching of fields.
[0140] As Figure 5 shown, it is a logical schematic diagram of storage and matching disclosed in an embodiment of the present invention.
[0141] The storage logic is as follows:
[0142] The structured data to be stored is connected to the data storage module and the data routing module through the data receiving interface of the first data receiver. In the data storage module, it is converted into the first dictionary data through the data object, and in the data routing module, it is converted into routing data based on the structured data to be stored, and the mapping relationship between the routing data and the local cache pool is established or obtained. If the mapping relationship between the routing data and the local cache pool is established, caching is performed; the first dictionary data is assigned to the storage object to obtain the local data, and the local data is stored in the local cache pool mapped by the routing data through the storage operator.
[0143] The matching logic is as follows:
[0144] The structured data to be completed is accessed to a data routing module and a data matching module through a second data receiver; in the data routing module, corresponding routing data is queried based on the structured data to be completed to determine the local cache pool corresponding to the routing data; in the data matching module, second dictionary data is constructed based on the structured data to be completed, and matching local data is extracted from the local cache pool based on the second dictionary data; the local data is indexed using pre-generated index data to obtain target data; and the target data is assigned to the structured data to be completed to complete data completion.
[0145] It should be noted that the first data receiver and the second data receiver can be message middleware such as Kafka, or directly provided in batches through an interface.
[0146] Based on the data completion method disclosed in the above embodiments of the present invention, the structured data to be stored is pre-parsed into first dictionary data, then assigned to a storage object to obtain local data, and the local data is stored in the corresponding local cache pool. Thus, when the structured data to be completed is received, it is parsed into second dictionary data, the matching local data is retrieved from the local cache pool, and the completion is completed using the local data. In this solution, the degree of intrusion code is low and there is no need to rely on a data warehouse, thus achieving the purpose of no operation and maintenance cost, out-of-the-box use, and ensuring high performance.
[0147] Corresponding to the data completion method disclosed in the above embodiments of the present invention, as Figure 6 shown, it is a structural diagram of a data completion device disclosed in an embodiment of the present invention. The device includes: a first receiving unit 601, a first parsing unit 602, a matching unit 603, a first constructing unit 604, a query unit 605, and a completion unit 606.
[0148] The first receiving unit 601 is used to receive the structured data to be completed.
[0149] The first parsing unit 602 is used to parse the structured data to be completed to obtain the field names and field contents of one or more target fields.
[0150] The matching unit 603 is used to query matching routing data from a pre-constructed routing set based on the field names of the target fields; each routing data corresponds to a local cache pool; the local cache pool is pre-created on the local disk using the CQEngine tool and the RocksDB tool; the local cache pool pre-stores local data; the local data is obtained by assigning the first dictionary data containing business data content to a storage object.
[0151] The first construction unit 604 is configured to construct a second dictionary data that does not contain business data content based on the field name and field content of the target field.
[0152] In one embodiment, the first construction unit 604 is specifically configured to:
[0153] Distinguish IP type fields and non-IP type fields from the conditional fields; sort the field names of each non-IP type field, and splice the field contents of each non-IP type field according to the sorting to obtain a string; sort the field names of each IP type field, and put the field contents of each IP type field into an array according to the sorting to obtain a target array; save the string and the target array to different fields respectively to obtain the second dictionary data.
[0154] The query unit 605 is configured to query local data that matches the second dictionary data from the local cache pool corresponding to the routing data.
[0155] In one embodiment, the query unit is specifically configured to:
[0156] Construct a query statement based on the second dictionary data; use the query statement and the CQEngine tool to query local data that matches the second dictionary data from the local cache pool corresponding to the routing data.
[0157] The completion unit 606 is configured to complete the structured data to be completed by using the local data.
[0158] In one embodiment, the completion unit 606 is specifically configured to:
[0159] Obtain index data corresponding to the local data from the pre-constructed index data cache; use the index data to query target data required for the structured data to be completed from the local data; use the target data to complete the structured data to be completed.
[0160] In one embodiment, the apparatus further includes:
[0161] The second receiving unit is configured to receive the structured data to be stored.
[0162] The second parsing unit is configured to parse the structured data to be stored, and distinguish conditional fields carrying matching conditions and business fields carrying business data content.
[0163] The second construction unit is configured to construct a first dictionary data based on the conditional fields and the business fields.
[0164] In one embodiment, the second construction unit is specifically configured to:
[0165] Distinguish the IP type fields and non-IP type fields from the condition fields; sort the field names of each non-IP type field, and splice the field contents of each non-IP type field according to the sort to obtain a string; sort the field names of each IP type field, and put the field contents of each IP type field into an array according to the sort to obtain a target array; save the string, the target array and the service data content carried by the service field into different fields respectively to obtain the first dictionary data.
[0166] The third construction unit is used to construct routing data based on the field names of the condition fields and the field names of the service fields, and add the routing data to the pre-constructed routing set.
[0167] The splicing unit is used to splice the field names of the condition fields and the field names of the service fields to obtain a key value.
[0168] The storage unit is used to determine whether there is a local cache pool corresponding to the key value; if so, assign the first dictionary data to the storage object to obtain local data, and store the local data in the local cache pool; if not, create a local cache pool corresponding to the key value, assign the first dictionary data to the storage object to obtain local data, and store the local data in the local cache pool.
[0169] Based on the data completion device disclosed in the above embodiments of the present invention, the structured data to be stored is pre-parsed into the first dictionary data, then assigned to the storage object to obtain local data, and the local data is stored in the corresponding local cache pool. Thus, when receiving the structured data to be completed, it is parsed into the second dictionary data, the matching local data is retrieved from the local cache pool, and the completion is completed using the local data. In this solution, the degree of intrusion code is low and there is no need to rely on a data warehouse, so as to achieve the purpose of no operation and maintenance cost, out-of-the-box use and high performance guarantee.
[0170] Each embodiment in this specification is described in a progressive manner. The same or similar parts between the embodiments can be referred to each other. The key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The systems and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement without creative work.
[0171] Those skilled in the art may further realize that the units and algorithm steps of each example described in connection with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0172] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather should be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data completion method, characterized in that, The method includes: Receiving the structured data to be completed; Parsing the structured data to be completed to obtain the field names and field contents of one or more target fields; Querying matching routing data from a pre-constructed routing set based on the field names of the target fields; each piece of the routing data corresponds to a local cache pool; the local cache pool is pre-created on the local disk using the CQEngine tool and the RocksDB tool; the local cache pool stores local data in advance; the local data is obtained by assigning the first dictionary data containing business data content to a storage object; Constructing a second dictionary data that does not contain business data content based on the field names and field contents of the target fields; Querying the local data matching the second dictionary data from the local cache pool corresponding to the routing data; Completing the structured data to be completed using the local data.
2. The method according to claim 1, wherein The process of storing the local data into the local cache pool includes: Receiving the structured data to be stored; Parsing the structured data to be stored to distinguish the conditional fields carrying matching conditions and the business fields carrying business data content; Constructing a first dictionary data based on the conditional fields and the business fields; Constructing routing data based on the field names of the conditional fields and the field names of the business fields, and adding the routing data to a pre-constructed routing set; Concatenating the field names of the conditional fields and the field names of the business fields to obtain a key value; Determining whether there is a local cache pool corresponding to the key value; If so, assigning the first dictionary data to a storage object to obtain local data, and storing the local data into the local cache pool; If not, creating a local cache pool corresponding to the key value, assigning the first dictionary data to a storage object to obtain local data, and storing the local data into the local cache pool.
3. The method according to claim 1, wherein The constructing the first dictionary data based on the conditional fields and the business fields includes: Distinguishing IP type fields and non-IP type fields from the conditional fields; Sorting the field names of each non-IP type field, and concatenating the field contents of each non-IP type field in the sorted order to obtain a string; Sorting the field names of each IP type field, and putting the field contents of each IP type field into an array in the sorted order to obtain a target array; Saving the string, the target array, and the business data content carried by the business fields into different fields respectively to obtain the first dictionary data.
4. The method according to claim 1, characterized in that, The constructing the second dictionary data that does not contain business data content based on the field names and field contents of the target fields includes: Distinguishing IP type fields and non-IP type fields from the conditional fields; Sorting the field names of each non-IP type field, and concatenating the field contents of each non-IP type field in the sorted order to obtain a string; Sort the field names of each of the IP type fields, and put the field contents of each of the IP type fields into an array according to the sorting to obtain a target array; Save the string and the target array into different fields respectively to obtain second dictionary data.
5. The method according to claim 1, characterized in that, The querying the local data matching the second dictionary data from the local cache pool corresponding to the routing data includes: Construct a query statement based on the second dictionary data; Use the query statement and the CQEngine tool to query the local data matching the second dictionary data from the local cache pool corresponding to the routing data.
6. The method according to any one of claims 1 to 5, characterized in that The using the local data to complete the structured data to be completed includes: Obtain the index data corresponding to the local data from a pre-constructed index data cache; Use the index data to query target data required for the structured data to be completed from the local data; Use the target data to complete the structured data to be completed.
7. A data completion device, characterized in that, The device includes: A first receiving unit, configured to receive structured data to be completed; A first parsing unit, configured to parse the structured data to be completed to obtain the field names and field contents of one or more target fields; A matching unit, configured to query matching routing data from a pre-constructed routing set based on the field names of the target fields; each of the routing data corresponds to a local cache pool; the local cache pool is pre-created on a local disk by using the CQEngine tool and the RocksDB tool; the local cache pool pre-stores local data; the local data is obtained by assigning first dictionary data including service data content to a storage object; A first constructing unit, configured to construct second dictionary data that does not include service data content based on the field names and field contents of the target fields; A querying unit, configured to query the local data matching the second dictionary data from the local cache pool corresponding to the routing data; A completing unit, configured to use the local data to complete the structured data to be completed.
8. The device according to claim 7, characterized in that, The device further includes: A second receiving unit, configured to receive structured data to be stored; A second parsing unit, configured to parse the structured data to be stored to distinguish conditional fields carrying matching conditions and service fields carrying service data content; A second constructing unit, configured to construct first dictionary data based on the conditional fields and the service fields; A third constructing unit, configured to construct routing data based on the field names of the conditional fields and the field names of the service fields, and add the routing data to a pre-constructed routing set; A splicing unit, configured to splice the field names of the conditional fields and the field names of the service fields to obtain a key value; A storage unit for determining whether there is a local cache pool corresponding to the key value; if so, assigning the first dictionary data to a storage object to obtain local data, and storing the local data in the local cache pool; if not, creating a local cache pool corresponding to the key value, assigning the first dictionary data to a storage object to obtain local data, and storing the local data in the local cache pool.
9. The device according to claim 8, characterized in that, The second construction unit is specifically used for: Distinguish IP type fields and non-IP type fields from the condition fields; Sort the field names of each non-IP type field, and splice the field contents of each non-IP type field according to the sorting to obtain a string; Sort the field names of each IP type field, and put the field contents of each IP type field into an array according to the sorting to obtain a target array; Save the string, the target array, and the service data content carried by the service field to different fields respectively to obtain first dictionary data.
10. The device according to claim 7, characterized in that, The first construction unit is specifically used for: Distinguish IP type fields and non-IP type fields from the condition fields; Sort the field names of each non-IP type field, and splice the field contents of each non-IP type field according to the sorting to obtain a string; Sort the field names of each IP type field, and put the field contents of each IP type field into an array according to the sorting to obtain a target array; Save the string and the target array to different fields respectively to obtain second dictionary data.