A data processing method, a computer device, a storage medium, and a program product

The data processing method improves system performance by categorizing data structures and using ChronicleMap caching to manage data operations, addressing memory pressure and flexibility issues in existing systems, ensuring efficient and reliable data handling.

CN119829632BActive Publication Date: 2025-07-15INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510296474.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-15
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

Existing data processing systems occupy too much memory when processing large amounts of data, lack performance, difficult to scale efficiently, and lack flexibility when connecting to multi-system services, resulting in high development costs and slowing down project progress.

Method used

By identifying the data structure type of the data object as the first structure data or the second structure data, different data cache objects and processing strategies are adopted, and ChronicleMap is used as the cache system to achieve efficient data processing.

Benefits of technology

It improves the performance of the data processing system, solves the problem of memory resource pressure, improves the system's flexibility and expansion capabilities, and reduces development costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119829632B_ABST
    Figure CN119829632B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method, a computer device, a storage medium, and a program product, relating to the field of computer technologies, including determining the data structure type of a data object according to a target field of the data object, where the data structure type includes first structured data and second structured data, the first structured data includes data definition language, and the second structured data includes the first structured data and data content, obtaining corresponding data processing strategies according to different data structure types, and processing data of different data structure types, which solves the problems of poor flexibility when docking system tasks and the pressure on memory resources caused by obtaining a large amount of data, and achieves the technical effect of improving the performance of the data processing system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a data processing method, a computer device, a storage medium, and a program product. Background Art

[0002] As a cutting-edge technology in the field of artificial intelligence, data processing has made remarkable progress in recent years. In the prior art, during the data preparation process of a data processing system, obtaining a large amount of data will occupy too much memory, causing great pressure on the system's memory resources. Moreover, traditional data processing systems rely on databases to load data, which not only has poor performance but also incurs high development costs and delays the overall project progress in the case of complex business processes. As a result, when docking with multiple system services, they lack flexibility and are difficult to expand efficiently. Therefore, there is an urgent need for a data processing method that can improve the performance of data processing systems. Summary of the Invention

[0003] This application provides a data processing method, a computer device, a storage medium, and a program product to at least solve the problem of system performance degradation caused by too many inference requests being processed simultaneously in the related art.

[0004] This application provides a data processing method, including:

[0005] Obtain the data corresponding to the target field in the data object, and determine the data structure type of the data object according to the data corresponding to the target field. The data structure type includes first structure data and second structure data. The first structure data includes data definition language, and the second structure data includes the first structure data and data content;

[0006] In response to the data structure type being the first structure data, obtain the first data cache object corresponding to the first structure data and the first data processing strategy, and process the first structure data based on the cache system parameters and the first data processing strategy, where the first data cache object includes cache system parameters;

[0007] In response to the data structure type being the second structure data, obtain the second data processing strategy and the first data cache object corresponding to the first structure data in the second structure data, and process the second structure data based on the cache system parameters of the first data cache object and the second data processing strategy.

[0008] This application also provides a computer device, including: a memory for storing a computer program; a processor for implementing the steps of the data processing method in the following embodiments when executing the computer program.

[0009] Obtain the data corresponding to the target field in the data object, and determine the data structure type of the data object according to the data corresponding to the target field. The data structure type includes first structure data and second structure data. The first structure data includes data definition language, and the second structure data includes the first structure data and data content;

[0010] In response to the data structure type being the first structure data, obtain the first data cache object corresponding to the first structure data and the first data processing strategy, and process the first structure data based on the cache system parameters and the first data processing strategy. Among them, the first data cache object includes cache system parameters;

[0011] In response to the data structure type being the second structure data, obtain the second data processing strategy and the first data cache object corresponding to the first structure data in the second structure data, and process the second structure data based on the cache system parameters of the first data cache object and the second data processing strategy.

[0012] This application also provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the data processing method in the following implementation manners are implemented.

[0013] Obtain the data corresponding to the target field in the data object, and determine the data structure type of the data object according to the data corresponding to the target field. The data structure type includes first structure data and second structure data. The first structure data includes data definition language, and the second structure data includes the first structure data and data content;

[0014] In response to the data structure type being the first structure data, obtain the first data cache object corresponding to the first structure data and the first data processing strategy, and process the first structure data based on the cache system parameters and the first data processing strategy. Among them, the first data cache object includes cache system parameters;

[0015] In response to the data structure type being the second structure data, obtain the second data processing strategy and the first data cache object corresponding to the first structure data in the second structure data, and process the second structure data based on the cache system parameters of the first data cache object and the second data processing strategy.

[0016] This application also provides a computer program product, including a computer program. When the computer program is executed by a processor, the steps of the data processing method in the following implementation manners are implemented.

[0017] Obtain the data corresponding to the target field in the data object, and determine the data structure type of the data object according to the data corresponding to the target field. The data structure type includes first structure data and second structure data. The first structure data includes data definition language, and the second structure data includes the first structure data and data content;

[0018] In response to the data structure type being the first structure data, obtain the first data cache object corresponding to the first structure data and the first data processing strategy, and process the first structure data based on the cache system parameters and the first data processing strategy, where the first data cache object includes cache system parameters;

[0019] In response to the data structure type being the second structure data, obtain the second data processing strategy and the first data cache object corresponding to the first structure data in the second structure data, and process the second structure data based on the cache system parameters of the first data cache object and the second data processing strategy.

[0020] Through the data processing method provided by this application, by determining the data structure type of the data object according to the target field of the data object, the data structure type includes first structure data and second structure data, the first structure data includes data definition language, and the second structure data includes the first structure data and data content, obtain the corresponding data processing strategy according to different data structure types, and process the data of different data structure types, which solves the problems of poor flexibility when docking system tasks and the pressure on memory resources caused by obtaining a large amount of data, and achieves the technical effect of improving the performance of the data processing system. Description of the Drawings

[0021] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0022] Figure 1 It is a schematic flow chart of the data processing method provided by the embodiment of the present application;

[0023] Figure 2 It is a schematic diagram of the first structure data provided by the embodiment of the present application;

[0024] Figure 3 It is a schematic diagram of the second structure data provided by the embodiment of the present application;

[0025] Figure 4 It is a schematic diagram of the first data cache object provided by the embodiment of the present application;

[0026] Figure 5Schematic flowchart of the data processing method provided by another embodiment of the present application;

[0027] Figure 6 Schematic structural diagram of the data processing system provided by an embodiment of the present application;

[0028] Figure 7 Block diagram of the structure of the data processing device provided by an embodiment of the present application;

[0029] Figure 8 Internal structure diagram of the computer device provided by an embodiment of the present application. Detailed implementation manners

[0030] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0031] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variation thereof are intended to cover a non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects and are not used to describe a specific order or sequence.

[0032] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific implementation manners.

[0033] As Figure 1 shown, an embodiment of the present application provides a data processing method, which specifically includes the following steps:

[0034] Step 101: Obtain the data corresponding to the target field in the data object, and determine the data structure type of the data object according to the data corresponding to the target field. The data structure type includes first structure data and second structure data. The first structure data includes data definition language, and the second structure data includes the first structure data and data content.

[0035] Specifically, before obtaining the data corresponding to the target field in the data object, it further includes: receiving the data object and determining whether the data object meets the preset structural requirements. Receiving the data object and determining whether the data object meets the preset structural requirements includes: obtaining the metadata information transmitted by the customer system; searching for the data object in the customer system according to the metadata information; in response to receiving the data object transmitted by the customer system, determining whether the structure of the data object is consistent with the preset data object structure; if it is consistent, it is considered that the data object whose structure is consistent with the preset data object structure meets the preset structural requirements.

[0036] This solution involves a customer system, a target database, and a data processing system set in this application formed by a message middleware and a ChronicleMap.

[0037] Among them, the data processing system caches the data of the target database. The customer system will send metadata information to the data processing system, and the data processing system searches for the data object according to the metadata information sent by the customer system. Here, the data object can be a changed data object, and the changed data object corresponds to the changed data information of the target database relative to the data of the target database cached by the data processing system. Here, the changed data information includes the field names and field contents included in the changed data.

[0038] The data processing system in this application defines a unified data structure. When the data object received from the customer system meets the unified data structure defined by the data processing system, it is considered that the data object is a data object that meets the preset requirements.

[0039] Here, the unified data structure can be a data structure in which the data object includes a set of multiple field names and multiple field contents.

[0040] If the received data object does not meet the requirements of the preset structure, the data object will not be processed and the reception of the data object will be rejected.

[0041] If the received data object meets the requirements of the preset structure, the data corresponding to the target field in the data object will be obtained.

[0042] Since a unified data structure is defined in this application, the data objects received by the data processing system all include multiple field names and multiple field contents, where the field contents can be description information of the changed data. Among them, the multiple field names include a data type field (the target field here). According to the field content corresponding to the data type field, the data structure type of the data object can be determined. In this application, the data structure type of the data object is set to include first structure data and second structure data. The first structure data includes data definition language, and the second structure data includes data definition language and data content. The data definition language (DDL) is an important part of the SQL language and is mainly responsible for data structure definition and database object definition. In this application, the first structure data can be expressed as DDL data.

[0043] In this application, different structure data types contain different data information and correspond to different data processing strategies, which can meet the requirements of more customer systems.

[0044] Specifically, in this application, the first structure data is set to include a first field set and the data corresponding to the first field set, and the second structure data includes a second field set and the data corresponding to the second field set. The method further includes: determining the first structure data and determining the second structure data. Determining the first structure data and determining the second structure data includes: obtaining the first field set, and determining the first structure data based on the first field set; the first field set includes data source, sub-data type group, main data type group, metadata ID, data operation type, and data type field; obtaining the second field set, and determining the second structure data based on the second field set; the second field set includes data source, sub-data type group, main data type group, metadata ID, data ID, data content, data operation type, and data type field.

[0045] Specifically, the data structure of the first structure data can be as Figure 2 shown: The first structure data (the illustrated DDL data object) is composed of the first field set and the business data corresponding to the first field set. The first field set includes source (data source, illustrated data source type), groupID (sub-data type group, illustrated data type group), parentID (main data type group, illustrated data type hierarchy relationship), metadata ID (metadata ID, illustrated metadata id, used to find metadata objects in the metadata pool), optionType (operation type, illustrated ADD, DELETE, RELOAD, CLEAR operations above), and data Type (data type, not shown in the figure), etc.

[0046] The data structure of the second structure data can be as Figure 3As shown: The second structure data (graphical data object) is composed of a second field set and the business data corresponding to the second field set. The second field set includes source (data source, graphical data source type), groupID (sub-data type group, graphical data type group), parentID (main data type group, graphical data type hierarchical relationship), metadata ID (metadata ID, graphical metadata id, used to find metadata objects in the metadata pool), data ID (graphical data id), data (graphical data content, mainly MAP type data, used to store metadata information), optionType (operation type, ADD, DELETE, RELOAD operations on the graph), and data Type (data type, not shown in the figure), etc.

[0047] Among them, source (data source) corresponds to its business type, and its function is to distinguish data from different business systems. For example, for the data in a certain data table in a certain business system, it can be marked in the form of "XXX (representing the business system) - XXX data table (representing the data table)". As long as the sources of each business system are different, they can be distinguished.

[0048] groupID (sub-data type group) is used to represent data groups, aiming to distinguish different data groups. Exemplarily, when receiving batches of data from different sources, data with the same data source can be grouped together. For example, the observation list data from the MySQL database can be grouped as one group of data; the transaction logs from the MySQL database can also be grouped as one group of data.

[0049] parentID (main data type group) is the primary ID of groupID, which enables different groups to be recursively nested to adapt to the situation where a business system has multiple master - sub - type data groups.

[0050] metadata ID (metadata ID) is used to obtain metadata information by virtue of the ID, which can describe the number of fields and field types in the data set, providing a functional basis for encapsulation and caching. Metadata information is data about data, mainly information describing data properties (properties), used to support functions such as indicating storage locations, historical data, resource search, and file records. Metadata is a kind of electronic directory. To achieve the purpose of compiling the directory, it is necessary to describe and collect the content or characteristics of the data, and then achieve the purpose of assisting data retrieval.

[0051] The optionType (operation type) refers to the operation on the cache construction information. Among them, the add operation and the reload operation have the same effect, both of which are used to load data.

[0052] The data Type is used to distinguish data types and complete different business operations according to different data types. For example, the data type part of the first structure data can be expressed as "data Type=DDL". In this example, it means that the target data type of the first structure data is DDL data. The data type part of the second structure data can be expressed as "data Type=data". It means that the target data type of the second structure data is data.

[0053] The data field is of Map type. The MAP type field is a data structure that stores key-value pairs, where the key and value can be of any type. It provides a flexible way to store and query structured data. Among them, the key is the field name and the value is the field content. When the first structure data does not reach the cache system, or the first structure data reaches the cache system after the second structure data reaches the cache system, at this time, when the cache information cannot be found using the "three elements", the field information is obtained through the metadata ID field, and a default first default structure data is created. At the same time, a cache Map and a cache file are created based on this default structure data for storing data. This belongs to a fault tolerance mechanism design. Here, the cache system is the system that caches the first structure data and the second structure data, and specifically can be a cache system formed based on ChronicleMap.

[0054] ChronicleMap is a high-performance Java container library based on memory-mapped files (mMap), designed for high-performance and low-latency application scenarios. It supports large data volumes and high-concurrency access, and is especially suitable for scenarios that require fast read and write operations. ChronicleMap supports persisting data to disk, and even if the process terminates, the data will not be lost. And it is designed considering multi-process access scenarios and is suitable for use in a multi-process environment.

[0055] The field names in data correspond to the field names in the metadata. The fields in data can only be fewer than those in the metadata. VALUE is the real value of the changed data corresponding to the data object. The metadata only describes the field names and contents. That is to say, the metadata is the descriptive information of the changed data.

[0056] As can be seen from the above description, the difference between the first field set and the second field set of the first structure data and the second structure data is that the second field set sets data ID (data ID) and data (data content) fields on the basis of the first field set.

[0057] The first structure data and the second structure data are set here so that they can still be used and still lead to data in the absence of specific data. That is, the first structure data is data description information, which can inform the customer system what data can be obtained from here, but it does not contain the content of the data information obtained by the customer system. The second structure data is the data description information and the content of the data information. Through the second structure data, the address of the specific data information required by the customer system can be obtained, and the real cached data content can be obtained based on this address.

[0058] Step 102: In response to the data structure type being the first structure data, obtain the first data cache object and the first data processing policy corresponding to the first structure data, and process the first structure data based on the cache system parameters and the first data processing policy, where the first data cache object includes the cache system parameters.

[0059] In this application, the first structure data will also be converted into a data cache object to obtain the first data cache object. The specific structure of the first data cache object can be as Figure 4 shown:

[0060] The first data cache object includes multiple fields such as source (data source), groupID (sub-data type group), parentID (main data type group), metadata ID (metadata ID), and ChronicleMap (cache system parameters), etc.

[0061] Among them, source (data source) corresponds to the business type, and its function is to distinguish data from different business systems.

[0062] groupID (sub-data type group) is used to represent the data group, and its purpose is to distinguish different data groups.

[0063] parentID (main data type group) is the main-level ID of groupID, which enables different groups to be recursively nested to adapt to the situation where a business system has multiple main and sub-level type data groups.

[0064] metadata ID (metadata ID) is used to obtain metadata information by virtue of the ID, which can describe the number of fields and the field types of the data set, and provides a functional basis for encapsulating the cache.

[0065] ChronicleMap (cache system parameters) are parameters crucial for creating a cache system, specifically including the maxBloatFactor parameter, representing the maximum expansion factor during an abnormal burst; the averageKeySize parameter, representing the average string size of each field name of the metadata sent by the client system; the averageValueSize parameter, representing the average string length of the description content of the metadata sent by the client system; the entries parameter, indicating the size of the storage space capacity of the cache system; and the createPersistedToPath parameter, representing the storage path of the first structured data.

[0066] When storing the first structured data for the first time, the storage path is usually the common parameter location specified by different projects. After the storage path is determined, it can be concatenated with the name of the cache file to complete the definition of the cache file path. The cache file name is composed of the concatenation of the strings of the three elements: source (data source), groupID (sub-data type group), and parentID (main data type group) corresponding to the cache file.

[0067] The cache pool for the first structured data, that is, the storage space for the first structured data, is established when the cache system starts. It is an independent ChronicleMap object specifically used to store DDL cache data. Since the cache system locally adopts a singleton system design, it ensures that there is only one DDL cache Map, and there is also only one corresponding local file. All data cache operations obtain data from the DDL cache Map, and then construct the corresponding ChronicleMap object. Each set of three elements corresponds to a DDL cache object, which is stored in the DDL cache Map, as well as a data cache ChronicleMap named after the three elements and the corresponding local file.

[0068] Please refer to Figure 5, obtain the cache system parameters in the first data cache object, where the cache system parameters include the first data cache address; obtain the data operation type corresponding to the first structured data, and the data operation type corresponding to the first structured data includes a first data clearing operation (clear), a first data deletion operation (delete), and a first data addition operation (add); if the data operation type corresponding to the first structured data is the first data clearing operation, then delete the cached data at the first data cache address; if the data operation type corresponding to the first structured data is the first data deletion operation, then delete the metadata in the first data cache object and delete the cached data at the first data cache address; if the data operation type corresponding to the first structured data is the first data addition operation, then obtain the cached data corresponding to the first structured data and add the cached data corresponding to the first structured data to the cache system based on the first data cache address.

[0069] Step 103: In response to the data structure type being the second structured data, obtain the second data processing policy and the first data cache object corresponding to the first structured data in the second structured data, and process the second structured data based on the cache system parameters of the first data cache object and the second data processing policy.

[0070] Please continue to refer to Figure 5 , obtain the cache system parameters in the first data cache object, where the cache system parameters include the first data cache address; determine the second data cache address based on the first data cache address; obtain the data operation type corresponding to the second structured data, and the data operation type corresponding to the second structured data includes a second data addition operation and a second data deletion operation; if the data operation type corresponding to the second structured data is the second data deletion operation, then obtain the second data cache address and delete the cached data at the second data cache address; if the data operation type corresponding to the second structured data is the second data addition operation, then obtain the second cached data corresponding to the second structured data and add the second cached data corresponding to the second structured data to the cache system based on the second data cache address.

[0071] In this application, after determining whether it is DDL data (the first structured data) or cached data (the second structured data), specific operations can be obtained according to the operation type field in the corresponding structured data. The figure shows three operation types of cleaning, deletion, and addition corresponding to DDL data, and two operation types of deletion and addition corresponding to cached data.

[0072] In this application, cache objects are respectively established for DDL data objects (first - structured data) to obtain DDL data cache objects (first - data cache objects), and data objects are established for cached data (second - structured data) to obtain second - data cache objects. Since the DDL data cache object includes the parameters of ChronicleMap, and this parameter includes the cache address, thus, the DDL data cache object can be added or deleted at the cache address according to the corresponding operation type.

[0073] There are two operations corresponding to the cached data, namely addition and deletion. If it is an addition, obtain the ChronicleMap object from the DDL data cache object. This object is the cache address of the DDL cached data. Add the DDL cached data cache address to the address - lookup instruction to obtain the address of the second - data cache object. Perform the deletion and addition of the second - cached data at this cache address.

[0074] In one embodiment, obtain the target cache - data size of the second - cached data and the remaining size of the target storage space of the cache system; compare the remaining size of the target storage space with the target cache - data size; if the remaining size of the target storage space is less than or equal to the target cache - data size, directly store the second - cached data into the target storage space; if the remaining size of the target storage space is greater than the target cache - data size, expand the target storage space based on the expansion algorithm.

[0075] Since the second - cached data includes data content, when adding the second - cached data, there will be a process of judging whether the added data will overflow, including the expansion algorithm and establishing a new cache address.

[0076] Specifically, the preset range of the first - storage - space size, the preset range of the second - storage - space size, and the preset range of the third - storage - space size can be obtained; compare the target cache - data size with the preset range of the first - storage - space size, the preset range of the second - storage - space size, and the preset range of the third - storage - space size respectively; if the target cache - data size is within the range of the first - storage - space size, expand the target storage space based on the first expansion formula; if the target cache - data size is within the range of the second - storage - space size, expand the target storage space based on the second expansion formula; if the target cache - data size is within the range of the third - storage - space size, expand the target storage space based on the third expansion formula;

[0077] Among them, the first expansion formula is as follows:

[0078] ;

[0079] The second expansion formula is as follows:

[0080] ;

[0081] The third expansion formula is as follows:

[0082] ;

[0083] Among them, C MAX is the target storage space size after expansion, B is the target storage space size before expansion, and D is a constant, which is set according to actual needs.

[0084] In this application, an expansion threshold is also set. The expansion threshold can be obtained. The expansion threshold is the maximum value of the target storage space size. When the target cache data size is greater than the expansion threshold, a new storage space is created, and a timestamp is added to the new storage space. The second cache data is imported into the new storage space.

[0085] When the cache data capacity overflows, an exception is obtained through the try method. The DDL cache object is obtained by looking up the content of the source (data source), groupID (sub-data type group), and parentID (main data type group) fields. A new ChronicleMap is rebuilt and the original data is copied to the new Map. After modifying the path name, the original object and file are deleted. The expansion algorithm is as follows: The size of the file storing the current data needs to be read each time for expansion.

[0086] If the current file size is below 128 kb, follow the expansion algorithm: max = max + (max * 0.5), where max is the current maximum storage quantity.

[0087] If the current file size is between 128 kb and 1 mb, max = max + (max * 0.25), where max is the current maximum storage quantity.

[0088] If the current file size is between 1 mb and 512 mb, max = max + 2000. max is the current maximum storage quantity. At this time, expansion will make the file too large, wasting disk space. At the same time, there will not be too much caching of this type of data, and frequent expansion operations will not be triggered.

[0089] If the current file size exceeds 512 mb, data reception will be refused. At this time, the data volume is generally in the tens of millions or even hundreds of millions, which puts great pressure on disk expansion. At the same time, it is no longer for caching purposes but for persistent data purposes, which is not within the scope of support of this system.

[0090] In one embodiment, the time when the first structure data arrives at the cache system is obtained, and it is judged whether the time when the first structure data arrives at the cache system meets the preset time requirement; if not, the second structure data is made to send metadata, a first default structure data is created based on the second structure sending metadata, and a default cache address and a default cache file are created based on the first default structure data.

[0091] In practical applications, when receiving the first structure data in response, it is judged whether the first default structure data corresponding to the first structure data is included in the cache system; if so, the first default structure data is verified based on the first structure data, and the difference information between the first structure data and the first default structure data is obtained; the first default structure information is corrected based on the difference information.

[0092] In this application, the received data objects all have a DDL object (the first structure data) and a cache object (the second structure data). Usually, the DDL object is first converted into a DDL cache object and stored in the cache system, and then the data object is stored in the cache system according to the cache address of this DDL cache object. In actual operation, when the DDL object of a data is not stored in the cache system, or the DDL data object of this data has not been stored yet, and the data object is stored in the cache system. At this time, the metadata is obtained from the data object, and a default DDL data object is created according to this metadata. After receiving the DDL data object of this data, the received DDL data object is compared with the DDL data object created according to the metadata, and the default DDL data object is corrected.

[0093] This is set because in actual operation, there is a situation where the first structure data cannot be received. After the first data arrives at the cache system, the absence of the first structure data will cause data loss. Since the first structure data is not received, the system will forcefully send a metadata ID. At this time, the metadata ID is obtained based on the data object, and the metadata ID includes metadata information. A first default structure data is created based on this metadata information. After the system receives the first structure data, the received first structure data is compared with the first default data object created according to the metadata, and the default DDL data object is corrected, which is equivalent to a fault tolerance mechanism to reduce the trial-and-error cost.

[0094] In one embodiment, in response to receiving a data object sent by a client system, the data object is stored in the message queue of the message middleware; the data object in the message queue is sent to the cache system for processing according to the level of the message queue.

[0095] The number of processor cores of a data processing system can be obtained, and the number of consumer threads is determined based on the number of processor cores of the data processing system; wherein, determining the number of consumer threads based on the number of processor cores of the data processing system includes: obtaining a first threshold range of the number of processor cores, a second threshold range of the number of processor cores, and a third threshold range of the number of processor cores; comparing the number of processor cores with the first threshold range of the number of processor cores, the second threshold range of the number of processor cores, and the third threshold range of the number of processor cores respectively; if the number of processor cores is within the first threshold range of the number of processor cores, determining the number of consumer threads based on the first calculation formula; if the number of processor cores is within the second threshold range of the number of processor cores, determining the number of consumer threads based on the second calculation formula; if the number of processor cores is within the third threshold range of the number of processor cores, determining the number of consumer threads based on the third calculation formula.

[0096] Wherein, the first calculation formula is as follows:

[0097] ;

[0098] The second calculation formula is as follows:

[0099] ;

[0100] The third calculation formula is as follows:

[0101] ;

[0102] Wherein, K is the number of consumer threads, and Y is the number of processor cores.

[0103] Multiple Kafka clients of the present application can be set up, which are uniformly managed by a thread pool and use only one topic name. All producers only need to send data to the same topic according to the specified data structure. The thread pool creates consumer threads based on the current number of CPU cores, monitors each thread, and outputs the log of the received data. Try to improve the data consumption ability and the rate of parsing local files by virtue of the improvement of the hardware configuration, which is particularly crucial in the scenarios of virtual machine deployment and cloud environment deployment, because such environments will frequently change the configuration as needed. In this way, while improving the configuration, the caching system can also provide corresponding performance accordingly.

[0104] Specifically, the algorithm for creating consumer thread data is as follows:

[0105] To ensure system performance and at the same time ensure the consumption performance of Kafka, it is necessary to set the number of partitions of Kafka topics according to the number of CPUs, and at the same time allocate the same number of consumer threads as the number of partitions.

[0106] k represents the current number of threads, and cpu represents the current number of cores.

[0107] If the current number of CPU cores is less than 32, the formula for the number of threads is: k = cpu / 8 + 1.

[0108] If the current CPU is greater than 32 but less than 128, the formula for the number of threads is: k = cpu / 32 + 4.

[0109] If the current CPU has more than 128 cores, which belongs to a supercomputer, the formula for the number of threads is: k = cpu / 128 + 8. Because too many consumer threads will not improve the consumption efficiency and will also slow down the efficiency of the CPU in executing the main business.

[0110] Let k represent the current number of threads and cpu represent the current number of cores, which can be expressed as the following piecewise function:

[0111]

[0112] In a feasible implementation, it is possible to determine whether the data object written into the message queue is complete; among them, when the data object written into the message queue is incomplete, the data compensation mechanism is triggered.

[0113] When using MQ systems such as RabbitMQ and RocketMQ to implement asynchronous processing for the message queue, although messages (events) can be stored on disk and the message data will not be lost even if the MQ has problems, message loss may occur in all links such as message sending, transmission, and processing in the asynchronous process. In addition, no MQ middleware can ensure 100% availability, and it is necessary to consider how the asynchronous process continues when it is unavailable. Therefore, for the asynchronous processing process, it is necessary to consider compensation or establish a primary and standby dual-active process. Based on this, the present invention sets up a data compensation mechanism, and when the data object written into the message queue is incomplete, the data compensation mechanism is triggered to compensate the data object written into the message queue.

[0114] Optionally, the data object that has not been written into the message queue can be compared with the data object after it is written into the message queue. When the comparison shows a difference between the two, it is considered that the data object written into the message queue is incomplete. And the system will feedback the specific difference content between the data object that has not been written into the message queue and the data object after it is written into the message queue to the user. It can be through an alarm or other means to remind the user that the data object written into the message queue is incomplete.

[0115] Exemplarily, when it is confirmed that the data object written into the message queue is incomplete, a data compensation mechanism is triggered to generate a data compensation instruction. In response to receiving the data compensation instruction, the difference information between the data object written into the message queue and the data object after it corresponding to the written message queue can be captured, and the scenario parameters of this difference information are matched according to the feedback information of the data object not written into the message queue itself. The difference information is compensated in a timely manner to a certain extent based on the information itself. By setting up the data compensation mechanism, the reliability of message transmission is improved, ensuring the correct transmission of messages. In addition, abnormal messages (data objects after being written into the message queue) can be stored, generally in the message middleware cache queue. According to needs, the messages can be retrieved and resent to avoid the loss of messages after an exception occurs, so as to improve the reliability and robustness of subsequent data transmission.

[0116] As Figure 6 shown, the present application provides a data processing system. The overall design of the data processing system is divided into four layers, including a data channel layer, a data conversion and encapsulation layer, a data cache layer, and a data storage layer.

[0117] Among them, the data channel layer includes a message middleware component (illustrated as a kafka channel), an interface channel, a unified communication data structure, and a thread pool management.

[0118] The data channel layer of the present application undertakes the task of receiving data, mainly relying on kafka to implement. Kafka is an open-source stream processing platform developed by the Apache Software Foundation, written in Scala and Java. Kafka is a high-throughput distributed publish-subscribe message system that can process all action stream data of consumers in a website. This kind of action (webpage browsing, searching, and other user actions) is a key factor in many social functions on modern networks. These data are usually solved by processing logs and log aggregation due to throughput requirements. For log data and offline analysis systems like Hadoop, but with the limitation of real-time processing requirements, this is a feasible solution. The purpose of Kafka is to unify online and offline message processing through Hadoop's parallel loading mechanism and also to provide real-time messages through a cluster.

[0119] Since the channel layer completely relies on kafka, in the present application, when synchronizing data in kafka, a timed operation of pushing full-volume data is set. To avoid the situation that after the kafka service starts, the cache system fails to start, and it takes too long, and the data cannot be pulled before the kafka data expires, resulting in data loss and missing cache data and being unusable.

[0120] The data channel layer is responsible for receiving data and mainly relies on Kafka to achieve this. It converts the received JSON-formatted data into a specified data structure. A unified data structure is defined in the data channel layer. When receiving data, the data at the structure layer directly uses this data structure and completes the reception operation through the API interface. The role of Kafka here is to convert the received JSON data into the aforementioned unified data structure.

[0121] The main task of the data conversion and encapsulation layer is to convert the received data into a data type for storage. After receiving the DDL data, it needs to be transformed into a DDL data cache object, which is used to describe the construction parameters of the ChronicleMap cache pool.

[0122] The data cache layer mainly undertakes three tasks: First, it calls the storage layer interface according to the DDL cache object, creates a ChronicleMap object and manages it; second, after receiving the cached data, it accurately finds the already created ChronicleMap object and writes the data into it; third, it can quickly locate, search for the cache and obtain the cache content.

[0123] The data storage layer directly uses the ChronicleMap object to implement its functions and achieves the purpose of local file caching by virtue of its ability to write files.

[0124] In this application, a design solution for localizing data caching is achieved by utilizing the file-writing characteristics and high-performance capabilities of ChronicleMap. And an optimized design is made for the use problems of ChronicleMap, such as the handling of overflow data, the handling of data files, the management design of multiple data caches; the design of a unified data reception method, making use of the distributed characteristics of Kafka, only by connecting to Kafka can the data reception and caching be completed, improving scalability and decoupling the data generation end and the data usage end.

[0125] An embodiment of this application provides a data processing device, and the data processing device is specifically as Figure 7 shown. The data processing device includes: a receiving module 20 and a processing module 21.

[0126] The receiving module 20 is used to obtain the data corresponding to the target field in the data object, determine the data structure type of the data object according to the data corresponding to the target field. The data structure type includes the first structure data and the second structure data. The first structure data includes the data definition language, and the second structure data includes the first structure data and the data content.

[0127] A processing module 21, configured to, when responding to the data structure type being the first structured data, obtain a first data cache object corresponding to the first structured data and a first data processing policy, and process the first structured data based on the cache system parameters and the first data processing policy, wherein the first data cache object includes the cache system parameters;

[0128] When responding to the data structure type being the second structured data, obtain a second data processing policy and the first data cache object corresponding to the first structured data in the second structured data, and process the second structured data based on the cache system parameters of the first data cache object and the second data processing policy.

[0129] For the description of the features in the embodiments corresponding to the data processing device, reference may be made to the relevant description in the embodiments corresponding to the data processing method, which will not be elaborated here one by one.

[0130] An embodiment of the present application further provides a computer device, as Figure 8 shown, the computer device includes a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any of the above embodiments of the data processing method.

[0131] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, wherein the computer program is configured to execute the steps in any of the above embodiments of the data processing method when running.

[0132] In an exemplary embodiment, the above computer-readable storage medium may include, but is not limited to: various media such as a USB flash drive, a read-only memory (ROM for short), a random access memory (RAM for short), a mobile hard disk, a magnetic disk, or an optical disc that can store a computer program.

[0133] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to their functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present application.

[0134] The above has introduced in detail a data processing method provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A data processing method, characterized in that Including: Obtain the data corresponding to the target field in the data object, and determine the data structure type of the data object according to the data corresponding to the target field. The data structure type includes first structure data and second structure data. The first structure data includes data definition language, and the second structure data includes the first structure data and data content; In response to the data structure type being the first structure data, obtain the first data cache object corresponding to the first structure data and the first data processing strategy, and process the first structure data based on the cache system parameters and the first data processing strategy, where the first data cache object includes the cache system parameters; In response to the data structure type being the second structure data, obtain the second data processing strategy and the first data cache object corresponding to the first structure data in the second structure data, and process the second structure data based on the cache system parameters of the first data cache object and the second data processing strategy; In response to the first structure data not reaching the cache system and the second structure data reaching the cache system, or the second structure data reaching the cache system earlier than the first structure data, enable the second structure data to send metadata, create the first default structure data based on the metadata sent by the second structure data, and create a default cache address and a default cache file based on the first default structure data; In response to receiving the first structure data and the first default structure data corresponding to the first structure data being included in the cache system; compare the first structure data and the first default structure data, obtain the difference information between the first structure data and the first default structure data; and correct the first default structure information based on the difference information.

2. The data processing method according to claim 1, wherein Before obtaining the data corresponding to the target field in the data object, it further includes: receiving the data object, and determining whether the data object meets the preset structure requirements. The receiving the data object and determining whether the data object meets the preset structure requirements includes: obtaining the metadata information transmitted by the client system; Search for the data object in the client system according to the metadata information; In response to receiving the data object transmitted by the client system, determine whether the data object structure is consistent with the preset data object structure; If it is consistent, it is considered that the data object whose data object structure is consistent with the preset data object structure meets the preset structure requirements.

3. The data processing method according to claim 1, wherein The first structure data includes a first field set and the data corresponding to the first field set, the second structure data includes a second field set and the data corresponding to the second field set, and the second field set includes the first field set. The method further includes: determining the first structure data and determining the second structure data: The determining the first structure data and determining the second structure data includes: Obtain the first field set, and determine the first structure data based on the first field set; The first field set includes data source, sub-data type group, main data type group, metadata ID, data operation type, and data type field; Obtain the second field set, and determine the second structure data based on the second field set; The second field set includes data source, sub-data type group, main data type group, metadata ID, data ID, data content, data operation type, and data type field.

4. The data processing method according to claim 1, wherein Obtain a first data cache object corresponding to the first structured data and a first data processing policy, and process the first structured data based on the cache system parameters and the first data processing policy, where the first data cache object includes the cache system parameters including: Obtain the cache system parameters in the first data cache object, and the cache system parameters include a first data cache address; Obtain the data operation type corresponding to the first structured data, and the data operation type corresponding to the first structured data includes a first data clearing operation, a first data deletion operation, and a first data addition operation; If the data operation type corresponding to the first structured data is a first data clearing operation, delete the cached data at the first data cache address; If the data operation type corresponding to the first structured data is a first data deletion operation, delete the metadata in the first data cache object and delete the cached data at the first data cache address; If the data operation type corresponding to the first structured data is a first data addition operation, obtain the cached data corresponding to the first structured data, and add the cached data corresponding to the first structured data to the cache system based on the first data cache address.

5. The data processing method according to claim 1, wherein The process of obtaining a second data processing policy and a first data cache object corresponding to the first structured data in the second structured data, and processing the second structured data based on the cache system parameters of the first data cache object and the second data processing policy includes: Obtain the cache system parameters in the first data cache object, and the cache system parameters include a first data cache address; determine a second data cache address based on the first data cache address; Obtain the data operation type corresponding to the second structured data, and the data operation type corresponding to the second structured data includes a second data addition operation and a second data deletion operation; If the data operation type corresponding to the second structured data is a second data deletion operation, obtain the second data cache address and delete the cached data in the second data cache address; If the data operation type corresponding to the second structured data is a second data addition operation, obtain the second cached data corresponding to the second structured data, and add the second cached data corresponding to the second structured data to the cache system based on the second data cache address.

6. The data processing method according to claim 5, wherein The process of adding the second cached data corresponding to the second structured data to the cache system based on the data cache address includes: Obtain the target cached data size of the second cached data and the remaining size of the target storage space of the cache system; Compare the remaining size of the target storage space with the target cached data size; If the remaining size of the target storage space is less than or equal to the target cached data size, directly store the second cached data in the target storage space; If the remaining size of the target storage space is greater than the target cached data size, expand the target storage space based on an expansion algorithm.

7. The data processing method according to claim 6, wherein The process of expanding the target storage space based on the expansion algorithm includes: Obtain the preset range of the size of the first storage space, the preset range of the size of the second storage space, and the preset range of the size of the third storage space; Compare the size of the target cached data with the preset range of the size of the first storage space, the preset range of the size of the second storage space, and the preset range of the size of the third storage space respectively; If the size of the target cached data is within the range of the size of the first storage space, expand the target storage space based on the first expansion formula; If the size of the target cached data is within the range of the size of the second storage space, expand the target storage space based on the second expansion formula; If the size of the target cached data is within the range of the size of the third storage space, expand the target storage space based on the third expansion formula; Among them, the first expansion formula is as follows: The second expansion formula is as follows: The third expansion formula is as follows: C MAX = B + D; Among them, C MAX is the size of the target storage space after expansion, B is the size of the target storage space before expansion, and D is a constant, which is set according to actual needs.

8. The data processing method according to claim 7, wherein The method further includes: Obtain an expansion threshold, where the expansion threshold is the maximum value of the size of the target storage space; When the size of the target cached data is greater than the expansion threshold, create a new storage space and add a timestamp to the new storage space; Import the second cached data into the new storage space.

9. The data processing method according to claim 1, wherein In response to receiving the first structured data and the cache system includes the first default structured data corresponding to the first structured data; Compare the first structured data and the first default structured data, and obtain the difference information between the first structured data and the first default structured data; Based on the difference information, the correction of the first default structured information includes: In response to receiving the first structured data, determine whether the cache system includes the first default structured data corresponding to the first structured data; If included, compare the first structured data and the first default structured data, and obtain the difference information between the first structured data and the first default structured data; Based on the difference information, correct the first default structured information.

10. The data processing method according to claim 2, characterized in that, The method further includes: In response to receiving a data object sent by the client system, store the data object in the message queue of the message middleware; Send the data object in the message queue to the cache system for processing according to the level of the message queue.

11. The data processing method according to claim 1, wherein The method further includes: Obtain the number of processor cores of the data processing system, and determine the number of consumer threads based on the number of processor cores of the data processing system; Among them, determining the number of consumer threads based on the number of processor cores of the data processing system includes: obtaining the first threshold range of the number of processor cores, the second threshold range of the number of processor cores, and the third threshold range of the number of processor cores; Compare the number of processor cores with the first threshold range of the number of processor cores, the second threshold range of the number of processor cores, and the third threshold range of the number of processor cores respectively; If the number of processor cores is within the first threshold range of the number of processor cores, determine the number of consumer threads based on the first calculation formula; If the number of processor cores is within the second threshold range of the number of processor cores, determine the number of consumer threads based on the second calculation formula; If the number of the processor cores is within the range of the third threshold of the number of processor cores, determine the number of consumer threads based on the third calculation formula; Among them, the first calculation formula is as follows: The second calculation formula is as follows: The third calculation formula is as follows: Among them, K is the number of consumer threads, and Y is the number of processor cores.

12. A computer device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the data processing method according to any one of claims 1 to 11 when executing the computer program.

13. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium, wherein the computer program implements the steps of the data processing method according to any one of claims 1 to 11 when executed by a processor.

14. A computer program product, comprising a computer program, characterized in that, The computer program implements the steps of the data processing method according to any one of claims 1 to 11 when executed by a processor.

Citation Information

Patent Citations

  • Data caching method and device, computer equipment and storage medium

    CN115525677A

  • Data synchronization method and system based on DDL monitoring processing

    CN117407463A