Data processing method and device, electronic equipment and storage medium
By identifying and classifying the original data of the income and expenditure detailed system, and storing it in the column database, the problem of low data processing efficiency in the existing technology is solved and efficient data query and modification is achieved.
Patent Information
- Application Number
- CN202510691018.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-09-05
AI Technical Summary
When processing data, the existing revenue and expenditure detailed system needs to frequently exchange data with the big data platform, resulting in low data query efficiency and inability to obtain statistical data in time, and the modification situation cannot be synchronized in time.
By identifying the original data obtained in real time, classifying and statistics, and storing the data in a column database, it is divided into three-layer structures: data addition identification, classification statistics, data storage, and data query modification, reducing dependence on the big data platform.
It improves data processing efficiency and applicability, reduces system complexity and hardware resource consumption, and realizes efficient data query and modification.
Smart Images

Figure CN120596531A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and in particular to a data processing method, device, electronic device and storage medium. Background Art
[0002] With the development of financial digitalization, the income and expenditure detail system is becoming more and more widely used. The income and expenditure detail system aggregates, processes, and categorizes raw data, and can provide multi-dimensional data query and modification functions.
[0003] Existing income and expenditure detail systems obtain raw data from real-time data transmission tools, add tags to the raw data, and write it to a database. They then periodically import incremental tagged data into a big data platform for classification and statistics. However, this approach places heavy demands on the big data platform, requiring frequent data exchange between the income and expenditure detail system and the big data platform. This can prevent the system from obtaining statistical data in a timely manner, resulting in low query efficiency. Furthermore, if the raw data is modified, the changes cannot be synchronized to the statistical data in a timely manner. Summary of the Invention
[0004] The present invention provides a data processing method, device, electronic device and storage medium to realize real-time processing of raw data, thereby improving the processing efficiency of data statistics, query and modification processes.
[0005] In a first aspect, an embodiment of the present invention provides a data processing method, the method comprising:
[0006] Determine the original data and add a data type identifier to the original data based on the pre-set data type;
[0007] Perform classification statistics on the original data after adding data type identifiers to obtain statistical data;
[0008] The raw data and statistical data are stored in a column database, and data query and / or data modification are performed based on the raw data and statistical data.
[0009] In a second aspect, an embodiment of the present invention further provides a data processing device, the device comprising:
[0010] An identification adding module is used to determine the original data and add a data type identification to the original data based on a preset data type;
[0011] The data statistics module is used to classify and count the original data after adding data type identifiers to obtain statistical data;
[0012] The data storage module is used to store the original data and statistical data in a column database, and perform data query and / or data modification based on the original data and statistical data.
[0013] In a third aspect, an embodiment of the present invention further provides an electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the data processing method as described in any one of the embodiments of the present invention is implemented.
[0014] In a fourth aspect, an embodiment of the present invention further provides a storage medium storing computer-executable instructions, wherein the computer-executable instructions, when executed by a computer processor, are used to execute the data processing method as described in any one of the embodiments of the present invention.
[0015] The technical solution of the embodiment of the present invention adds data type identifiers to raw data obtained in real time, classifies and counts the identified raw data to obtain statistical data, stores the raw data and statistical data in a column database, and implements data query and / or data modification based on the raw data and statistical data. This embodiment does not rely on a big data platform and divides data processing into a three-tiered structure: data identification and classification statistics, data storage, and data query and modification. This facilitates horizontal expansion during data processing and improves data processing efficiency and applicability.
[0016] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0018] Figure 1 This is a flow chart of a data processing method provided by the first embodiment of the present invention;
[0019] Figure 2 is a flow chart of another data processing method provided by the second embodiment of the present invention;
[0020] Figure 3 This is a flow chart of another data processing method provided by the third embodiment of the present invention;
[0021] Figure 4 1 is a structural diagram of a data processing device provided by a fourth embodiment of the present invention;
[0022] Figure 5 This is a structural diagram of an electronic device provided in Example 5 of the present invention. DETAILED DESCRIPTION
[0023] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.
[0024] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. In the embodiments of the present application, certain software, components, models and other existing solutions in the industry may be mentioned, and they should be considered as exemplary. Their purpose is merely to illustrate the feasibility of the implementation of the technical solution of the present application, but it does not mean that the applicant has or will necessarily use the solution.
[0025] The acquisition, transmission, storage, use, and processing of data in the technical solution of this application comply with the relevant provisions of national laws and regulations.
[0026] Example 1
[0027] Figure 1 A flowchart of a data processing method is provided for the first embodiment of the present invention. This embodiment is applicable to data processing in an income and expenditure detail system. The method can be executed by a data processing device, which can be implemented in the form of hardware and / or software, and can be configured in a server.
[0028] like Figure 1 As shown, the method includes:
[0029] S110: Determine the original data, and add a data type identifier to the original data based on a preset data type.
[0030] Among them, the original data refers to the original detailed data obtained in real time by the income and expenditure details system from the data transmission platform. The data transmission platform is connected to multiple different application systems and can obtain the income and expenditure details data generated by multiple application systems in real time.
[0031] The data type is a pre-set, enumerable type. For example, the data type may include whether it is included in income and expenditure, revenue, expenditure, etc., and may also include detailed classifications such as living expenses and catering.
[0032] In this embodiment, after obtaining the original data from the data transmission platform, the original data is marked according to the preset data type so that the data can be subsequently classified based on the preset data type to achieve real-time statistics of the original data.
[0033] Furthermore, before adding data type identifiers to the original data, the original data can also be pre-processed, including data deduplication, data filtering, and data merging. Among them, data deduplication removes duplicate data during data transmission, data filtering refers to filtering invalid fields in the original data, and data merging refers to merging original data with the same attributes or with an associated relationship to obtain full data, which is convenient for the subsequent addition of data type identifiers. For example, for data of the type of living payment, the log number can be used as the basis for merging, and the original data from multiple channels can be merged to generate the full amount of original data.
[0034] Data deduplication, data filtering, and data merging avoid data duplication, improve the accuracy of data statistical results, and at the same time, ensure the efficiency of data processing and implementation.
[0035] S120: Perform classification statistics on the original data after adding the data type identifier to obtain statistical data.
[0036] Statistical data is the result of classifying and counting the original data. In the income and expenditure details system, statistical data not only includes the statistical values obtained after classification and counting, but also includes the detailed data corresponding to the statistical dimensions.
[0037] In this embodiment, each original data is classified into different dimensions according to the data type identifier of each original data, and real-time statistics are performed on each dimension.
[0038] In a specific application scenario, the original data corresponding to the "not included in income and expenditure" mark can be directly stored in the column database, and the original data corresponding to the "included in income and expenditure" mark can be further classified and counted. For each original data corresponding to the "included in income and expenditure" mark, statistics can be performed in different dimensions such as income, expenditure, and detailed classification according to its mark. Furthermore, each statistical dimension and each statistical dimension and at least one time dimension can be superimposed. For example, for each original data corresponding to the "included in income and expenditure" mark, the statistical results of the income dimension within a week, a month, a quarter, and a year can be obtained according to its mark, or the statistical results of the living expense dimension in the expenditure dimension can be obtained.
[0039] In this embodiment, the income and expenditure details system can implement data processing through the coordinated operation of a three-tier structure. The three-tier structure can include a data preprocessing module, a data storage module, and a data query and modification module. The data preprocessing module is used to add data type identifiers to the real-time raw data in the data transmission platform, perform classification statistics based on the data type identifiers, and store the resulting statistical data and raw data.
[0040] The technical solution of this embodiment, through the labeling and real-time statistics of the real-time raw data in the data transmission platform, the resulting statistical data becomes the basis for subsequent data query and modification, thereby improving the efficiency of data query and modification. At the same time, the classification statistics of the raw data no longer rely on the big data platform, avoiding frequent interactions between the big data platform and the income and expenditure detailed system, and improving data processing efficiency. The three-tier structure realizes data processing and is widely applicable to cross-system data consumption scenarios. It can effectively expand horizontally to address data processing bottlenecks and improve the execution speed of the income and expenditure detailed system.
[0041] S130: Storing the original data and the statistical data in a column database, and performing data query and / or data modification based on the original data and the statistical data.
[0042] Among them, a column database is a database system that uses columns rather than rows as the basic unit of data storage and management.
[0043] In this embodiment, data is stored based on a column database. The reason is that, firstly, the column database's ability to support massive amounts of data is far greater than that of a traditional relational database. When the amount of data is the same, the overhead of establishing indexes in a traditional relational database will lead to index expansion and high hardware resource consumption. Therefore, in the income and expenditure details system of this embodiment, the following database is more suitable for scenarios with large amounts of data and multiple data sources. Secondly, in the applicable scenarios of the income and expenditure details system, there are more raw data flows, while data queries and modifications are relatively less. In this scenario with more writes and less reads, as long as the configuration, quantity, and partition distribution of the column database are reasonable, its read and write performance can fully support the TPS (Transactions Per Second) of the income and expenditure details system. Thirdly, the column database does not have a fixed table structure and is not limited by the multi-table connection query mechanism of the traditional relational database, and has better scalability.
[0044] In this embodiment, when acquiring real-time raw data from the data transmission platform, data preprocessing is performed using steps S110-S120, and statistical data classified and counted based on pre-set data types is saved. Therefore, when performing data queries, data query efficiency is improved and query response time is reduced. At the same time, since frequent interaction with the big data platform is not required, system complexity is reduced and hardware resource consumption is reduced.
[0045] For data queries based on pre-defined data types and dimensions, you can directly obtain the corresponding data query results based on the statistical data. For example, if the data query request is to query expenditures within a week, the data query results can be directly obtained based on the statistical data of the income dimension within a week and its corresponding raw data.
[0046] For data queries that don't use pre-defined data types or dimensions, or for queries that combine multiple conditions, further data filtering and querying can be performed based on the closest statistical data, thereby improving query speed and reducing wait time. For example, if the data query request is for expenditures from xA-day, xx-month, xx-year to xB-day, xx-month, xx-year, the statistical data with the smallest amount of data covering that time period can be obtained, such as the total expenditure data for x-month, xx-year. Based on this total expenditure data, expenditures from xA-day to xB-day can be further filtered.
[0047] In this embodiment, data modification is implemented based on a column database. Statistical data in the column database can be directly located based on the row key, thereby modifying both the statistical data and the corresponding original data. This avoids inverse processing on the big data platform, reduces the complexity of data modification, and improves the response speed of data modification.
[0048] Furthermore, in addition to data preprocessing, data storage, modification, and querying in this embodiment, batch data exchange between the income and expenditure detail system and the big data platform can be performed via a batch data bus to implement data mining and analysis. For example, the income and expenditure detail system can send end-of-day raw data and statistical data to the big data platform. The big data platform then performs data mining and analysis and feeds the results back to the income and expenditure detail system, enabling functions such as income and expenditure forecasting and investment and marketing product recommendations.
[0049] The technical solution of the embodiment of the present invention adds data type identifiers to raw data obtained in real time, classifies and counts the identified raw data to obtain statistical data, stores the raw data and statistical data in a column database, and implements data query and / or data modification based on the raw data and statistical data. This embodiment does not rely on a big data platform and divides data processing into a three-tiered structure: data identification and classification statistics, data storage, and data query and modification. This facilitates horizontal expansion during data processing and improves data processing efficiency and applicability.
[0050] Example 2
[0051] Figure 2 This is a flow chart of a data processing method provided in the second embodiment of the present invention. Based on the above embodiments, this embodiment of the present invention further specifies the process of data query.
[0052] like Figure 2 As shown, the method includes:
[0053] S210: Determine the original data, and add a data type identifier to the original data based on a preset data type.
[0054] S220: Perform classification statistics on the original data after adding the data type identifier to obtain statistical data.
[0055] S230: Store the original data and statistical data in a column database.
[0056] The above embodiments have specifically described the process of data preprocessing, including real-time acquisition of raw data, adding data type identifiers to the raw data, performing classification statistics based on the data type identifiers, and storing the raw data and statistical data. This embodiment will not be repeated here.
[0057] S240: Determine whether the data query request matches a preset data type. If so, execute S250; otherwise, execute S260.
[0058] The data query request matches the pre-set data type. Since the data type is a pre-set enumerable data type, the data query request is an enumerable filtering condition. For example, data types include income, expenditure, and different time dimensions such as within three days, within a week, and within a month. If the data query request is for expenditure data within three days, then the data query request matches the pre-set data type.
[0059] A data query request that doesn't match the pre-defined data type means the request contains non-enumerable filter criteria or multiple combined filter criteria. For example, detailed data with an amount range of xx to xxxx yuan, expenditure data with a payee name of xx, or expenditure data with a note of xx within a month are all data query requests that don't match the pre-defined data type.
[0060] Understandably, data query requests that don't match the pre-set data types are typically user-defined and difficult to enumerate individually. Furthermore, given that the income and expenditure detail system itself is primarily write-heavy and read-less, pre-enumerating and pre-categorizing these requests would actually consume computing and hardware resources. Therefore, for data query requests that don't match the pre-set data types, providing an immediate response based on the closest statistical data can reduce computing and hardware resource consumption while ensuring data query response speed.
[0061] S250: Determine query data that matches the data query request based on statistical data, and determine original data that matches the query data in the original data.
[0062] In this embodiment, for data query requests that match pre-set data types, that is, enumerable filtering conditions, since statistical data have been obtained by classification statistics in advance in the data preprocessing stage, the statistical data corresponding to the data query request (that is, query data) and the original data corresponding to the query data can be quickly filtered out directly in the column database through filters.
[0063] S260: Filter the statistical data according to the data query request to obtain original data that matches the statistical data corresponding to the data query request.
[0064] In this embodiment, a distributed search and data analysis engine can be used to respond to data query requests that do not match pre-set data types, such as non-enumerable filtering conditions or multiple combined filtering conditions.
[0065] Specifically, this embodiment adopts a combination of fixed data processing and streaming data processing to achieve responses to data query requests that do not match pre-set data types, thereby providing better data query performance.
[0066] Firstly, the idea of fixed data processing is adopted. According to the data query request, the statistical data that best meets the data query request is screened out from the statistical data in the column database, and the corresponding original data is obtained based on the statistical data.
[0067] For example, taking the data query request for querying the expenditure data with the remark information of xx within one month as an example, the distributed search and data analysis engine can be used to first query the column database to obtain the row key (Row Key) of the most suitable statistical data (expenditure data within one month), and then obtain the corresponding original data based on the row key of the statistical data.
[0068] S270: Screen the original data according to the data query request to obtain original data matching the data query request.
[0069] After obtaining statistical data that meets the data query request conditions based on fixed data processing and screening, further screening is performed according to the data query request in the original data corresponding to the statistical data based on the idea of streaming data processing.
[0070] Still taking the above example, after querying and obtaining the original data corresponding to the most suitable statistical data (expenditure data within one month), the original data with the remark information of xx is filtered out from the original data as the original data that matches the data query request.
[0071] The data query method of combining fixed data processing with streaming data processing in this embodiment improves the response speed of data query and realizes efficient and accurate data query.
[0072] Furthermore, if it is determined that the data query request does not match the preset data type, the original data and / or statistical data are filtered according to the data query request based on the sliding window to obtain the query data and the original data matching the query data.
[0073] The sliding window is used to dynamically partition the currently processed data within the data sequence. The sliding window can be moved in fixed steps, dividing a single data query task into multiple queries. Each data query only filters the raw data and / or statistical data within the sliding window, eliminating the need to filter the entire data set. The results of multiple queries are accumulated to obtain the complete query data and its corresponding raw data for the current data query task.
[0074] The advantage of setting a sliding window for multiple data queries is that it can ensure the response speed of data queries, achieve efficient and reliable real-time data queries, and reduce computing resource consumption.
[0075] The technical solution of the embodiment of the present invention adds data type identification to the raw data obtained in real time, classifies and counts the raw data with the identification, obtains statistical data, stores the raw data and statistical data in a column database, and implements data query and / or data modification based on the raw data and statistical data. This embodiment does not rely on a big data platform and divides data processing into a three-layer structure of data addition identification and classification statistics, data storage, and data query and modification, which facilitates horizontal expansion during data processing and improves data processing efficiency and applicability. When performing a data query, for enumerable filtering conditions that match a pre-set data type, since preliminary classification statistics have been performed during data preprocessing, the query data and its corresponding raw data can be quickly filtered out in the column database through a filter. For filtering conditions that do not match the pre-set data type, a data query method that combines fixed data processing and streaming data processing is adopted. The distributed search and data analysis engine first searches for statistical data that meets the data query request conditions, and then further filters the raw data in the raw data corresponding to the statistical data, thereby improving the response speed of the data query and achieving efficient and accurate data query.
[0076] Example 3
[0077] Figure 3 This is a flow chart of a data processing method provided in the third embodiment of the present invention. Based on the above embodiments, this embodiment of the present invention further specifies the process of data modification.
[0078] like Figure 3 As shown, the method includes:
[0079] S310: Determine the original data, and add a data type identifier to the original data based on a preset data type.
[0080] Furthermore, S310 may include: adding the original data to the original data table;
[0081] In this embodiment, when data type identification is added and classification statistics are performed on real-time original data, it is achieved by maintaining an original data table and an original data supplementary table.
[0082] Specifically, the original data table is used to store the original data obtained from the data transmission platform in real time, while the original data supplementary table is used to store the modified data within this classification statistical period.
[0083] In this embodiment, since the original data is acquired, labeled, and classified and counted in real time, directly modifying the original data will overwrite the original data in the database, which not only reduces the modification efficiency but also wastes computing and hardware resources. Therefore, this embodiment performs data modification based on the original data supplementary table, and implements data classification and statistics and storage based on the original data table and the original data supplementary table, eliminating the need for database reverse transactions. This not only improves data processing efficiency but also saves computing and hardware resources.
[0084] S320: Perform classification statistics on the original data after adding the data type identifier to obtain statistical data.
[0085] Furthermore, S320 may include: performing classification statistics on the original data in the original data table after the data type identifier is added, obtaining statistical data, and adding the statistical data to the statistical data table.
[0086] The statistical data table is used to store the statistical data obtained after classifying and counting the data in the original data table and the original data supplementary table.
[0087] In this embodiment, when there is no need to modify the original data within the current classification statistics period, the data preprocessing module of the income and expenditure details system reads the original data in the original data table, adds identifiers, and performs classification statistics.
[0088] S330: Store the original data and statistical data in a column database.
[0089] S340: Determine whether the data modification request is to modify statistical data or to modify original data. If it is to modify statistical data, execute S350; otherwise, execute S360.
[0090] In this embodiment, the data modification types are divided into two categories: modification of statistical data before the current classification statistical period, and modification of original data within the current classification statistical period.
[0091] Understandably, no new real-time raw data is generated before the current statistical period. Therefore, the statistical data and its corresponding raw data are permanently stored in the column database. Modifying either the statistical data or the raw data before the current statistical period requires modifying both the statistical data and the corresponding raw data. However, during the current statistical period, new raw data may be generated at any time, and the statistical data is constantly changing. Statistical data itself is inherently fluid, so modifications to data within the current statistical period are typically modifications to the raw data.
[0092] S350: Modify the statistical data that matches the data modification request, and modify the original data that matches the statistical data.
[0093] When modifying the statistical data before the current statistical period, the row key corresponding to the statistical data is determined according to the data modification request, the position of the statistical data in the column database is located, and the statistical data is modified. At the same time, the original data corresponding to the statistical data is also modified.
[0094] S360: Update the original data matching the data modification request into the original data supplement table.
[0095] In this embodiment, if the original data within the current statistical period needs to be modified, the modified original data is updated in the original data supplementary table for the current statistical period, rather than modifying the original data table and the statistical data table. This has the advantage of not modifying the original data table, thus avoiding reverse transactions and the inefficiency caused by concurrent data insertion. The statistical data table is not modified because the statistical data before the end of the current statistical period is already dynamically changing.
[0096] After the current classification statistics cycle ends, the original data in the original data table, the updated original data in the original data supplementary table, and the statistical data in the statistical data table can be stored in the column database. The original data table, the original data supplementary table, and the statistical data table are cleared to facilitate data classification statistics for the next classification statistics cycle.
[0097] S370: Perform classification statistics on the original data after the data type identifier is added in the original data table and the updated original data in the original data supplementary table.
[0098] After modifying the original data in the original data supplement table within this classification statistical period, data classification statistics are performed based on the original data in the original data table and the updated original data in the original data supplement table to obtain statistical data, which are stored in the statistical data table.
[0099] The technical solution of the embodiment of the present invention is to add data type identification to the original data obtained in real time, classify and count the original data with the identification added, obtain statistical data, store the original data and statistical data in a column database, and implement data query and / or data modification based on the original data and statistical data. This embodiment does not rely on the big data platform, and divides data processing into a three-layer structure of data addition identification and classification statistics, data storage, and data query and modification, which is convenient for horizontal expansion during data processing and improves data processing efficiency and applicability. When modifying data, for data modification requests for statistical data before this classification statistical period, the statistical data is located according to the row key of the statistical data, and the statistical data and its corresponding original data are modified to ensure data consistency. For data modification requests for original data within this classification period, the modified original data is updated to the original data supplementary table, rather than directly modifying the original data table and the statistical data table, thereby avoiding data overwriting, avoiding the inefficiency caused by concurrent data insertion, and improving the efficiency of data modification.
[0100] Example 4
[0101] Figure 4 This is a structural diagram of a data processing device provided by the fourth embodiment of the present invention. Figure 4 As shown, the device includes:
[0102] The identification adding module 410 is used to determine the original data and add a data type identification to the original data based on a preset data type;
[0103] The data statistics module 420 is used to perform classification statistics on the original data after adding the data type identifier to obtain statistical data;
[0104] The data storage module 430 is used to store the original data and statistical data in a column database, and perform data query and / or data modification based on the original data and statistical data.
[0105] The technical solution of the embodiment of the present invention adds data type identifiers to raw data obtained in real time, classifies and counts the identified raw data to obtain statistical data, stores the raw data and statistical data in a column database, and implements data query and / or data modification based on the raw data and statistical data. This embodiment does not rely on a big data platform and divides data processing into a three-tiered structure: data identification and classification statistics, data storage, and data query and modification. This facilitates horizontal expansion during data processing and improves data processing efficiency and applicability.
[0106] Based on the above embodiment, optionally, the data storage module 430 includes:
[0107] The first data query unit is configured to determine, based on statistical data, query data that matches the data query request if it is determined that the data query request matches a preset data type, and determine raw data that matches the query data in the raw data.
[0108] Based on the above embodiment, optionally, the data storage module 430 includes:
[0109] The second data query unit is used to filter the original data and / or statistical data based on the data query request based on the sliding window if it is determined that the data query request does not match the preset data type, to obtain the query data and the original data matching the query data.
[0110] Based on the above embodiment, optionally, the device further includes:
[0111] The statistical data screening module is used to screen the statistical data according to the data query request and obtain the original data that matches the statistical data corresponding to the data query request;
[0112] The original data screening module is used to screen the original data according to the data query request to obtain the original data that matches the data query request.
[0113] Based on the above embodiment, optionally, the data storage module 430 includes:
[0114] The statistical data modification module is configured to modify the statistical data that matches the data modification request if it is determined that the data modification request is to modify the statistical data, and to modify the original data that matches the statistical data.
[0115] Based on the above embodiment, optionally, the identification adding module 410 includes:
[0116] The original data table adding unit is used to add the original data into the original data table;
[0117] The data statistics module 420 includes:
[0118] The classification statistics unit is used to perform classification statistics on the original data after the data type identifier is added in the original data table, obtain statistical data, and add the statistical data to the statistical data table.
[0119] Based on the above embodiment, optionally, the data storage module 430 includes:
[0120] an original data modification unit, configured to update the original data matching the data modification request into the original data supplement table if it is determined that the data modification request is to modify the original data;
[0121] The device further comprises:
[0122] The classification statistics module is used to perform classification statistics on the original data after the data type identifier is added in the original data table and the updated original data in the original data supplementary table.
[0123] The data processing device provided by the embodiment of the present invention can execute the data processing method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.
[0124] Example 5
[0125] Figure 5 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0126] like Figure 5 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0127] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0128] The processor 11 may be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 11 performs the various methods and processes described above, such as the data processing method.
[0129] In some embodiments, the data processing method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the data processing method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the data processing method in any other suitable manner (e.g., by means of firmware).
[0130] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0131] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0132] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0133] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0134] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.
[0135] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.
[0136] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.
[0137] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.
Claims
1. A data processing method, characterized in that: include: Determine the original data and add a data type identifier to the original data based on the pre-set data type; Perform classification statistics on the original data after adding data type identifiers to obtain statistical data; The raw data and statistical data are stored in a column database, and data query and / or data modification are performed based on the raw data and statistical data.
2. The method according to claim 1, characterized in that Data query based on raw data and statistical data, including: If it is determined that the data query request matches the preset data type, query data matching the data query request is determined based on statistical data, and raw data matching the query data is determined in the raw data.
3. The method according to claim 1, characterized in that Data query based on raw data and statistical data, including: If it is determined that the data query request does not match the preset data type, the original data and / or statistical data are filtered according to the data query request based on the sliding window to obtain the query data and the original data matching the query data.
4. The method according to claim 3, characterized in that After determining that the data query request does not match the preset data type, the following is further included: Filtering statistical data according to the data query request to obtain original data matching the statistical data corresponding to the data query request; The original data is screened according to the data query request to obtain the original data that matches the data query request.
5. The method according to claim 1, wherein Data modification based on raw data and statistical data, including: If it is determined that the data modification request is to modify the statistical data, the statistical data matching the data modification request is modified, and the original data matching the statistical data is modified.
6. The method according to claim 1, characterized in that Identify the original data, including: Add the original data to the original data table; Perform classification statistics on the original data after adding data type identifiers to obtain statistical data, including: Perform classification statistics on the original data after adding data type identifiers in the original data table to obtain statistical data, and add the statistical data to the statistical data table.
7. The method according to claim 6, characterized in that Data modification based on raw data and statistical data, including: If it is determined that the data modification request is to modify the original data, then updating the original data matching the data modification request in the original data supplement table; The method further comprises: Classify and count the original data after adding the data type identifier in the original data table and the updated original data in the original data supplementary table.
8. A data processing device, characterized in that: include: An identification adding module is used to determine the original data and add a data type identification to the original data based on a preset data type; The data statistics module is used to classify and count the original data after adding data type identifiers to obtain statistical data; The data storage module is used to store the original data and statistical data in a column database, and perform data query and / or data modification based on the original data and statistical data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the data processing method according to any one of claims 1 to 7 is implemented.
10. A storage medium storing computer executable instructions, characterized in that: When the computer executable instructions are executed by a computer processor, they are used to perform the data processing method according to any one of claims 1 to 7.