Data processing method, system and equipment and storage medium
The data processing method addresses manual configuration complexity and enhances flexibility by generating operation logs, parsing parameters, and converting data to structured form with user-defined rules, improving data quality and efficiency.
Patent Information
- Application Number
- CN202510359617.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-25
AI Technical Summary
The existing data processing system lacks an effective data verification mechanism, resulting in insufficient data consistency and accuracy. Manual configuration items are cumbersome and complicated, making it difficult to meet flexible and changeable business needs. The serial processing method has become a performance bottleneck and lacks automation and flexibility.
By generating operation logs, analyzing key parameters, selecting rules and filtering conditions based on pre-set configuration items, automatically processing data operation instructions, converting them into structured data and filtering into standardized data, supporting incremental data table management.
It realizes flexible configuration and data screening according to business needs, reduces operational complexity, improves data processing efficiency and accuracy, and supports data processing in different business scenarios.
Smart Images

Figure CN120316166A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a data processing method, system, device, and storage medium. Background Art
[0002] Existing data processing systems use database management systems such as SQL Server, MySQL, Oracle, etc. to store and manage data, but lack an effective data verification mechanism, which affects data consistency and accuracy. The frequent adjustment of data structures due to changes in enterprise business requirements increases the complexity of data processing.
[0003] To solve the data verification problem, researchers have proposed a data verification method based on rules and transformations, which preprocesses and standardizes data by defining rules and transformation logics to improve data quality and reliability. However, the existing technology faces challenges in implementation, especially the configuration items are cumbersome and complex, and users need to manually set parameters and options, which increases the operation complexity and error rate.
[0004] In the existing technology, the cumbersome and complex manual setting of configuration item parameters is the main problem. In addition, the serial processing method becomes a performance bottleneck. The implementation scheme of the customized functions of the data middle platform depends on fixed processes and rules, lacking flexibility and automation. There are deficiencies in aspects such as data synchronization, filtering and screening, and log processing, and it is difficult to meet the complex and changeable business requirements. These problems limit the efficiency and accuracy of the data processing system and affect the performance of the data middle platform in complex business environments. Summary of the Invention
[0005] In view of this, this application provides a data processing method, system, device, and storage medium to solve the technical problem of the cumbersome and complex manual setting of configuration item parameters in the existing technology.
[0006] According to the first aspect of this application, a data processing method is provided, including: Generating an operation log in response to a data operation instruction sent by a user, where the data operation instruction represents a data processing rule selected by the user from pre-set configuration items; Parsing key parameters in the operation log to determine the target data identifier information and operation type in the data operation instruction; Fetching target data according to the target data identifier information, where the target data represents the entire data entry corresponding to the target data identifier information; Responding to a configuration specified rule set by the user, and converting the target data into structured data according to the configuration specified rule and the operation type, where the configuration specified rule represents at least one configuration condition selected from pre-set configuration items; Filter the structured data in response to the data filtering rules set by the user to obtain standardized data, where the data filtering rules refer to at least one data filtering condition selected from the pre-set configuration items.
[0007] Preferably, parsing the key parameters in the operation log to determine the target data identifier information and operation type in the data operation instruction includes the following steps: Parse the operation log to confirm the instruction trigger scenario; In the case where the instruction trigger scenario is a first-run trigger, create a temporary data table with the same structure as the original data table. The first-run trigger indicates that the operation log is generated for the first time, and the temporary data table is used to temporarily store and process the data in the original data table; Confirm the identifier information of all data in the temporary data table as the target data identifier information, and confirm the operation type as adding all data in the temporary data table.
[0008] Preferably, the data operation instruction includes at least one sub-transaction instruction, and each sub-transaction instruction includes at least a sub-transaction operation type, a sub-transaction operation object, and a sub-transaction operation content; Parsing the key parameters in the operation log to determine the target data identifier information and operation type in the data operation instruction includes the following steps: Parse the key parameters in the operation log to determine the sub-transaction instruction; Classify the sub-transaction instructions according to the sub-transaction operation type to generate at least one group of sub-transaction groups; Analyze the association relationship between each sub-transaction instruction in the sub-transaction group at least according to the sub-transaction operation object and the sub-transaction operation content; Generate a sub-transaction merging strategy according to the association relationship between each sub-transaction instruction in the sub-transaction group; Merge the sub-transaction instructions in each group of sub-transaction instruction groups according to the sub-transaction merging strategy to generate an instruction merging set; Analyze the instruction merging set to determine the target data identifier information and the operation type.
[0009] Preferably, in response to the configuration specifying rules set by the user, and according to the configuration specifying rules and the operation type, convert the target data into structured data, including the following steps: Construct a structured data framework according to the operation log and the target data. The structured data framework includes at least the identifier information of the operation log, the target data identifier information, the operation type, and the operation time; In response to the configured specified rules, identify and delete the specified fields in the target data to generate data to be structured; Extract all fields and the current value information of each field from the data to be structured; Fill the all fields and the current value information of each field into the structured data framework to generate structured data.
[0010] Preferably, after constructing the structured data framework according to the operation log and the target data, the following steps are further included: In the case where the operation type is to update the target data, extract the first impact fields from the target data, where the first impact fields represent the fields directly associated with the updated target data; Identify the second impact fields in the target data that are associated with the first impact fields; Merge the first impact fields and the second impact fields to generate an impact field list, where the impact field list includes at least the first impact fields and the second impact fields.
[0011] Preferably, after filtering the structured data according to the data filtering rules set by the user to obtain standardized data, the following is further included: Obtain the filtered data in the structured data, where the filtered data represents the data filtered by the data filtering rules; Write the filtered data into the filtered data table; Record the operation data related to the filtered data, where the operation data includes at least the data filtering rules and the time information when the filtered data is filtered.
[0012] Preferably, the method further includes: Create the incremental data table in the case where the incremental data table does not exist; In the case where the incremental data table exists, write the standardized data into the incremental data table according to the standardized data writing rules set by the user, where the incremental data table represents data increment records; Map the standardized data to the incremental data table.
[0013] According to the second aspect of the present application, a data processing system is provided, including: A log generation module, configured to generate an operation log in response to a data operation instruction sent by a user, where the data operation instruction represents a data processing rule selected by the user from pre-set configuration items; An analysis module, configured to analyze the key parameters in the operation log to determine the target data identifier information and the operation type in the data operation instruction; A data scraping module, configured to scrape target data according to the target data identifier information, where the target data represents the entire data of the data entry corresponding to the target data identifier information; A data conversion module, configured to respond to a configuration specification rule set by a user, and convert the target data into structured data according to the configuration specification rule and the operation type, where the configuration specification rule represents at least one configuration condition selected from pre-set configuration items; A filtering module, configured to filter the structured data in response to a data filtering rule set by a user to obtain standardized data, where the data filtering rule refers to at least one data filtering condition selected from the pre-set configuration items; A writing module, configured to write the standardized data into an incremental data table, where the incremental data table represents data increment records.
[0014] According to a third aspect of the present application, there is provided an electronic device, including: A processor; A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the instructions to implement the above data processing method.
[0015] According to a fourth aspect of the present application, there is provided a computer-readable storage medium, when instructions in the computer-readable storage medium are executed by a processor of a terminal, enabling the terminal to execute the above data processing method.
[0016] Compared with the prior art, the present application has the following advantages: First, the user selects at least one configuration condition from pre-set configuration items as a configuration specification rule according to actual needs, and selects at least one data filtering condition as a data filtering rule. By parsing key parameters in the operation log, the target data identifier information and the operation type in the data operation instruction are determined. The target data scraping module scrapes the target data representing the entire data of the data entry corresponding to the target data identifier information. According to the configuration specification rule and the data filtering rule, the target data is processed to generate standardized data. By flexibly selecting the configuration specification rule and the data filtering rule according to actual needs, data processing is realized, which has good flexibility. Configuration conditions and data filtering conditions can be customized according to different business scenarios, thereby effectively processing the original data generated in different business scenarios. It solves the technical problem of cumbersome manual setting of configuration item parameters in the prior art and reduces the cumbersome degree of data processing. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 is a step flowchart of a data processing method provided by an embodiment of the present application; Figure 2 It is a flowchart of a method for parsing operation logs provided by an embodiment of the present application; Figure 3 It is a flowchart of a method for merging sub-transaction instructions provided by an embodiment of the present application; Figure 4 It is a flowchart of a method for generating structured data provided by an embodiment of the present application; Figure 5 It is a flowchart of a method for generating a list of impact fields provided by an embodiment of the present application; Figure 6 It is a flowchart of a method for writing filtered data into a filtered data table provided by an embodiment of the present application; Figure 7 It is a flowchart of a method for creating an incremental data table in the case where the incremental data table does not exist provided by an embodiment of the present application; Figure 8 It is a flowchart of a method for mapping standardized data to an incremental data table provided by an embodiment of the present application; Figure 9 It is a first example diagram of a setting interface for standardized data writing rules provided by an embodiment of the present application; Figure 10 It is a second example diagram of a setting interface for standardized data writing rules provided by an embodiment of the present application; Figure 11 It is a schematic structural diagram of a data processing system provided by an embodiment of the present application; Figure 12 It is a structural block diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0018] To make the above objects, features, and advantages of the present application more obvious and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific implementation manners.
[0019] The following is the relevant background introduction: In existing data processing systems, database management systems such as SQL Server, MySQL, Oracle, etc. are usually used to store and manage massive data. However, these systems often expose the defect of lacking an effective data verification mechanism in actual applications, and cannot guarantee the consistency and accuracy of data. Due to the lack of data verification, data errors and outliers are difficult to be discovered and processed in a timely manner, thereby affecting the accuracy of data analysis and business decisions.
[0020] In addition, as the business needs of enterprises continue to change, the structure and format of data are also frequently adjusted. This dynamic nature not only increases the complexity of data processing but also poses higher adaptability requirements for data management systems. Traditional database management systems often struggle to cope with such frequent changes and are unable to provide flexible and efficient data processing solutions.
[0021] To address this issue, some researchers have proposed data validation methods based on rules and transformations. The core of this method lies in preprocessing and standardizing input data by defining a series of refined rules and transformation logics. During this process, the system can intelligently identify and handle common problems such as outliers and missing values, thus significantly improving the quality and reliability of data. However, although this method has enhanced data processing capabilities to a certain extent, it still faces some challenges.
[0022] Firstly, the configuration items of many existing systems are quite cumbersome and complex, and users need to manually set various parameters and options. This manual operation not only increases the complexity of the operation but also easily leads to configuration errors due to human mistakes. Secondly, for different types of data sources (such as table structures, views, XML files, etc.), existing systems often lack a unified data reading interface. This lack of uniformity makes system integration extremely difficult and requires a large amount of time and effort for adaptation and integration.
[0023] Finally, when dealing with large-scale data sets, the traditional serial processing method often becomes a performance bottleneck. Since serial processing cannot fully utilize the parallelism of modern computing resources, the processing speed is limited, which in turn affects the operating efficiency of the entire system.
[0024] As the core component of enterprise data management and analysis, the data middle platform undertakes the heavy responsibility of processing a large amount of data from different data sources. With the increasing complexity and diversification of enterprise business, the demand for customized functions of the data middle platform is becoming increasingly urgent. However, the existing solutions for implementing customized functions in the data middle platform usually rely on fixed data processing processes and rules, which results in the system lacking sufficient flexibility and automation when facing changing business needs.
[0025] In the existing technology, the closest implementation solutions include aspects such as data synchronization, data filtering and screening, and log processing. However, these solutions all have obvious deficiencies. In terms of data synchronization, existing technologies usually adopt full or incremental synchronization methods but lack fine-grained control at the data field level and are difficult to meet complex and changing business needs. In terms of data filtering and screening, the rules of existing solutions are relatively fixed and difficult to adapt to the rapidly changing market environment and business needs. In terms of log processing, existing technologies often require manual intervention for log analysis and processing, lacking automated and intelligent processing processes.
[0026] In summary, existing data processing and data middle platform technologies have obvious problems in terms of flexibility, automation, and compatibility. These problems limit the efficiency and accuracy of data processing systems and also affect the performance of data middle platforms in complex business environments. Therefore, developing more flexible, efficient, and intelligent data processing and data middle platform technologies has become an urgent issue to be solved.
[0027] Embodiment 1 See also Figure 1 As shown, a data processing method provided by the present application includes: S101. Generate an operation log in response to a data operation instruction sent by a user, where the data operation instruction represents a data processing rule selected by the user from pre-set configuration items.
[0028] Specifically, data operation instructions refer to the operation instructions actually executed by the user, including at least the addition, deletion and modification of data. Operation logs refer to detailed documents that record the user's data operation behavior, including key parameters such as operation time, operation user, operation instruction, and operation result. Operation time records the specific time point when the operation occurs, which helps to understand the sequence of operations and the activities of the system in different time periods. Operation user indicates the identity of the user who performs the operation, which is very important for permission auditing and responsibility tracing. Operation instructions are the core part of the log, recording the operation commands or requests actually executed by the user. When parsing, it is necessary to pay attention to the specific content of the instructions to determine the user's operation intention. The target data identifier information is used to specify the specific object of the operation, such as the record ID in the database, the path name in the file, etc. When parsing the log, it is necessary to identify and extract these identifiers in order to understand the specific target of the operation. Operation results record whether the operation is successfully executed and the result status after execution, which helps to judge the effectiveness of the operation and take corresponding remedial measures when necessary.
[0029] Exemplarily, the data operation instruction characterizes the data processing rules selected by the user from the pre-set configuration items. The data processing rules at least include a configuration specification rule, a data screening rule, and a standardized data writing rule. The configuration specification rule characterizes the rule for converting target data into structured data. The data screening condition characterizes the specific rule for screening structured data. The standardized data writing rule characterizes the specific rule for writing the target data as standardized data after processing the target data. The pre-set configuration items include a source database, a source data table, an exclusion field list, a data screening rule, a target database, a target data table, and a filtered data processing flow. The source database and the source data table characterize the database and the data table for storing the target data. The target database and the target data table characterize the database and the data table into which the standardized data obtained after processing the target data is written. Before the user sends an operation instruction, an operation instruction is generated by selecting a data processing rule from the pre-set configuration items. By setting the data processing rules through the pre-set configuration items, the situation where manual re-setting is required every time data processing is performed due to different data processing requirements is avoided. By selecting specific options in the pre-set configuration items to generate data processing rules, the data processing efficiency can be improved.
[0030] S102. Analyze the key parameters in the operation log to determine the target data identifier information and the operation type in the data operation instruction.
[0031] Specifically, the operation type refers to the specific operation behavior performed by the user. When analyzing the operation log, determining the operation type helps to understand the user's operation intention. When determining the operation type, it is first necessary to identify the data operation instruction part in the operation log. By searching for specific keywords or commands, the data operation instruction is identified. According to the specific content of the data operation instruction and the pre-set operation specifications, the operation type corresponding to the data operation instruction is determined. Exemplarily, when the data operation instruction is INSERT, the operation type is determined as a data insertion operation. When the data operation instruction is UPDATE, the operation type is determined as a data update operation.
[0032] S103. Fetch the target data according to the target data identifier information. The target data characterizes the entire data entry corresponding to the target data identifier information.
[0033] Specifically, the target data identifier information serves as a symbol or code that uniquely identifies the target data. Through the target data identifier information, the target data is located in the database or data system. Exemplarily, by comparing the parsed target data identifier information with the records in the database, the location of the target data can be determined, thereby fetching the target data.
[0034] The captured target data includes all relevant information of data entries, such as field values, associated data, etc. Each data entry usually contains multiple fields, and each field is used to store specific information. When capturing target data, it is necessary to ensure all field values of the data entry to ensure that the target data is completely characterized.
[0035] S104. In response to the configuration specification rules set by the user, and based on the configuration specification rules and operation types, convert the target data into structured data. The configuration specification rules represent at least one configuration condition selected from the pre-set configuration items.
[0036] Specifically, the configuration specification rules are set by the user according to actual needs and are a series of conditions and operations for guiding how to convert target data into structured data. Structured data refers to data with a fixed format and a finite set, which has the advantages of being easy to query and integrate. The configuration specification rules are usually based on pre-set configuration items, and the user can select at least one configuration condition from them to build the rules. The pre-set configuration items at least include data type, data format, and conversion logic. The specific data types at least include text, numerical value, and date. The data formats at least include JSON, CSV, and XML. The conversion logic at least includes data cleaning, data screening, data splitting, data mapping, etc.
[0037] According to the configuration specification rules set by the user, generate conversion rules for converting target data into structured data. The conversion rules are used to guide how to process the target data to achieve the conversion of target data into structured data. Before converting the target data into structured data, first preprocess the target data to ensure the integrity, accuracy, and suitability of the target data to meet expectations, so as to ensure the smooth progress of the conversion process. During the conversion process, according to the conversion rules and operation types, perform operations such as cleaning, mapping, and splitting on the target data to generate structured data. Exemplarily, the structured data at least includes operation log ID, target data identifier information, operation type, operation time, and current values of each field.
[0038] S105. In response to the data screening rules set by the user, filter the structured data to obtain standardized data. The data screening rules refer to at least one data screening condition selected from the pre-set configuration items.
[0039] Specifically, the data screening rules set by the user are set by the user according to actual needs. Standardized data refers to a data set that conforms to specific standards or specifications and is used to ensure data consistency, accuracy, and comparability. The data screening rules are selected by the user from the pre-set configuration items. According to the data screening rules set by the user, clarify the screening conditions and filter the structured data to obtain standardized data.
[0040] Exemplarily, the pre-set configuration items at least include database reading rules, data table reading rules, field list to be excluded, data screening conditions, database writing rules, data table writing rules, and data processing flow rules. More specifically, the data screening conditions at least include data type screening conditions, data format screening conditions, specific value screening conditions, data quality screening conditions, and business logic screening conditions. The data type screening conditions refer to the conditions for screening based on the essential type of data, including screening rules for different types of data such as dirty data, ignored data, and confidential data. The data format screening conditions are set based on the storage format of the data. The specific value screening conditions are for screening based on the values of specific fields in the data, such as only retaining or excluding data containing specific keywords, numerical ranges, or date intervals. The business logic screening is set according to business logic or rules. In the data table writing rules, there are valid data table writing rules and filtered data table writing rules. The valid data table is used to write the standardized data after data processing, and the filtered data table is used to write the filtered data. The data processing flow rules at least include the target person in charge, configuration specification rules, and trigger time. Among them, the target person in charge refers to the list of users who have the authority to send data operation instructions.
[0041] In this embodiment, the user first selects at least one configuration condition from the pre-set configuration items as the configuration specification rule according to the actual needs, and selects at least one data screening condition as the data screening rule. By parsing the key parameters in the operation log, the target data identifier information and operation type in the data operation instruction are determined. The target data identifier information grabs the target data representing the entire data entry corresponding to the target data identifier information. According to the configuration specification rule and the data screening rule, the target data is processed to generate standardized data. Flexibly selecting the configuration specification rule and the data screening rule according to the actual needs to achieve data processing has good flexibility. It is possible to customize the configuration conditions and data screening conditions according to different business scenarios, thereby effectively processing the original data generated in different business scenarios. This solves the technical problem of the cumbersome and complex manual setting of configuration item parameters in the prior art and reduces the cumbersome degree of data processing.
[0042] To better understand this application, the following will be elaborated in detail in combination with specific implementation manners.
[0043] See Figure 2 As shown, S101, parse the key parameters in the operation log to determine the target data identifier information and operation type in the data operation instruction, including: S201, parse the operation log to confirm the instruction trigger scenario.
[0044] Specifically, the instruction trigger scenarios include first run and daily real-time trigger. By parsing the operation logs, the specific time for responding to the data operation instructions sent by the user can be determined, and understanding under what circumstances the data operation instructions are triggered helps to determine the subsequent processing flow.
[0045] S202. In the case where the instruction trigger scenario is the first run trigger, create a temporary data table with the same structure as the original data table. The first run trigger indicates the first generation of operation logs, and the temporary data table is used to temporarily store and process the data in the original data table.
[0046] Specifically, when it is confirmed that the instruction trigger scenario is the first run, first, data table synchronization initialization will be performed. By creating a temporary data table with the same structure as the original data table, it is used to temporarily store and process the data in the original data table. Specifically, the temporary data table at least includes the same definitions of relevant fields, data types, and indexes as the original data table. By using the temporary data table, it is possible to avoid directly performing operations on the original data table that may cause problems, thereby reducing the risk of data loss or damage and avoiding directly affecting the original data table.
[0047] S203. Confirm the identifier information of all data in the temporary data table as the target data identifier information, and confirm the operation type as adding all data in the temporary data table.
[0048] Specifically, after the temporary data table is created and filled with data, it is necessary to determine the identifier information of all data in the temporary data table, and confirm the identifier information of all data as the target data identifier information. Confirming the operation type as adding all data in the temporary data table means regarding all data in the temporary data table as new data, that is, adding all data in the temporary data table to the target location.
[0049] By confirming the target data identifier information, it helps to ensure the uniqueness and accuracy of the data, avoid data duplication and conflicts. In the case where the instruction trigger scenario is the first trigger, confirming the operation type as adding new data provides clear guidance for subsequent data processing and ensures the accuracy and effectiveness of data processing.
[0050] See Figure 3 As shown, the data operation instruction includes at least one sub-transaction instruction, and each sub-transaction instruction at least includes a sub-transaction operation type, a sub-transaction operation object, and a sub-transaction operation content.
[0051] S101. Parse the key parameters in the operation log to determine the target data identifier information and operation type in the data operation instruction, including the following steps: S301. Parse the key parameters in the operation log to determine the sub-transaction instruction.
[0052] Specifically, the data operation instruction includes at least one sub-transaction instruction, and each sub-transaction instruction records a specific data operation. By parsing the key parameters in the operation log, the sub-transaction instruction is determined to support subsequent classification, correlation analysis, and merging.
[0053] S302. Classify the sub-transaction instructions according to the sub-transaction operation type to generate at least one group of sub-transaction groups.
[0054] Specifically, different sub-transaction instructions may belong to different operation types. Through classification, the instructions with similar operation types are grouped together to facilitate subsequent analysis and processing.
[0055] S303. Analyze the correlation relationship between each sub-transaction instruction in the sub-transaction group at least according to the sub-transaction operation object and the sub-transaction operation content.
[0056] S304. Generate a sub-transaction merging strategy according to the correlation relationship between each sub-transaction instruction in the sub-transaction group.
[0057] S305. Merge the sub-transaction instructions in each group of sub-transaction instruction groups according to the sub-transaction merging strategy to generate an instruction merging set.
[0058] S306. Analyze the instruction merging set to determine the target data identifier information and the operation type.
[0059] Specifically, by analyzing the correlation relationship between each sub-transaction instruction in the sub-transaction group, a sub-transaction merging strategy is formulated. According to the sub-transaction merging strategy, the sub-transaction instructions that meet the merging conditions are merged according to the sub-transaction merging strategy. After forming an instruction merging set, the operation type and the target identifier information of the instruction merging set are determined. Using the sub-transaction merging strategy to merge logically related or sequentially continuous sub-transaction instructions into a larger operation unit can simplify the operation log, reduce the redundancy of data operations, avoid the performance overhead caused by frequent writes to the database, reduce the number of disk input / output times, and improve the overall throughput and data processing efficiency.
[0060] See Figure 4 As shown in the figure, S104. In response to the configuration specified rule set by the user and according to the configuration rule and the operation type, convert the target data into structured data, including the following steps: S401. Construct a structured data framework according to the operation log and the target data. The structured data framework includes at least the identifier information of the operation log, the target data identifier information, the operation type, and the operation time.
[0061] Specifically, during the process of constructing the structured data framework, first, define the fields in the structured data framework based on the data operation instructions and target data in the operation log. These fields should cover all the key information in the operation log and target data. Then, determine the appropriate data type for each field to ensure data accuracy and consistency. Exemplarily, the operation time field uses the date and time type, while the operation type uses the enumeration type. Finally, clarify the relationships between multiple fields to generate the structured data framework. By constructing the structured data framework, a unified data model is provided, enabling the operation log and target data to be stored and managed in a standardized manner, which helps with subsequent data processing and analysis.
[0062] S402. In response to the configured specified rule, identify and delete the specified fields in the target data to generate the data to be structured.
[0063] Specifically, the specified fields in the target data represent the unnecessary fields, and the configured specified rule is selected according to the preset configuration items. Specifically, the configuration items are preset and generated, and the configured specified rule is determined based on factors such as field name, data type, value range, etc. After deleting the specified fields, the remaining data is the data to be structured.
[0064] By deleting the specified fields in the target data, data redundancy and complexity can be reduced, ensuring that only relevant and valuable data for analysis is retained, improving the efficiency and accuracy of data processing.
[0065] S403. Extract all fields and the current value information of each field from the data to be structured.
[0066] Specifically, by traversing the data to be structured, extract the name of each field and the corresponding current value information, and extract all fields and the current value information of each field to fill the corresponding fields in the structured data framework, converting the data to be structured into a structured form to ensure data accuracy and consistency.
[0067] S404. Fill all fields and the current value information of each field into the structured data framework to generate structured data.
[0068] Specifically, fill the name and value of each field correspondingly into the structured data framework to generate structured data. Using structured data ensures that data is stored and presented in a standardized, easy-to-manage, and analyzable manner, facilitating further analysis, querying, and report generation. Through structured data, users can more easily understand the meaning and structure of the data.
[0069] See Figure 5 As shown, after S401. Construct a structured data framework according to the operation log and target data, the following steps are further included: S501. When the operation type is to update target data, extract the first impact fields from the target data, where the first impact fields represent the fields directly associated with updating the target data.
[0070] Specifically, when the operation type is to update target data, extract the first impact fields directly associated with updating the target data from the target data. Determining the first impact fields can clarify which data has changed directly during the update operation, which helps with subsequent data processing and analysis.
[0071] S502. Identify the second impact fields in the target data that are associated with the first impact fields.
[0072] Specifically, determine the second impact fields by analyzing the relationships between various fields. Exemplarily, the association relationships between various fields are determined according to the data model or business logic. By determining the second impact fields, ensure that all fields that have changed indirectly due to the update operation are taken into consideration, further improving the accuracy and integrity of data processing.
[0073] S503. Merge the first impact fields and the second impact fields to generate an impact field list, where the impact field list includes at least the first impact fields and the second impact fields.
[0074] Specifically, during the process of merging the first impact fields and the second impact fields, ensure that there are no duplicate fields in the impact field list. By generating the impact field list, show which fields are affected when updating the target data, providing data support for subsequent data verification, auditing, and report generation, and ensuring the accuracy and integrity of the data are effectively maintained during the update operation.
[0075] See Figure 6 As shown, after S105. In response to the data filtering rule set by the user, filter the structured data to obtain the standardized data, it further includes: S601. Obtain the filtered data in the structured data, where the filtered data represents the data filtered by the data filtering rule.
[0076] Specifically, apply the data filtering rule to query or process the structured data, identify the data records that meet the filtering conditions, and form the filtered data by separating the selected data. Customize the filtering rule according to actual needs to retain high-quality data useful for processing or analysis, ensuring the accuracy of subsequent processing or analysis.
[0077] S602. Write the filtered data into the filtered data table.
[0078] Specifically, the filtered data table is used to store filtered data. Among them, the structure of the filtered data table is compatible with the format of the filtered data. By writing the filtered data into the filtered data table, it is convenient to review or analyze the filtered data, ensuring the transparency of the filtering operation during the data processing process, and facilitating auditing and traceability.
[0079] S603. Record the operation data related to the filtered data. The operation data includes at least the data screening rule and the time information when the filtered data is filtered.
[0080] Specifically, by recording the operation data related to the filtered data, it meets the requirements of data governance, compliance, and auditing. When there are data quality problems or system failures, the operation data related to the filtered data serves as an important basis for diagnosis and analysis. By analyzing the operation data, the bottlenecks and deficiencies in the data processing process can be identified, guiding the optimization and improvement of the data processing process.
[0081] See Figures 7 - 10 As shown, it also includes: S701. Create an incremental data table when the incremental data table does not exist.
[0082] Specifically, when the incremental data table does not exist, according to business requirements and data specifications, design the fields and data types of the incremental data table, and create the table name, field names, data types, and constraint conditions of the incremental data table. The incremental data table is used to store the data records that are newly added or updated over time, which helps in the effective management and analysis of data. By creating a dedicated incremental data table, the integrity and consistency of the data can be ensured, avoiding confusion with the original data or data in other processing processes.
[0083] S801. When the incremental data table exists, write the standardized data into the incremental data table according to the standardized data writing rule set by the user. The incremental data table represents the data increment records.
[0084] Specifically, the standardized data writing rule refers to the rule for writing the standardized data into the incremental data table according to a certain format, standard, and specification. The incremental data table is used to store the standardized data. According to the standardized data writing rule, writing the standardized data into the incremental data table can update the data records in real-time or periodically, reflecting the latest status of the data. In addition, the data in the incremental data table can be integrated with the original data or other data to support more complex data analysis and processing tasks. By setting the standardized data writing rule, the efficient and accurate writing of the standardized data into the incremental data table can be achieved, supporting the real-time update and integration of the data, improving data consistency and data processing efficiency, and providing strong support for complex data analysis and processing tasks.
[0085] Exemplarily, Figure 9 and Figure 10Displays the setting interface for the standardized data writing rules. The source table refers to the standardized data, and the target table represents the data table into which the standardized data is to be written. The specific process for the user to set the standardized data writing rules includes selecting the target database and target data table, configuring the data increment table, defining the fields of the increment data table, and setting the data writing rules. During the process of selecting the target database and target data table, the user first determines the database type into which the standardized data is to be written. Then, the user selects the target data table under this database type. When there is no increment data table under this database type, the setting interface for the standardized data writing rules provides an option to automatically create a table. By selecting to automatically create an increment data table and naming the increment data table to be created, the storage location of the standardized data can be determined.
[0086] After determining the increment data table, by defining the fields of the increment data table, the information that the increment data table needs to record is determined. When the standardized data is inconsistent with the fields of the increment data table, in a mapping manner, the standardized data is made to correspond one by one with the fields of the increment data table. In addition, the standardized data can also generate the corresponding fields of the increment data table through an integration method. After setting the database type, increment data table, and fields, by setting the timestamp, the type of change of the standardized data, and the specific content of the standardized data, the standardized data writing rules are generated.
[0087] S802. Map the standardized data to the increment data table.
[0088] Specifically, during the mapping process, first, the mapping relationship between each field of the standardized data and the corresponding field of the increment data table is clarified. The mapping relationship includes at least matching fields or data type conversion. By defining and executing the data mapping rules, it is ensured that the standardized data can be accurately and consistently written into the increment data table, avoiding data loss or errors. Through reasonable mapping, the data in the increment data table is made more understandable and analyzable, supporting more efficient data processing tasks.
[0089] Embodiment 2 See Figure 11 As shown, the present application provides a data processing system, including: A log generation module 400, configured to generate an operation log in response to a data operation instruction sent by a user, where the data operation instruction represents a data processing rule selected by the user from pre-set configuration items.
[0090] An analysis module 500, configured to analyze key parameters in the operation log to determine the target data identifier information and operation type in the data operation instruction.
[0091] A data capture module 600, configured to capture target data according to the target data identifier information, where the target data represents the entire data entry corresponding to the target data identifier information.
[0092] A data conversion module 700, configured to respond to a configuration specification rule set by a user, and convert target data into structured data according to the configuration specification rule and an operation type, where the configuration specification rule represents at least one configuration condition selected from pre-set configuration items.
[0093] A filtering module 800, configured to filter the structured data in response to a data filtering rule set by a user to obtain standardized data, where the data filtering rule refers to at least one data filtering condition selected from pre-set configuration items.
[0094] In some embodiments, the data scraping module 600 includes: A parsing unit, configured to parse an operation log to confirm an instruction trigger scenario.
[0095] A creating unit, configured to create a temporary data table having the same structure as an original data table when the instruction trigger scenario is a first-run trigger, where the first-run trigger represents that an operation log is generated for the first time, and the temporary data table is used to temporarily store and process data in the original data table.
[0096] A confirming unit, configured to confirm identifier information of all data in the temporary data table as target data identifier information, and confirm the operation type as adding all data in the temporary data table.
[0097] In some embodiments, a data operation instruction includes at least one sub-transaction instruction, and each sub-transaction instruction includes at least a sub-transaction operation type, a sub-transaction operation object, and sub-transaction operation content.
[0098] The data scraping module 600 further includes: parsing key parameters in the operation log to determine sub-transaction instructions.
[0099] A classifying unit, configured to classify the sub-transaction instructions according to the sub-transaction operation type to generate at least one group of sub-transaction groups.
[0100] An analyzing unit, configured to analyze an association relationship between each sub-transaction instruction in the sub-transaction group at least according to the sub-transaction operation object and the sub-transaction operation content.
[0101] A policy generating unit, configured to generate a sub-transaction merging policy according to the association relationship between each sub-transaction instruction in the sub-transaction group.
[0102] A merging unit, configured to merge the sub-transaction instructions in each group of sub-transaction instruction groups according to the sub-transaction merging policy to generate an instruction merging set.
[0103] A determining unit, configured to analyze the instruction merging set to determine target data identifier information and an operation type.
[0104] In some embodiments, the data conversion module 700 includes: A construction unit, configured to construct a structured data framework according to the operation log and the target data. The structured data framework at least includes identifier information of the operation log, target data identifier information, operation type, and operation time.
[0105] A deletion unit, configured to identify and delete specified fields in the target data in response to a configured specified rule, and generate data to be structured.
[0106] An extraction unit, configured to extract all fields and current value information of each field from the data to be structured.
[0107] A structured data unit, configured to fill all fields and current value information of each field into the structured data framework to generate structured data.
[0108] In some embodiments, it further includes: An associated field module, configured to extract first affected fields from the target data when the operation type is to update the target data. The first affected fields represent fields directly associated with the updated target data.
[0109] An identification module, configured to identify second affected fields in the target data that are associated with the first affected fields.
[0110] An affected field list module, configured to merge the first affected fields and the second affected fields to generate an affected field list. The affected field list at least includes the first affected fields and the second affected fields.
[0111] In some embodiments, it further includes: A filtered data module, configured to obtain filtered data in the structured data. The filtered data represents data filtered by a data screening rule.
[0112] A writing module, configured to write the filtered data into a filtered data table.
[0113] A recording module, configured to record operation data related to the filtered data. The operation data at least includes the data screening rule and time information when the filtered data is filtered.
[0114] For the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For related parts, refer to the partial description of the method embodiment.
[0115] Each embodiment in this specification is described in a progressive manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0116] Embodiment III Please refer toFigure 12 , embodiments of the present application further provide an electronic device, including: A processor.
[0117] A memory for storing executable instructions of the processor.
[0118] Wherein, the processor is configured to execute instructions to implement any one of the data processing methods.
[0119] In this embodiment, the computer device includes a processor, a memory, and a network interface connected through a system bus.
[0120] Wherein, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data samples. The network interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements any one of the data processing methods.
[0121] Those skilled in the art can understand that Figure 12 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0122] Embodiment 4 Embodiments of the present application further provide a computer-readable storage medium. When the instructions in the computer-readable storage medium are executed by a processor of a terminal, the terminal can execute any one of the data processing methods.
[0123] The above-mentioned computer-readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a static random access memory, an electrically erasable programmable read-only memory, an erasable programmable read-only memory, a programmable read-only memory, a read-only memory, a magnetic memory, a flash memory, a magnetic disk, or an optical disk. The readable storage medium can be any available medium accessible by a general-purpose or special-purpose computer.
[0124] Optionally, a readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be an integral part of the processor. The processor and the readable storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor and the readable storage medium can also exist as discrete components in a device.
[0125] Embodiment 5 The embodiments of the present application also provide a computer program product. The computer program product includes a computer program which, when executed by a processor, implements any of the data processing methods.
[0126] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program code.
[0127] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0128] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for realizing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or the combination of blocks.
[0129] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, causing a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, so that the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 or steps for implementing the functions specified in one block or a plurality of blocks.
[0130] Although the preferred embodiments of the present invention have been described, additional changes and modifications can be made by those skilled in the art once they learn of the basic inventive concept. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.
[0131] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
[0132] The above has introduced in detail the data processing method, system, device, and storage medium provided in this application. Specific examples are used in this article to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application; at the same time, for those of ordinary skill in the art, according to the idea of this application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to this application.
Claims
1. A data processing method, characterized in that, It includes the following steps: In response to a data operation instruction sent by a user, generate an operation log, where the data operation instruction characterizes a data processing rule selected by the user from pre-set configuration items; Parse key parameters in the operation log to determine the target data identifier information and operation type in the data operation instruction; Fetch target data according to the target data identifier information, where the target data characterizes the entire data entry corresponding to the target data identifier information; In response to a configuration specification rule set by the user, and according to the configuration specification rule and the operation type, convert the target data into structured data, where the configuration specification rule characterizes at least one configuration condition selected from pre-set configuration items; In response to a data screening rule set by the user, filter the structured data to obtain standardized data, where the data screening rule refers to at least one data screening condition selected from the pre-set configuration items.
2. The data processing method according to claim 1, wherein The step of parsing key parameters in the operation log to determine the target data identifier information and operation type in the data operation instruction includes the following steps: Parse the operation log to confirm the instruction trigger scenario; In the case where the instruction trigger scenario is a first-run trigger, create a temporary data table with the same structure as the original data table. The first-run trigger indicates that the operation log is generated for the first time, and the temporary data table is used to temporarily store and process data in the original data table; Confirm the identifier information of all data in the temporary data table as the target data identifier information, and confirm the operation type as adding all data in the temporary data table.
3. The data processing method according to claim 1, characterized in that The data operation instruction includes at least one sub-transaction instruction, and each sub-transaction instruction includes at least a sub-transaction operation type, a sub-transaction operation object, and a sub-transaction operation content; The step of parsing key parameters in the operation log to determine the target data identifier information and operation type in the data operation instruction includes the following steps: Parse key parameters in the operation log to determine the sub-transaction instruction; Classify the sub-transaction instruction according to the sub-transaction operation type to generate at least one group of sub-transaction groups; Analyze the association relationship between each sub-transaction instruction in the sub-transaction group at least according to the sub-transaction operation object and the sub-transaction operation content; Generate a sub-transaction merging strategy according to the association relationship between each sub-transaction instruction in the sub-transaction group; Merge the sub-transaction instructions in each group of sub-transaction instruction groups according to the sub-transaction merging strategy to generate an instruction merging set; Analyze the instruction merging set to determine the target data identifier information and the operation type.
4. The data processing method according to claim 1, wherein The step of, in response to a configuration specification rule set by the user, and according to the configuration specification rule and the operation type, converting the target data into structured data includes the following steps: Construct a structured data framework according to the operation log and the target data. The structured data framework includes at least the identifier information of the operation log, the target data identifier information, the operation type, and the operation time; In response to the configuration-specified rule, identify and delete the specified fields in the target data to generate data to be structured; Extract all fields and the current value information of each field from the data to be structured; Fill the all fields and the current value information of each field into the structured data framework to generate structured data.
5. The data processing method according to claim 1, wherein After constructing the structured data framework according to the operation log and the target data, the following steps are further included: In the case where the operation type is to update the target data, extract the first affected fields from the target data, and the first affected fields represent the fields directly associated with the updated target data; Identify the second affected fields in the target data that are associated with the first affected fields; Merge the first affected fields and the second affected fields to generate a list of affected fields, and the list of affected fields includes at least the first affected fields and the second affected fields.
6. The data processing method according to claim 1, wherein After filtering the structured data in response to the data filtering rules set by the user to obtain standardized data, the following is further included: Obtain the filtered data in the structured data, and the filtered data represents the data filtered by the data filtering rules; Write the filtered data into the filtered data table; Record the operation data related to the filtered data, and the operation data includes at least the data filtering rules and the time information when the filtered data is filtered.
7. The data processing method according to claim 1, wherein The method further includes: Create the incremental data table in the case where the incremental data table does not exist; In the case where the incremental data table exists, write the standardized data into the incremental data table according to the standardized data writing rules set by the user, and the incremental data table represents data increment records; Map the standardized data to the incremental data table.
8. A data processing system, characterized in that, Includes: A log generation module, configured to generate an operation log in response to a data operation instruction sent by a user, and the data operation instruction represents a data processing rule selected by the user from pre-set configuration items; A parsing module, configured to parse the key parameters in the operation log to determine the target data identifier information and the operation type in the data operation instruction; A data scraping module, configured to scrape target data according to the target data identifier information, and the target data represents the entire data entry corresponding to the target data identifier information; A data conversion module, configured to respond to the configuration-specified rules set by the user, and convert the target data into structured data according to the configuration-specified rules and the operation type, and the configuration-specified rules represent at least one configuration condition selected from pre-set configuration items; A filtering module, configured to filter the structured data in response to the data filtering rules set by the user to obtain standardized data, and the data filtering rules refer to at least one data filtering condition selected from the pre-set configuration items.
9. An electronic device, characterized in that, Includes: A processor; A memory, configured to store instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the data processing method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by a processor of the terminal, the terminal is enabled to execute the data processing method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Data processing device and method
CN103020088A
Log data screening method and device
CN106874354A
Data capturing method and device and computer readable storage medium
CN109614539A
Streaming data processing method and device
CN115062002A
Real-time analytical database system for querying data of transactional systems
US11327962B1