Batch data processing method, device and electronic equipment
By analyzing the business scenarios and types of batch metadata and determining and applying appropriate data processing modes, the problem of low data processing reliability in the energy industry is solved, and efficient and accurate data processing is achieved.
Patent Information
- Application Number
- CN202011596568.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-29
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2040-12-29
AI Technical Summary
In the energy industry, the lack of batch data processing for different business scenarios leads to low reliability of data processing.
By obtaining batch metadata, analyzing its business scenario, collection time and data type, determining the corresponding data processing mode, and processing the data based on this mode, including format normalization, abnormal data removal, data merging and missing data filling.
It improves the reliability of data processing, reduces manual processing, reduces labor costs, and ensures the accuracy and reliability of the data set.
Smart Images

Figure CN113761103B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device, electronic device, storage medium, and computer program product for processing batch data. Background Art
[0002] With the development of big data and artificial intelligence, the relationship between the two is becoming increasingly close. The implementation of big data and artificial intelligence is inseparable from the support of massive amounts of data. Big data in the energy industry refers to large amounts of structured, semi-structured, and unstructured data obtained from various heterogeneous sources. This data is often believed to contain valuable information, but it requires considerable effort and resources to collect and process.
[0003] The energy industry places high demands on the authenticity and reliability of datasets generated for application scenarios. Therefore, reliably generating datasets that meet these requirements after acquiring raw data for different application scenarios is crucial for improving the intelligent control of the energy industry. However, related technologies lack batch data processing tailored to specific business scenarios, resulting in low reliability. Summary of the Invention
[0004] The present application provides a method, device and electronic device for processing batch data.
[0005] According to a first aspect of the present application, a method for processing batch data is provided, comprising:
[0006] Get the batch metadata to be processed;
[0007] Parsing the batch metadata to determine the business scenario to which the batch metadata belongs, the collection time and data type of each metadata;
[0008] Determining a first data processing mode according to the business scenario, the collection time and data type of each metadata;
[0009] The batch metadata is processed based on the first data processing mode.
[0010] According to a second aspect of the present application, a batch data processing device is provided, comprising:
[0011] A first acquisition module is used to acquire batch metadata to be processed;
[0012] A first determination module is configured to parse the batch metadata to determine the business scenario to which the batch metadata belongs, the collection time and data type of each metadata;
[0013] A second determination module is configured to determine a first data processing mode according to the business scenario, the collection time and the data type of each metadata;
[0014] A first processing module is configured to process the batch metadata based on the first data processing mode.
[0015] According to a third aspect of the present application, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the batch data processing method described in the above-mentioned embodiment.
[0016] According to a fourth aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, on which a computer program is stored. The computer instructions are used to enable the computer to execute the batch data processing method described in the embodiment of the above one aspect.
[0017] According to a fifth aspect of the present application, a computer program product is provided. When the computer program is executed by a processor, the batch data processing method described in the embodiment of the first aspect above is implemented.
[0018] The technical solution of the present application processes batch metadata according to the business scenario to which the batch metadata belongs, the collection time and data type of each metadata, thereby not only realizing the processing of batch metadata for different business scenarios and improving the reliability of data processing, but also helping to reduce manual processing links and reduce labor costs.
[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The accompanying drawings are provided to facilitate a better understanding of the present invention and do not constitute a limitation of the present application.
[0021] Figure 1 A flow chart of a method for processing batch data provided in an embodiment of the present application;
[0022] Figure 2A A schematic diagram of data before format normalization provided in an embodiment of the present application;
[0023] Figure 2B A schematic diagram of data after format normalization provided in an embodiment of the present application
[0024] Figure 3 A schematic diagram of configuration information Data_bound.csv provided in an embodiment of the present application;
[0025] Figure 4 A schematic diagram of configuration information Data_config.csv provided in an embodiment of the present application;
[0026] Figure 5 A comparison diagram before and after data filling provided in an embodiment of the present application;
[0027] Figure 6 A schematic diagram of a process for performing data processing based on a first data processing mode provided in an embodiment of the present application;
[0028] Figure 7 A schematic diagram of data statistics provided in an embodiment of the present application;
[0029] Figure 8 A schematic diagram of another process for performing data processing based on the first data processing mode provided in an embodiment of the present application;
[0030] Figure 9 A schematic diagram of a process for indexing and searching data processing results provided in an embodiment of the present application;
[0031] Figure 10 A schematic diagram of a batch data processing pipeline provided in an embodiment of the present application;
[0032] Figure 11 A schematic diagram of a final data processing result provided in an embodiment of the present application;
[0033] Figure 12 A schematic diagram of the structure of a batch data processing device provided in an embodiment of the present application;
[0034] Figure 13 The block diagram is a block diagram of an electronic device used to implement the batch data processing method of an embodiment of the present application. DETAILED DESCRIPTION
[0035] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, including various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present application. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0036] The following describes a batch data processing method, device, and electronic device according to an embodiment of the present application with reference to the accompanying drawings.
[0037] Figure 1 A flowchart of a method for processing batch data provided in an embodiment of the present application.
[0038] It should be noted that the execution subject of the batch data processing method in the embodiment of the present application can be an electronic device. Specifically, the electronic device can be but not limited to a server or a terminal, and the terminal can be but not limited to a personal computer, a smart phone, an IPAD, etc.
[0039] In the embodiment of the present application, a method for processing batch data is configured in a batch data processing device as an example. The device can be applied to an electronic device so that the electronic device can execute the method for processing batch data.
[0040] like Figure 1 As shown, the batch data processing method includes the following steps:
[0041] S101: Obtain batch metadata to be processed.
[0042] In the embodiments of this application, any raw batch data that needs to be processed in any business scenario is referred to as batch metadata to be processed. These business scenarios include, but are not limited to, complex industrial systems, such as combustion optimization, coal mill optimization, cold end optimization, and sootblowing optimization models in thermal power generation. They also encompass various business scenarios, including the application of digital twin technology in complex industrial systems.
[0043] There may be multiple metadata, which may be structured data, semi-structured data and / or unstructured data.
[0044] Specifically, if batch metadata needs to be processed, then the batch metadata to be processed is obtained, wherein the batch metadata may be collected by a sensor.
[0045] S102: parsing the batch metadata to determine the business scenario to which the batch metadata belongs, the collection time and data type of each metadata.
[0046] It should be noted that the data types of batch metadata in different business scenarios may be different. For example, the data types in industrial environment monitoring scenarios may include temperature data, humidity data, etc. The data types in industrial power statistics scenarios may include voltage data, current data, etc. The data types in equipment operation status monitoring scenarios may include operating time, operating efficiency, etc.
[0047] Specifically, after obtaining the batch metadata to be processed, the batch metadata can be parsed using relevant technical parsing methods, such as statistics, merging, and analysis, to determine the business scenario to which the batch metadata belongs, the collection time of each metadata item, and the data type. The parsing method can be determined based on the business scenario.
[0048] It should be noted that the method of parsing batch metadata in the embodiment of the present application can also be other methods in the relevant technology. It is sufficient to intelligently determine the business scenario to which the batch metadata belongs, the collection time and data type of each metadata. The embodiment of the present application does not limit this.
[0049] S103: Determine a first data processing mode according to the business scenario, the collection time and the data type of each metadata.
[0050] The first data processing mode may refer to a mode for processing batch metadata, which may include a processing method (such as merging, statistics, filtering, etc.) and a processing time.
[0051] Specifically, after determining the business scenario to which the batch metadata belongs, the collection time and data type of each metadata, the first data processing mode is determined according to the business scenario, the collection time and data type of each metadata.
[0052] Specifically, by querying preset data graphs, data tables or data texts, etc., the first data processing mode corresponding to the business scenario, the collection time and data type of each metadata can be obtained. Among them, the data graph, data table or data text can record the correspondence between the business scenario and the data processing mode, between the data collection time and the data processing mode, and between the data type and the data processing mode.
[0053] S104: Processing batch metadata based on the first data processing mode
[0054] Specifically, after the first data processing mode is determined, the batch data is processed based on the first data processing mode.
[0055] For example, for temperature monitoring data collected in the same business scenario or environment, some data can be merged according to the positional relationship of the monitoring points. The specific merging method may be related to the business scenario. If the business scenario requires prediction based on data and the use of a model, the merging method can be selected according to the type of model input data to merge the data according to the merging method; if the business scenario only requires simple statistical data for later application, the metadata can be counted according to the corresponding collection time and stored for subsequent calls.
[0056] By executing the above steps S101 to S104, when the original data in different application scenarios or business scenarios is obtained, the original data can be processed according to different application scenarios or business scenarios, so as to quickly and accurately generate a data set that meets the actual needs of the business scenario (such as those required by the relevant algorithm model).
[0057] It should be noted that after the batch metadata is processed, the processed batch metadata can be called according to the needs of the business scenario. For example, the processed batch metadata can be input into the corresponding algorithm model for model prediction.
[0058] The batch data processing method of the embodiment of the present application processes batch metadata according to the business scenario to which the batch metadata belongs, the collection time and data type of each metadata, thereby not only realizing batch data processing for different business scenarios and improving the reliability of data processing, but also helping to reduce manual processing links and reduce labor costs.
[0059] The first data processing mode in the above step S103 may be the data processing rules and the execution order of each processing rule, or may be the data processing rules and the processing time of each processing rule.
[0060] In one embodiment of the present application, when the first data processing mode is the data processing rules and the execution order between each processing rule, determining the first data mode in the above step S103 may include: determining the current data processing rules and the execution order between each processing rule.
[0061] In the embodiment of the present application, a variety of data processing rules with a wide coverage and the execution order of each processing rule can be pre-configured and stored so that they can be called when needed later.
[0062] It should be noted that the number of determined data processing rules can be one or more. When the number is one, only the current data processing rule is determined; when the number is multiple, the current data processing rule and the execution order of each processing rule need to be determined.
[0063] Specifically, after determining the business scenario to which the batch metadata belongs, the collection time and data type of each metadata item, a current data processing rule can be selected from multiple data processing rules, and the execution order of each processing rule can be determined based on the business scenario, the collection time and data type of each metadata item. The selection criteria can be determined based on actual needs. Then, step S104 described above can be executed to process the batch metadata based on the current data processing rule and the execution order of each processing rule, thereby achieving batch data processing.
[0064] For example, consider a boiler temperature monitoring scenario where individual temperature data needs to be fed into an algorithmic model for prediction to determine the boiler's operating status. Two processing rules can be selected: "Data Format Normalization" and "Multiple Data Merge." Furthermore, the execution order of "Data Format Normalization" and "Multiple Data Merge" must be determined, such as restrictive "Data Format Normalization" followed by "Multiple Data Merge." Subsequently, the individual temperature data can be normalized before the multiple data are merged, achieving complete processing of the individual temperature data.
[0065] Therefore, based on the current data processing rules and the execution order between each processing rule, batch metadata is processed, thereby improving the accuracy of data processing.
[0066] By analyzing and summarizing specific issues that arise during the data collection process for multiple sensors, we have determined that noise data can occur between sensor acquisition and data storage due to sensor failure, data transmission errors, and other factors. This noise data can include shutdown data that fails to accurately display industrial scenes, as well as abnormal, missing, and invalid data that occurs during normal operation. In embodiments of the present application, data processing rules can be configured to address this noise data.
[0067] That is, in one embodiment of the present application, the data processing rules may include at least one of the following rules: format normalization, abnormal data elimination based on boundary limits, data elimination based on specified time intervals, merging of data from multiple monitoring points based on similar data, abnormal data elimination based on business scenarios, and missing data filling based on known data.
[0068] Among them, format normalization can refer to converting different types of data formats into a unified data format, abnormal data elimination based on boundary limits can refer to eliminating abnormal data that exceeds the boundary limits (upper and lower limits), data elimination based on specified time intervals can refer to eliminating duplicate data within the specified time interval, merging of multiple monitoring point data based on similar data can refer to merging multiple monitoring point data with similar meanings, abnormal data elimination based on business scenarios can refer to eliminating abnormal data corresponding to business scenarios, and missing data filling based on known data can refer to filling data of different missing types based on existing data.
[0069] Specifically, after determining the business scenario to which the batch metadata belongs, the collection time and data type of each metadata, a current data processing rule can be selected from multiple data processing rules based on the business scenario, the collection time and data type of each metadata, and the execution order of each processing rule can be determined. The data processing rules may include at least one of format normalization, outlier data removal based on boundary limits, data removal based on specified time intervals, merging data from multiple monitoring points based on similar data, outlier data removal based on business scenarios, and missing data filling based on known data.
[0070] After that, the above step S104 can be executed, that is, the batch metadata is processed based on the current data processing rules and the execution order of each processing rule. The above step S104 is explained below based on each of the above data processing rules:
[0071] Specifically, when data processing rules include format normalization, sensor data (metadata) in different formats, such as text data, image data, or single sensor point data, can be normalized. For example, based on business scenario requirements, effective feature extraction can be performed on text data and image data to convert them into a single sensor data set. It should be noted that different data processing modes can be set for different data types, such as using image digital recognition technology to convert data information in an image into unified and standardized data.
[0072] Among them, during format normalization processing, data of the same type can be format normalized, for example, time data can be unified into the same format, weight data can be unified into the same format, etc.
[0073] For example, if the combustible matter of fly ash and slag of coal mill is counted, the metadata before and after format normalization can be compared as follows: Figure 2A and Figure 2B As shown, refer to Figure 2A Before formatting, the metadata includes time data, load data, fly ash combustible data and slag combustible data. At this time, the time data can be normalized into the same format, the load data can be normalized into the same format, and the fly ash combustible and slag combustible data can be normalized into the same format. Figure 2B After format normalization, the format of time data can be "×××× year ×× month ××", the format of related value (value) can be "×:××:××", and the format of related quality (quality) data can be "××.××××××××".
[0074] It should be noted that in an embodiment of the present application, a data processing configuration file can also be determined based on multiple batch metadata (monitoring point data). The following data processing rules: abnormal data elimination based on boundary limits, data elimination based on specified time intervals, merging of multiple monitoring point data based on similar data, abnormal data elimination based on business scenarios, and missing data filling based on known data can all be executed through the data processing configuration file.
[0075] That is, when the data processing rule includes the elimination of abnormal data based on the boundary limit, the abnormal data exceeding the boundary limit can be eliminated or deleted according to the upper and lower boundaries in the configuration information Data_bound.csv in the data configuration file, wherein the configuration information Data_bound.csv can be as follows: Figure 3 shown.
[0076] When the data processing rule includes data removal based on a specified time interval, a single sensor can be merged at the specified time interval to remove duplicate data and retain only the data from the same sensor at the same time point, thus merging single sensor data from multiple time periods into a single file. Data can also be resampled according to the configuration file.
[0077] When the data processing rules include the merging of multiple monitoring point data based on similar data, wherein the multiple monitoring point data can be measured by multiple single sensors, in order to make better use of the data, when there are multiple sensor monitoring point data belonging to the same type of data (data with similar meanings) in the data, the multiple sensor monitoring point data are merged to merge the multiple single sensor data files into a single multi-sensor data file, thereby reducing the data dimension and reducing the impact of single sensor anomalies. Among them, the multiple sensor data can be merged according to time, and at the same time, according to the grouping between sensors (for example, grouping according to sensor type), the multiple sensor monitoring point data can be merged to form a merge point. Furthermore, when merging multiple monitoring points, the data of 2 monitoring points can be merged to take the average, the data of 3 monitoring points can be taken to take the median, and the data of more than 3 monitoring points need to be merged and the average value is taken after removing the maximum and minimum values. This data processing rule can be implemented through the configuration information Data_config.csv in the data configuration file. The example of Data_config.csv can be as follows Figure 4 As shown, the Data_config.csv also includes the data feature dimensions (kks) included in the generated features. The rule merges multiple monitoring point data representing the same information to generate new feature point data.
[0078] When data processing rules include business scenario-based abnormal data removal, rules can be set to identify invalid or abnormal data based on business knowledge (such as common sense or general operating conditions) for each specific business scenario. For example, at coal mill startup, if the main motor current exceeds 200A (amperes), the entire main motor current data entry will be deleted. In a boiler efficiency management scenario, if the boiler efficiency exceeds 95%, the entire boiler efficiency data entry will be deleted. If the furnace negative pressure is outside of ±300°, it is considered an abnormal operating condition and the entire furnace negative pressure data entry will be deleted. Furthermore, data anomalies in multiple sensors can be identified based on specific data requirements. For example, if the coal feed rate for each coal mill is less than 1 and the total coal feed rate is less than 1, the data entry is considered invalid and deleted. This data processing rule can be implemented using the configuration information in the data configuration file, Shut_down.csv.
[0079] When the data processing rule includes filling missing data based on known data, filling operations are performed on the existing data for different types of missing data. For example, missing data caused by sensor damage or anomalies, or anomalies during data transmission and storage, are filled. This data processing rule can be implemented using the configuration information in the data configuration file, Multi_point.csv.
[0080] For example, Figure 5 From the comparison before and after data filling, we can see that there is missing data in the left picture. After data filling, the missing data in the left picture is filled, and the data shown in the right picture after filling is obtained.
[0081] Based on the above data processing rules, in order to reliably remove noise data and further improve the accuracy of data processing, the data processing rules may include all of the above rules. In this case, the execution flow or execution order of the data processing rules may be:
[0082] The first step is to standardize the format.
[0083] The second step is to eliminate abnormal data based on boundary limits.
[0084] The third step is to eliminate data based on the specified time interval.
[0085] The fourth step is to merge the data of multiple monitoring points based on the same type of data.
[0086] The fifth step is to eliminate abnormal data based on business scenarios.
[0087] Step 6: Fill in missing data based on known data.
[0088] Among them, when the sensor monitoring point data is abnormal or missing due to sensor damage or deterioration, it will be processed (i.e., eliminated) in the third and fifth steps to eliminate all missing or abnormal data. At the same time, the missing data will be filled by the filling algorithm in the sixth step. Common missing value filling methods include rule-based filling algorithms, such as forward filling, mean filling, etc., and model-based filling algorithms, such as Multivariate Imputation model.
[0089] Therefore, configuring multiple data processing rules enriches the data processing methods and can reliably eliminate noise data, thereby further improving the accuracy of data processing.
[0090] As described above, the specific first data processing mode is described, as well as how to process batch metadata based on the specific data processing mode in the above step S104.
[0091] It should be noted that when executing the above step S104, after data processing, in order to further improve the accuracy and reliability of data processing, the processed data may be verified or deviation detected. Based on this, the embodiments of the present application propose the following embodiments:
[0092] It should be noted that embodiments of the present application may pre-set data thresholds to determine whether the data processed by the data processing mode contains deviation data based on the data thresholds. The data thresholds may be determined based on the distribution of sensor monitoring point data, the business knowledge required for the business scenario, and the application requirements of the data.
[0093] In one embodiment of the present application, Figure 6 As shown, the above step S104 may include the following steps S601 to S603:
[0094] S601 , performing deviation detection on first output data processed in a first data processing mode to determine whether the first output data contains deviation data.
[0095] Specifically, after determining the first data processing mode, the batch metadata can be first input into the first data processing mode so that the first data processing mode can be used for processing, and then the first data processing mode can output the first output data after processing. At this time, it can be detected whether the first output data exceeds the data threshold. If the first output data exceeds the data threshold, then the first output data has deviation data; if the first output data does not exceed the data threshold, then the first output data does not have deviation data.
[0096] It can be understood that the deviation data may refer to the absolute value of the difference between the first output data and the data threshold.
[0097] S602 : When there is deviation data in the first output data, determine a second data processing mode according to the deviation data, the business scenario, the collection time and the data type of each metadata.
[0098] Specifically, when there is deviation data in the first output data, the second data processing mode is determined according to the deviation data and the business scenario determined in the above step S102, the collection time and data type of each metadata.
[0099] It should be noted that the second data processing mode may include at least one of eliminating data, reacquiring similar data, and refilling data. When the second data processing mode includes multiple processing rules, the execution order of the multiple processing rules may also be determined.
[0100] S603: Process the batch metadata based on the second data processing mode.
[0101] Specifically, after the second processing mode is determined, the batch metadata is processed based on the second data processing mode.
[0102] Specifically, in the case where the second processing mode includes eliminating data, reacquiring similar data, and refilling data, if a certain first output data exceeds the data threshold, then the first output data can be eliminated, and then similar data can be reacquired to refill the acquired similar data.
[0103] For example, if the first output data is Figure 7 The data in Figure 7As shown in the figure, the amount of monitoring point data at the monitoring point W.UNIT1.AMO2MANCO (the total number of data in the figure) is much less than that of other monitoring points. Therefore, there is deviation in the data (corresponding maximum value, minimum value, mean value, standard deviation, total number of data, non-duplicate data, lower quartile and upper quartile). In this case, the data can be deleted and similar data can be obtained again to supplement the obtained similar data. The monitoring points W.UNIT1.HAD11CT229 and W.UNIT1.HAD11CT224 are different from the two sensor monitoring points (W. UNIT1.HAD11CT217 and W.UNIT1.HAD13CT216) are similar monitoring points, but the monitoring point data ranges of monitoring points W.UNIT1.HAD11CT229 and W.UNIT1.HAD11CT224 are different from those of monitoring points W.UNIT1.HAD11CT217 and W.UNIT1.HAD13CT216, indicating that there are problems with the data of these two monitoring points, that is, there are deviated data. At this time, the data of these two monitoring points can be deleted, and similar data can be re-acquired to supplement the acquired similar data.
[0104] Therefore, when the first output data after being processed by the first data processing mode has deviations, the batch metadata is reprocessed according to the deviation data and the business scenario, the collection time and data type of each metadata, thereby avoiding deviations in data processing and further improving the accuracy and reliability of data processing.
[0105] In another embodiment of the present application, Figure 7 As shown, the above step S104 may include the following steps S801 to S804:
[0106] S801 : Based on the processing rules and execution order in the first data processing mode, batch metadata are processed in sequence to obtain input data and output data corresponding to each processing rule.
[0107] It is understood that the processing rules and execution order of the first data processing mode have been described above and will not be repeated here to avoid redundancy. The specific implementation method for sequentially processing batch metadata based on the processing rules and execution order of the first data processing mode has been described above and will not be repeated here to avoid redundancy.
[0108] It should be noted that the input data and output data corresponding to the processing rules may refer to the data before processing and the data after processing, respectively.
[0109] S802 , performing deviation detection on the second output data processed in the first data processing mode to determine whether the second output data contains deviation data.
[0110] Specifically, after determining the processing rules and execution order in the first data processing mode, the batch metadata can first be input into each processing rule according to the execution order, and then the second output data can be output after processing by each processing rule. At this time, it can be detected whether the second output data exceeds the data threshold. If the second output data exceeds the data threshold, then the second output data has deviation data; if the second output data does not exceed the data threshold, then the second output data does not have deviation data.
[0111] It can be understood that the deviation data may refer to the absolute value of the difference between the first output data and the data threshold.
[0112] S803: When there is deviation data in the second output data, determine a third data processing mode according to the deviation data, the business scenario, the collection time and the data type of each metadata.
[0113] Specifically, when there is deviation data in the first output data, the third data processing mode is determined according to the deviation data and the business scenario determined in the above step S102, the collection time and data type of each metadata.
[0114] It should be noted that the third data processing mode may also include various processing rules and the execution order between the various processing rules. The third data processing mode may contain the same processing rules as the first data processing mode.
[0115] S804. When the first N processing rules in the third data processing mode are the same as the first N processing rules in the first data processing mode and have the same execution order, the output data corresponding to the Nth processing rule in the first data processing mode is used as the input data for the N+1th processing rule in the third data processing mode to complete the processing process of the batch metadata in the third data processing mode, where N is a positive integer.
[0116] It should be noted that if the first N processing rules in the third data processing mode are the same as the first N processing rules in the first data processing mode and are executed in the same order, since both data processing modes process the same batch of metadata, the output data corresponding to the first N processing rules in the third data processing mode is the same as the output data corresponding to the first N processing rules in the first data processing mode. If the first N processing rules in the third data processing mode are the same as the first N processing rules in the first data processing mode but are executed in a different order, the output data corresponding to the first N processing rules in the third data processing mode may be different from the output data corresponding to the first N processing rules in the first data processing mode.
[0117] Specifically, after determining the third data processing mode, that is, determining the various processing rules and execution order therein, the various processing rules in the third data processing mode can be compared with the various processing rules and execution data in the first data processing mode. When the first N processing rules in the third data processing mode are the same as the first N processing rules in the first data processing mode and the execution order is the same, in order to avoid processing the same batch metadata again based on the same processing rules, the output data corresponding to the Nth processing rule in the first data processing mode can be obtained as the input data of the N+1th processing rule in the third data processing mode to complete the processing process of the batch metadata by the third data processing mode.
[0118] For example, if the third data processing mode includes 5 processing rules (Rule G1, Rule G2, Rule G3, Rule G4, Rule G5), and the execution order of the 5 processing rules is: Rule G1 → Rule G2 → Rule G3 → Rule G4 → Rule G5, and the first data processing mode includes 4 processing rules (Rule G1, Rule G2, Rule G3, Rule G6), and the execution order of the 4 processing rules is: Rule G1 → Rule G2 → Rule G3 → Rule G6, then the first 3 processing rules in the third data processing mode and the first 3 processing rules in the first data processing mode are the same. The processing rules are the same and the execution order is also the same (all execute rules G1, G2, and G3 in sequence), then the output data Y3 corresponding to the third processing rule G3 in the first data processing mode is obtained as the input data of the fourth processing rule G4 in the third data processing mode, that is, the input data Y3 is processed based on the fourth processing rule G4, and the output data Y4 is output after processing, and the output data Y4 is processed based on the fifth processing rule G5 to obtain the final processed data, thereby completing the processing process of batch metadata in the third data processing mode.
[0119] Furthermore, in the above step S804, before using the output data N corresponding to the Nth processing rule in the first data processing mode as the input data of the N+1th processing rule in the third data processing mode, it can also include: using the identifiers of the previous N processing rules and the corresponding execution order as indexes, retrieving and obtaining the output data corresponding to the Nth processing rule from the database.
[0120] It should be noted that the input data and output data corresponding to each processing rule obtained in step S801 can be stored in the database according to the identification of the corresponding processing rule and the corresponding execution order, for subsequent call.
[0121] Among them, the identifier of the processing rule can refer to any identifier that can be used as an index of the processing rule, which can be the content of the data file or configuration file of the processing rule, such as keywords such as "removal", "filling", "exception", etc., or it can be the intermediate result of the processing rule (input data and output data during the processing process).
[0122] Specifically, if the first N processing rules in the third data processing mode are the same as the first N processing rules in the first data processing mode and are executed in the same order, the identifiers of the first N processing rules can be obtained, and the output data corresponding to the Nth processing rule can be retrieved from the database using the identifiers of the first N processing rules and the corresponding execution order as indexes. During the retrieval, a reverse search can be performed in the database to find the output data corresponding to the Nth processing rule.
[0123] Specifically, if Figure 9 As shown, first, the file index index of all processing rules can be obtained, where the index can be a data file or configuration file index, or an intermediate result index. The data file and configuration file index represents the index operation of the content therein, and the intermediate result index is indexed by the input data file index and the configuration file in the corresponding process; then, a reverse search is performed. In the reverse search process, starting from the last step of data processing, the index of the current processing intermediate result is searched in the processing file save list; if it has been searched, it can be directly returned as a result; if it has not been searched, then enter the previous step in the data processing process to perform the same search process.
[0124] Therefore, by borrowing the output results of the same function in the previous processing process, it is possible to avoid repeating the same data processing, reduce data processing time, save processing resources, and improve data processing efficiency.
[0125] It should be noted that after the batch metadata is processed based on the third data processing mode, deviation detection may be performed on the processed data to further determine whether there is any deviation in the processed data.
[0126] That is, in one embodiment of the present application, after the above-mentioned step S804, it may also include: performing deviation detection on the third output data processed by the third data processing mode to determine whether there is deviation data in the third output data; if there is no deviation data in the third output data, deleting the input data and output data corresponding to each processing rule in the database.
[0127] Specifically, after the third data processing mode completes the processing of the batch metadata, the processed third output data is obtained, and then the third output data can be checked to see if it exceeds a data threshold. If the third output data exceeds the data threshold, then the third output data contains deviation data; if the third output data does not exceed the data threshold, then the third output data does not contain deviation data. If the third output data does not contain deviation data, the input data and output data corresponding to each processing rule in the database are deleted, thereby deleting the intermediate results of each processing process after the same metadata processing is completed.
[0128] That is to say, if the same batch of metadata undergoes N rounds of processing, then all new intermediate processing results in the first N-1 rounds of processing can be stored for subsequent calls.
[0129] In an embodiment of the present application, after each batch of metadata is processed based on the data processing mode, deviation detection can be performed on the processed data to perform corresponding processing based on the results of the deviation detection, thereby greatly improving the authenticity and effectiveness of data processing.
[0130] The following combination Figure 10 Describe the specific implementation of a specific embodiment of the present application:
[0131] like Figure 10 As shown, monitoring point data or metadata is collected through sensors A, B, and C, and batch metadata is obtained. The collected batch metadata is then checked and processed in a unified format and then processed. At the same time, a data processing configuration file can be determined and obtained based on the sensor to implement the data processing process based on the data processing configuration file. For example, the upper and lower boundaries in the configuration information Data_bound.csv are used to remove or delete abnormal data that exceeds the upper and lower limits; the configuration information Data_config.csv is used to merge a single monitoring point data list into a multi-point data set; the shutdown data is deleted through the configuration information Shut_down.csv; and the abnormal data is deleted and missing data is filled through the configuration information Multi_point.csv, thereby completing a data processing operation.
[0132] After the data processing is completed, the processed output data is tested for deviations. If there are deviations in the output data, the data processing fails because it is unqualified, and the next data processing is continued until it passes. If there are no deviations in the output data, it passes, indicating that the data processing is passed. When the data passes, the processed data is collected to form an algorithm data set for subsequent use. The final data set can be found in Figure 11 As shown, there is no invalid or abnormal data, and all data are real and valid.
[0133] In summary, the embodiments of the present application provide a data processing flow for the overall process of converting sensor point data into the required data set. Based on this flow, a data processing pipeline is established to optimize both time and labor costs. By encoding files and operation methods to compile and save data processing results, human involvement in the data processing process is reduced, ensuring consistency between configuration files and data processing results. Furthermore, by leveraging the reusability of intermediate data processing results, the time cost of repeated data processing is reduced, improving data processing efficiency.
[0134] The present application also proposes a batch data processing device. Figure 12 A schematic diagram of the structure of a batch data processing device provided in an embodiment of the present application.
[0135] like Figure 12 As shown, the batch data processing device 1200 includes: a first acquisition module 1210 , a first determination module 1220 , a second determination module 1230 and a third determination module 1240 .
[0136] Among them, the first acquisition module 1210 is used to obtain batch metadata to be processed; the first determination module 1220 is used to parse the batch metadata to determine the business scenario to which the batch metadata belongs, the collection time and data type of each metadata; the second determination module 1230 is used to determine the first data processing mode based on the business scenario, the collection time and data type of each metadata; the first processing module 1240 is used to process the batch metadata based on the first data processing mode.
[0137] In one embodiment of the present application, the second determining module 1230 may include: a first determining unit, configured to determine the current data processing rule and the execution order of each processing rule.
[0138] In one embodiment of the present application, the data processing rules include at least one of the following rules: format normalization, abnormal data elimination based on boundary limits, data elimination based on specified time intervals, merging of data from multiple monitoring points based on similar data, abnormal data elimination based on business scenarios, and missing data filling based on known data.
[0139] In one embodiment of the present application, the first processing module 1240 may include: a second determination unit, used to perform deviation detection on the first output data after being processed by the first data processing mode to determine whether there is deviation data in the first output data; a third determination unit, used to determine the second data processing mode according to the deviation data and business scenario, the collection time and data type of each metadata when there is deviation data in the first output data; and a first processing unit, used to process batch metadata based on the second data processing mode.
[0140] In one embodiment of the present application, the first processing module 1240 may include: a first acquisition unit, used to process the batch metadata in sequence based on the processing rules and execution order in the first data processing mode to obtain input data and output data corresponding to each processing rule; a fourth determination unit, used to perform deviation detection on the second output data processed by the first data processing mode to determine whether there is deviation data in the second output data; a fifth determination unit, used to determine the third data processing mode based on the deviation data and business scenario, the collection time and data type of each metadata when there is deviation data in the second output data; a second processing unit, used to use the output data corresponding to the Nth processing rule in the first data processing mode as the input data for the N+1th processing rule in the third data processing mode when the first N processing rules in the third data processing mode are the same as the first N processing rules in the first data processing mode and the execution order is the same, so as to complete the processing process of the batch metadata by the third data processing mode, where N is a positive integer.
[0141] In one embodiment of the present application, the first processing module 1240 may further include: a second acquisition unit, configured to retrieve and acquire output data corresponding to the Nth processing rule from the database using the identifiers and corresponding execution orders of the first N processing rules as indexes.
[0142] In one embodiment of the present application, the first processing module 1240 may also include: a sixth determination unit, used to perform deviation detection on the third output data after processing by the third data processing mode to determine whether there is deviation data in the third output data; a first deletion unit, used to delete the input data and output data corresponding to each processing rule in the database when there is no abnormal data in the third output data.
[0143] It should be noted that other specific implementations of the batch data processing device of the embodiment of the present application can refer to the specific implementations of the aforementioned batch data processing method, and to avoid redundancy, they are not repeated here.
[0144] The batch data processing device of the embodiment of the present application processes batch metadata according to the business scenario to which the batch metadata belongs, the collection time and data type of each metadata, thereby not only realizing the processing of batch metadata for different business scenarios and improving the reliability of data processing, but also helping to reduce manual processing links and reduce labor costs.
[0145] According to the embodiments of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product for a method for processing batch data. Figure 13 Provide explanation.
[0146] Figure 13 4 is a block diagram of an electronic device according to a method for processing batch data according to an embodiment of the present application.
[0147] like Figure 13 As shown, the electronic device 1300 may include: a memory 1310 and at least one processor 1320, and a bus 1330 connecting different components (including the memory 1310 and the processor 1320).
[0148] The memory 1310 stores a computer program, and when the processor 1320 executes the program, the batch data processing method of the embodiment of the present application is implemented.
[0149] Bus 1330 represents one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processor, or a local bus using any of a variety of bus architectures. Examples of these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnection (PCI) bus.
[0150] The electronic device 1300 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the electronic device 1300, including volatile and non-volatile media, removable and non-removable media.
[0151] The memory 1310 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 1340 and / or cache memory 1350. The electronic device 1300 may further include other removable / non-removable, volatile / non-volatile computer system storage media. For example only, the storage system 1360 may be used to read and write non-removable, non-volatile magnetic media ( Figure 13 Not shown, often called a "hard drive"). Although Figure 13 Not shown, a disk drive for reading and writing to a removable non-volatile disk (e.g., a floppy disk) and an optical disk drive for reading and writing to a removable non-volatile optical disk (e.g., a Compact Disc Read Only Memory (CD-ROM), a Digital Video Disc Read Only Memory (DVD-ROM), or other optical media) may be provided. In these cases, each drive may be connected to bus 1330 via one or more data medium interfaces. Memory 1310 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.
[0152] A program / utility 1380 having a set (at least one) of program modules 1370 may be stored, for example, in memory 1310. Such program modules 1370 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each of which, or some combination thereof, may include an implementation of a network environment. Program modules 1370 generally implement the functions and / or methods of the embodiments described herein.
[0153] The electronic device 1300 can also communicate with one or more external devices 1390 (e.g., a keyboard, pointing device, display 1391, etc.), one or more devices that enable a user to interact with the electronic device 1300, and / or any device that enables the electronic device 1300 to communicate with one or more other computing devices (e.g., a network card, a modem, etc.). This communication can be performed via an input / output (I / O) interface 1392. Furthermore, the electronic device 1300 can also communicate with one or more networks (e.g., a local area network (LAN), a wide area network (WAN), and / or a public network such as the Internet) via a network adapter 1393. As shown, the network adapter 1393 communicates with other modules of the electronic device 1300 via a bus 1330. It should be understood that, although not shown in the figure, other hardware and / or software modules can be used in conjunction with the electronic device 1300, including but not limited to microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0154] The processor 1320 executes various functional applications and data processing by running the programs stored in the memory 1310, such as implementing the methods mentioned in the above embodiments.
[0155] In the description of this specification, the reference terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" mean that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.
[0156] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of the present invention, "plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.
[0157] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, segment or portion of code comprising one or more executable instructions for implementing the steps of a custom logical function or process, and the scope of the preferred embodiments of the present invention includes alternative implementations in which functions may be performed out of the order shown or discussed, including performing functions in a substantially simultaneous manner or in the reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present invention pertain.
[0158] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). Furthermore, the computer-readable medium may even be paper or other suitable medium on which the program is printed, since the program may be obtained electronically, for example, by optically scanning the paper or other medium and then editing, interpreting or processing it in another suitable manner if necessary, and then storing it in a computer memory.
[0159] It should be understood that various parts of the present invention can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having a logic gate circuit for implementing a logic function on a data signal, an application-specific integrated circuit having a suitable combination of logic gate circuits, a programmable gate array (PGA), a field programmable gate array (FPGA), etc.
[0160] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.
[0161] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing module, or each unit may exist physically separately, or two or more units may be integrated into a single module. The aforementioned integrated modules may be implemented in the form of hardware or in the form of software functional modules. If the integrated modules are implemented in the form of software functional modules and sold or used as independent products, they may also be stored in a computer-readable storage medium.
[0162] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and are not to be construed as limiting the present invention. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.
Claims
1. A method for processing batch data, characterized in that: include: Get the batch metadata to be processed; Parsing the batch metadata to determine the business scenario to which the batch metadata belongs, the collection time and data type of each metadata; Determining a first data processing mode according to the business scenario, the collection time and data type of each metadata, wherein the first data processing mode refers to a mode for processing batch metadata; Processing the batch metadata based on the first data processing mode; The processing of the batch metadata based on the first data processing mode includes: performing a deviation detection on the first output data processed by the first data processing mode to determine whether the first output data has deviation data, where the deviation data refers to an absolute value of a difference between the first output data and a data threshold; If there is deviation data in the first output data, determine a second data processing mode based on the deviation data, the business scenario, the collection time and data type of each metadata; The batch metadata is processed based on the second data processing mode.
2. The method according to claim 1, wherein The determining of the first data processing mode includes: Determine the current data processing rules and the execution order of each processing rule.
3. The method according to claim 2, wherein The data processing rules include at least one of the following rules: format normalization, abnormal data elimination based on boundary limits, data elimination based on specified time intervals, merging of data from multiple monitoring points based on similar data, abnormal data elimination based on business scenarios, and missing data filling based on known data.
4. The method according to any one of claims 2-3, characterized in that: The processing of the batch metadata based on the first data processing mode includes: Based on the processing rules and execution order in the first data processing mode, the batch metadata is processed in sequence to obtain input data and output data corresponding to each processing rule; performing deviation detection on the second output data processed by the first data processing mode to determine whether the second output data has deviation data; If there is deviation data in the second output data, determine a third data processing mode based on the deviation data, the business scenario, and the collection time and data type of each metadata; When the first N processing rules in the third data processing mode are the same as the first N processing rules in the first data processing mode and are executed in the same order, the output data corresponding to the Nth processing rule in the first data processing mode is used as the input data of the N+1th processing rule in the third data processing mode to complete the processing process of the batch metadata by the third data processing mode, where N is a positive integer.
5. The method according to claim 4, wherein Before using the output data N corresponding to the Nth processing rule in the first data processing mode as the input data for the N+1th processing rule in the third data processing mode, the method further includes: The output data corresponding to the Nth processing rule is retrieved and obtained from the database using the identifiers of the first N processing rules and the corresponding execution order as indexes.
6. The method according to claim 4, wherein After completing the processing of the batch metadata in the third data processing mode, the method further includes: performing deviation detection on the third output data processed by the third data processing mode to determine whether the third output data has deviation data; When there is no deviation data in the third output data, the input data and output data corresponding to each processing rule in the database are deleted.
7. A batch data processing device, characterized in that: include: A first acquisition module is used to acquire batch metadata to be processed; A first determination module is configured to parse the batch metadata to determine the business scenario to which the batch metadata belongs, the collection time and data type of each metadata; A second determining module is configured to determine a first data processing mode according to the business scenario, the collection time and the data type of each metadata, wherein the first data processing mode is a mode for processing batch metadata; a first processing module, configured to process the batch metadata based on the first data processing mode; The first processing module includes: a second determining unit, configured to perform deviation detection on the first output data processed by the first data processing mode to determine whether the first output data has deviation data, where the deviation data refers to an absolute value of a difference between the first output data and a data threshold; a third determining unit, configured to, when there is deviation data in the first output data, determine a second data processing mode based on the deviation data, the business scenario, the collection time and the data type of each metadata; The first processing unit is configured to process the batch metadata based on the second data processing mode.
8. The device according to claim 7, wherein The second determining module includes: The first determining unit is used to determine the current data processing rule and the execution order of each processing rule.
9. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor. The instructions are executed by the at least one processor to enable the at least one processor to perform the batch data processing method according to any one of claims 1 to 6.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the batch data processing method according to any one of claims 1 to 6.
11. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the batch data processing method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Electric power statistical management system and method
CN106649880A
Method and device for website data acquisition
CN106776693A
Fault judging method, device and system for ship machine equipment, and storage medium
CN110044586A
Data collection method and device
CN110457310A
Disease prediction system
CN110957033A