Data processing method based on flag bit
By using flags in data processing to identify data missing and cleaning according to the degree of missing and data type, the problem of missing data processing in data cleaning is solved, the quality and reliability of data are improved, and a good foundation is laid for data analysis and mining.
Patent Information
- Application Number
- CN201910566238.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2019-06-27
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2039-06-27
AI Technical Summary
During the data cleaning process, how to effectively deal with the problem of missing data, especially for data with time attributes and data containing intrinsic connections, how to use time-related data to fill and eliminate missing data.
The data processing method based on flag bits is adopted. By obtaining data including flag bits, the flag bits are used to identify whether the data is missing, and the data is cleaned according to the degree of missing data and the data type. Specific steps include obtaining the degree of missing data, deciding the cleaning method based on the degree of missing data and the type of data, such as discarding data, filling the missing data using relevant measurements, etc.
It improves the robustness and reliability of data, ensures the efficiency and accuracy of data cleaning, and provides a solid foundation for subsequent data analysis and data mining.
Smart Images

Figure CN110457293B_ABST
Abstract
Description
Technical Field
[0001] The invention relates to a data processing method based on a flag bit. Background Art
[0002] After years of information system construction, many institutions have their own relatively complete information systems to meet the different information needs in various fields. In particular, some large institutions have completed the construction of big data centers, achieved unified data sharing and data integration, accumulated massive data in production and operation, and comprehensively promoted the construction of public data resource pools, providing favorable conditions for data centralized sharing, analysis and utilization. With the continuous deepening of information construction and application, the data generated by information systems has become a valuable asset for various institutions. Therefore, how to clean the data generated by various information systems, improve data quality, and tap the value of data resources has become one of the important tasks in the informationization projects of major institutions.
[0003] Among these, from the perspective of time, data cleaning is the first task. As the name suggests, data cleaning means to "clean away" the "dirty", which refers to the last procedure of discovering and correcting identifiable errors in data files, including checking data consistency, handling invalid values and missing values, etc. Because the data in the data warehouse is a collection of data for a certain subject, these data are extracted from multiple business systems and contain historical data, so it is inevitable that some data are wrong data and some data conflict with each other. These wrong or conflicting data are obviously what we don't want, and are called "dirty data". We need to "clean away" the "dirty data" according to certain rules, which is data cleaning.
[0004] In the process of data cleaning, in order to improve the cleaning efficiency and prevent the loss of important data, it is also necessary to treat the data differently according to its attributes. Some unimportant data can be simply discarded. For some data with time attributes, such as data collected from multiple measurements, if a certain data is missing, how to use time-related data to fill it is a technical issue worthy of study. For some data with intrinsic connections, such as continuously changing temperature and humidity, how to find its boundary value through historical data and logical reasoning to eliminate possible errors in order to better clean the data is also a technical problem in the current data cleaning field. Summary of the invention
[0005] In order to solve the above technical problems, the present invention proposes a data processing method based on a flag bit, comprising:
[0006] Step 1000: Acquire data including a flag bit, wherein the flag bit is used to identify whether the data is missing.
[0007] Step 2000: Obtain the degree of data missing according to the flag bit of the data.
[0008] Step 3000: clean the data according to the degree of missingness and data type of the data.
[0009] The present invention can clean the data according to the degree of missing data and the type of data, thereby improving the robustness and reliability of the data and laying a solid foundation for subsequent data analysis and data mining. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] Figure 1 It is a flowchart of a specific implementation mode 1 of the present invention.
[0011] Figure 2 This is a schematic diagram of collecting data for the first measurement of the present invention.
[0012] Figure 3 This is a schematic diagram of collecting data for the second measurement of the present invention.
[0013] Figure 4 This is a flow chart of a second specific implementation mode of the present invention. DETAILED DESCRIPTION
[0014] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail in conjunction with the accompanying drawings. This description introduces specific embodiments consistent with the principles of the present invention by way of example and not limitation, and the description of these embodiments is sufficiently detailed to enable those skilled in the art to practice the present invention, and other embodiments may be used and the structure of each element may be changed and / or replaced without departing from the scope and spirit of the present invention. Therefore, the following detailed description should not be understood in a restrictive sense.
[0015] like Figure 1 As shown, according to a first aspect of the present invention, a data processing method based on a flag bit is proposed, comprising:
[0016] Step 1000: Obtain data including a flag bit, wherein the flag bit is used to identify whether a corresponding field of the data is missing.
[0017] In one embodiment, the flag bit uses 0 and 1 to represent whether the corresponding field in the data is missing. For example, when the third flag bit of the data is 0, it means that the third field of the data is missing, and when the third flag bit of the data is 1, it means that the third field of the data is complete. For example, a metadata description method can be used to set a corresponding flag bit for each field of the data.
[0018] In another embodiment, a data set including a flag bit is obtained. ,in, , n For the dataset Z The number of data in for z i The j data fields, m for z i The number of data fields in for Flag bit, when the field When missing, is 0; when the field When complete, is 1; c i For data types; for example, when =0, represents z i The third field of is missing, which is =1, represents z i The third field of is complete; t i for z i The timestamp of the measurement acquisition data.
[0019] According to the present invention, the data includes measurement and collection data, and the measurement and collection data is obtained by manual or automated measurement and collection. The measurement and collection data includes first measurement and collection data and second measurement and collection data. Figure 2 As shown, the first measurement and acquisition data is the measurement and acquisition data including periodic measurement and acquisition data in the data field part; Figure 3 As shown, the second measurement and collection data is recorded as measurement and collection data consisting of periodic measurement data. Periodic measurement data refers to data measured and collected at fixed time intervals, such as daily humidity, hourly temperature, and network connection heartbeats per minute. The first measurement and collection data and the second measurement and collection data also include a timestamp, which is used to record the time of the measurement and collection data. In particular, the timestamp can also be used to identify the measurement and collection period of the periodic measurement data.
[0020] Step 2000: Obtain the degree of data missing according to the flag bit of the data.
[0021] Furthermore, the degree of deficiency includes a first degree of deficiency and a second degree of deficiency, wherein the first degree of deficiency s 1 = m 0 , the second missing degree ,in m 0is the number of all flag bits of the data whose value is 0; m 1 The number of all flag bits of the data whose values are 1.
[0022] Using the first missing degree and the second missing degree to measure the data missing situation respectively can examine the data missing situation from two aspects: the number of missing fields and the ratio of the number of missing fields to the number of complete fields. It can more accurately determine the severity of the data missing situation and point out the direction for the next step of data processing.
[0023] Step 3000: clean the data according to the degree of missingness and data type of the data.
[0024] Furthermore, if s 1 Greater than the first threshold k 1 or s 2 Greater than the second threshold k 2 , then discard the data directly.
[0025] if s 1 Less than or equal to the first threshold k 1 and s 2 Less than or equal to the second threshold k 2 ,and z i Collect data for the first measurement, then ,in, v for z i Missing measurement acquisition data fields, v w The timestamp of the record is less than t i , v w The field and v same, u The number of data points collected for the randomly selected first measurement, u∈N , preferably u ≥3, more preferably 5.
[0026] in, k 1 ∈N +, k 1 ≥2, preferably, k 1 The number of data fields m A function, more preferably, ,in is the rounding symbol; k 2 is a positive real number, k 2 ∈R + , k 2 ≥1, more preferably 2.
[0027] In another embodiment, if s 1 Less than or equal to the first threshold k 1 and s 2 is less than or equal to the second threshold, and z i Collect data for the first measurement, v for z i If the measurement acquisition data field is missing in v=f(v 1 , v 2 ) ,Right now v yes v 1 and v 2 The function of v , v 1 and v 2 The records are different, but the fields are the same, and , j ∈[1, n 1 ], n 1 is the number of first measurement data in the data set, t 1 and t 2 They are v 1 and v 2 The timestamp of the record. t j is the first j The timestamp of the first measurement data.
[0028] Using the method of collecting the mean of related measurement data to fill in missing data can improve the reliability of data filling by adjusting the number of randomly selected data; using related measurement data to directly fill in missing data can improve the efficiency of data filling, reduce the complexity of data processing, and provide the possibility for the application of data processing methods in a big data environment.
[0029] like Figure 4 As shown, according to the second aspect of the present invention, a data processing method based on a flag bit is further provided, wherein step 1000 is the same as the first aspect of the present invention, and steps 2000 and 3000 of the first aspect of the present invention are replaced by:
[0030] Step 2000: Obtain the degree of data missing according to the flag bit of the data.
[0031] Furthermore, the degree of missing ,in m 1 is the number of all flag bits of the data whose value is 1, m 0 is the number of all flag bits of the data whose value is 0, α , β is the weighting coefficient, α + β=1 , preferably α , β=0.5 .
[0032] Use the missing degree s To measure the data missing situation, it can judge the data missing situation more quickly and accurately, and can indicate the direction for the next step of data processing. At the same time, the weighting coefficient can be adjusted according to different data situations to improve the adaptability of the data processing method, so that the data processing method of the present invention can be more widely used.
[0033] Step 3000: clean the data according to the degree of missingness and data type of the data.
[0034] Furthermore, if s Greater than the third threshold k 3 , discard the data.
[0035] if s Less than or equal to the third threshold k 3 ,and z i Collect data for the first measurement, v for z i If the measurement acquisition data field is missing in v=f(v1 , v 2 ) ,Right now v yes v 1 and v 2 The function of v , v 1 and v 2 The records are different, but the fields are the same, and , j ∈[1, n 1 ], n 1 is the number of first measurement data in the data set, t 1 and t 2 They are v 1 and v 2 The timestamp of the record. t j is the first j The timestamp of the first measurement data.
[0036] if s Less than or equal to the third threshold k 3 ,and z i To collect data for the second measurement, the following method is used z i Perform data cleaning:
[0037] if ,So ,if ,So ;if Missing, then ,and i ≠ j . k is a natural number, k∈[1,10], preferably 3. k 3 is a positive real number, k 3 ≥0.8, more preferably 1.5.
[0038] The advantage of using this data cleaning method is that it can eliminate more than 99% of abnormal data, ensuring the validity of the data. At the same time, it can fill in missing data without the need for data search, thereby improving the efficiency of data processing.
[0039] In addition, other implementations of the present invention are obvious to those skilled in the art based on the disclosed description of the present invention. The embodiments and / or various aspects of the embodiments may be used in the system and method of the present invention alone or in any combination. The description and the examples therein should be regarded as exemplary only, and the actual scope and spirit of the present invention are represented by the appended claims.
Claims
1. A data processing method based on a flag bit, characterized in that: include: Step 1000, obtain data set , , n For the dataset Z The number of data in for z i The j data fields, m for z i The number of data fields in for Flag bit, when the field When missing, is 0; when the field When complete, is 1; c i is the data type; t i for z i The timestamp of the measurement data; Step 2000: according to the flag bit of the data , obtain the degree of missing data; Degree of missing ,in m 1 For the value of 1 Number, m 0 The value is 0 Number, α , β is the weighting coefficient, ; Step 3000: based on the degree of missing data and data type c i , process the data; The step 3000 further comprises: if s Greater than the third threshold k 3 , discard the data; if s Less than or equal to the third threshold k 3 ,and c i instruct z i Collect data for the first measurement, then v=f(v 1 , v 2 ) ,Right now v yes v 1 and v 2 The function of v for z i The missing measurement acquisition data fields in ; where v , v 1 and v 2 The records are different, but the fields are the same, and , , n 1 is the number of first measurement data in the data set, t 1 and t 2 They are v 1 and v 2 The timestamp of the record in which it is located, t j is the first j The timestamp of the first measurement data; if s Less than or equal to the third threshold k 3 ,and c i instruct z i To collect data for the second measurement, the following method is used z i Perform data cleaning: if ,So ,if ,So ;if Missing, then ,and i ≠ j .
2. The data processing method based on flag bit according to claim 1, characterized in that: k 3 is a positive real number, k 3 ≥0.
8.
3. The data processing method based on flag bit according to claim 1, characterized in that: k is a natural number, k∈[1,10].
Citation Information
Patent Citations
Curve fitting-based time series data processing method
CN106844290A
Mass information rating method, equipment and system
CN107480249A