Code set extraction method, device, server and storage medium
By automatically determining the credibility using linear regression and logarithmic regression equation models during the code set extraction process, the problem of poor timeliness of code set updates is solved, and faster and more accurate code set extraction is achieved.
Patent Information
- Application Number
- CN202210995145.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-08-18
AI Technical Summary
In the prior art, the update timeliness of code sets are poor, which is mainly due to the long extraction time of relying on manual credibility setting, especially when metadata is updated.
By scanning metadata, marking the target fields, and automatically determining the credibility using linear regression and logarithmic regression equation models, it is directly stored in the code set library, reducing manual intervention.
Improve the update timeliness and extraction efficiency of the code set, ensuring timely update and accuracy of the code set.
Smart Images

Figure CN115291943B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present invention relate to the technical field of data management, and in particular, to a method, device, server and storage medium for extracting a code set. Background Art
[0002] Currently, the code set is usually based on business descriptions, and data managers analyze and sort out the code set list from database design documents, meta-databases, and data dictionaries. Establishing and timely maintaining and updating the code set is a long-term and cumbersome task, which is used to improve the availability and quality of reference data, and plays an important role in enterprises to improve the overall data quality and cope with the increasing business volume.
[0003] In the prior art, when extracting the code set, technical personnel with professional capabilities are required to manually mark the text data according to experience first, and then set a specific credibility for the marked text data for analysis. Finally, the code set is obtained and updated to the original code set library.
[0004] However, the inventors found that the prior art has at least the following technical problems: relying on technical personnel with professional capabilities to set the credibility according to experience to complete the extraction of the code set. In this process, due to the long time for manually determining the credibility, or due to the update of data fields, the originally manually set credibility needs to be reset, resulting in a poor timeliness of code set update. Summary of the Invention
[0005] The embodiments of the present application provide a method, device, server and storage medium for extracting a code set to solve the problem of poor timeliness of code set update.
[0006] In a first aspect, the present invention provides a method for extracting a code set, including:
[0007] Obtain metadata to be recognized, where the metadata to be recognized includes multiple fields, and each field includes multiple data values;
[0008] By scanning the metadata to be recognized, mark the target fields that match the preset information;
[0009] Scan the data values of the target fields, record the corresponding relationship between the number of scanned records and the number of distinct data values corresponding to the target fields, and store it in a log table; and after the scanning is completed, record the corresponding relationship between the field name, data value, and the number of occurrences of the data value of each target field, and store it in a statistical table;
[0010] According to the corresponding relationship between the number of all scanned records and the number of distinct data values of each target field, solve the linear regression equation to obtain a linear regression equation model;
[0011] If it is determined that the linear fitting degree of the linear regression equation model does not meet the preset conditions, then according to the corresponding relationship between the total number of scanned records and the number of distinct data values for each target field, a logarithmic regression equation is solved to obtain a logarithmic regression equation model;
[0012] If it is determined that the credibility of the logarithmic regression equation model is greater than the preset credibility threshold, then the target field corresponding to the logarithmic regression equation model is stored in the classification table of the code set library, and the field name, data value, and the number of occurrences of the data value corresponding to the target field are stored in the code value table of the code set library.
[0013] In a possible implementation manner of the present application, the code set extraction method further includes a process of determining whether the linear fitting degree of the linear regression equation model meets the preset conditions, as follows:
[0014] Obtain the total number of scanned records, the current number of scanned records for each target field, and the number of distinct data values;
[0015] Substitute the total number of scanned records and the current number of scanned records into the linear regression equation model to obtain the observed value of the number of distinct data values for each target field when the total number of scanned records is i and the preset default level number is j, and the first average value of the number of distinct data values for all target fields when the total number of scanned records is i, where i is less than or equal to the total number of scanned records, and j is a natural number greater than 0;
[0016] According to the preset default level number, the total number of scanned records for each target field, the observed value of the number of distinct values, and the first average value, determine the within-group error for each target field, and according to the within-group error and the preset default level number, determine the within-group mean square error for each target field;
[0017] According to the number of distinct data values for each target field, obtain the first overall average value of the number of distinct values for all target fields;
[0018] According to the first average value, the first overall average value, the total number of scanned records, and the preset default level number, determine the between-group error among the target fields, and according to the between-group error and the preset mode level number, determine the between-group mean square error among the target fields;
[0019] According to the within-group mean square error and the between-group mean square error, determine the test statistic;
[0020] If it is determined that the test statistic is greater than the preset test statistic, then determine that the linear fitting degree of the linear regression equation model does not meet the preset conditions; if it is determined that the test statistic is less than or equal to the preset test statistic, then determine that the linear fitting degree of the linear regression equation model meets the preset conditions.
[0021] In a possible implementation of the present application, the formula for determining the within-group error of each target field based on the preset default level number, the total number of scan records of each target field, the observed value of the number of deduplicated values, and the first average value, and determining the within-group mean square deviation of each target field based on the within-group error and the preset default level number is as follows:
[0022]
[0023] In the formula, k is the preset default level number, n is the total number of scan records, and X ji is the observed value of the number of deduplicated data values of each target field when the total number of scan records is i and the preset default level number is j, is the first average value of the number of deduplicated data values of all target fields when the total number of scan records is i, SSE is the within-group error of each target field, and MSE is the within-group mean square deviation of each target field;
[0024] The formula for determining the between-group error between each target field based on the first average value, the first overall average value, the total number of scan records, and the preset default level number, and determining the between-group mean square deviation between each target field based on the between-group error and the preset mode level number is as follows:
[0025]
[0026] In the formula, k is the preset default level number, n is the total number of scan records, and n i is the total number of all target fields, is the first average value of the number of deduplicated data values of all target fields when the total number of scan records is i, is the average value of the number of deduplicated values of all target fields, SSA is the between-group error between each target field, and MSA is the between-group mean square deviation between each target field;
[0027] The formula for determining the test statistic based on the within-group mean square deviation and the between-group mean square deviation is as follows:
[0028]
[0029] In the formula, k is the preset default level number, SSA is the between-group error between each target field, and F is the test statistic.
[0030] In a possible implementation of the present application, it further includes the process of obtaining the credibility of the logarithmic regression equation model as follows:
[0031] Obtain the total number of scan records, the current scanned number of each target field, and the number of deduplicated data values;
[0032] Substitute the total number of scanned records, the current number of scanned records for each target field, and the number of deduplicated data values into the logarithmic regression equation model to obtain the second overall average of the number of deduplicated data values for the target field and the number of deduplicated data values for all target fields when the total number of scanned records is i, where i is less than or equal to the total number of scanned records;
[0033] Determine the sum of squared residuals according to the total number of scanned records, the number of deduplicated data values for the target field when the total number of scanned records is i, and the predicted value of the number of preset deduplicated data values, and determine the total sum of squared deviations according to the total number of scanned records, the number of deduplicated data values for the target field when the total number of scanned records is i, and the second overall average;
[0034] Determine the credibility of the logarithmic regression equation model according to the sum of squared residuals and the total sum of squared deviations.
[0035] In a possible implementation manner of the present application, the formula for determining the sum of squared residuals according to the total number of scanned records, the number of deduplicated data values for the target field when the total number of scanned records is i, and the predicted value of the number of preset deduplicated data values, where i is less than or equal to the total number of scanned records, is:
[0036]
[0037] In the formula, n is the total number of scanned records, y i is the number of deduplicated data values for the target field when the total number of scanned records is i, is the predicted value of the number of preset deduplicated data values, and SSE R is the sum of squared residuals;
[0038] The formula for determining the total sum of squared deviations according to the total number of scanned records, the number of deduplicated data values for the target field when the total number of scanned records is i, and the second overall average is:
[0039]
[0040] In the formula, y i is the number of deduplicated data values for the target field when the total number of scanned records is i, is the second overall average, and SST is the total sum of squared deviations;
[0041] The formula for determining the credibility of the logarithmic regression equation model according to the sum of squared residuals and the total sum of squared deviations is:
[0042]
[0043] In the formula, SSE Ris the sum of squared residuals, SST is the total sum of squared deviations, and R is the confidence level of the logarithmic regression equation model.
[0044] In a possible implementation manner of the present application, according to the corresponding relationship between the number of all scanned records and the number of deduplicated data values of each target field, a linear regression equation is solved to obtain a linear regression equation model, and then:
[0045] If it is determined that the linear fitting degree of the linear regression equation model meets a preset condition,
[0046] then deposit the target field into the preselection table of the code set, and deposit the corresponding relationship between the name, data value, and the number of occurrences of each data value of the target field into the temporary statistical table of the code set;
[0047] If it is determined that the target field in the preselection table of the code set has a code, deposit the target field with a code in the preselection table of the code set into the classification table of the code set library, and deposit the corresponding relationship between the field name, data value, and the number of occurrences of the data value recorded in the temporary statistical table into the code value table of the code set library.
[0048] In a possible implementation manner of the present application, the code set extraction method further includes:
[0049] According to the data value verification range corresponding to each target field and each field name in the code value table of the code set library, determine the verification rule corresponding to each target field, and deposit the verification rule into the rule library;
[0050] Obtain target data during the data quality management process, and perform quality verification on each field in the target data according to the verification rule;
[0051] If it is determined that the field in the target data does not meet the verification rule, deposit the field that does not meet the verification rule into the statistical library as a field to be extracted;
[0052] Scan the field to be extracted. If it is determined that the field to be extracted has a code, deposit the field to be extracted with a code into the classification table of the code set library.
[0053] In a second aspect, the present application provides a code set extraction device, including:
[0054] An acquisition module, configured to acquire metadata to be recognized, where the metadata to be recognized includes multiple fields, and each field includes multiple data values;
[0055] A scanning module, configured to scan the metadata to be recognized and mark out target fields that match preset information;
[0056] The scanning module is further configured to scan the data values of the target fields, record the corresponding relationship between the number of scanned records and the number of deduplicated data values corresponding to the target fields, and store the relationship in a log table; and after the scanning is completed, record the corresponding relationship between the field names, data values, and the occurrence times of the data values of each target field, and store the relationship in a statistical table.
[0057] An operation module is configured to solve a linear regression equation according to the corresponding relationship between the number of all scanned records and the number of deduplicated data values of each target field, and obtain a linear regression equation model.
[0058] A judgment module is configured to determine that the linear fitting degree of the linear regression equation model does not meet a preset condition.
[0059] The operation module is further configured to solve a logarithmic regression equation according to the corresponding relationship between the number of all scanned records and the number of deduplicated data values of each target field, and obtain a logarithmic regression equation model.
[0060] The judgment module is further configured to determine that the credibility of the logarithmic regression equation model is greater than a preset credibility threshold.
[0061] A code set extraction module is configured to store the target fields corresponding to the logarithmic regression equation model in a classification table of a code set library, and store the field names, data values, and the occurrence times of the data values corresponding to the target fields in a code value table of the code set library.
[0062] In a third aspect, the present application provides a server, including: at least one processor and a memory;
[0063] The memory stores computer-executable instructions.
[0064] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor executes the code set extraction method according to any one of the possible implementation manners in the first aspect.
[0065] In a fourth aspect, a computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the code set extraction method according to any one of the possible implementation manners in the first aspect is implemented.
[0066] The present application provides a method, apparatus, server, and storage medium for code set extraction. By recording the information obtained by scanning target fields and storing it in a log table and a statistical table, and successively using a linear regression equation model fitting and a logarithmic regression equation fitting based on the information in the log table, a more accurate credibility is obtained. Moreover, a part of the target fields to be extracted as the code set is deleted during the two fittings, improving the efficiency of code set extraction for some business codes with a relatively fast change speed, and thus solving the problem of poor timeliness of code set updates. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0068] Figure 1 Schematic diagram of the application scenario of the code set extraction method provided by the embodiment of the present application;
[0069] Figure 2 Schematic diagram of the process of the code set extraction method provided by the embodiment of the present application;
[0070] Figure 3 Schematic diagram of the structure of a code set extraction apparatus further provided by the embodiment of the present application;
[0071] Figure 4 Schematic diagram of the hardware structure of a server provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0073] At present, many enterprises have accumulated a large amount of data during the business transformation process, and there will also be corresponding data dictionaries. These data dictionaries are used to describe data information. The code sets available for reference in the data dictionary often differ from the business data after transformation, and it is impossible to ensure that the code values in the code sets referenced between different businesses are consistent and up-to-date. Therefore, it is necessary to establish and maintain an updated code set. To update the code set, it is necessary to extract the code set from different data fields and store it in the original code set. This is an important task for improving the overall quality of data and enhancing the reference quality of data. In the prior art, the process of extracting the code set is as follows. First, data managers extract and identify text fields containing the code value range of the original code set from the meta-database, then scan the text fields to count the corresponding data values, and rely on business experience to manually set the credibility for the code set corresponding to each text field. Finally, the code set with high credibility is stored in the original code set library. Since each time the code set is extracted, technicians with professional capabilities need to rely on business experience to set the credibility, this will increase the time for code set extraction. Especially when new data fields appear in the metadata, it will take longer to extract the code set for the new data fields, resulting in a poor timeliness of code set update during the code set extraction process.
[0074] To solve the above technical problems, the inventor proposed the following idea for solving the problem: First, scan and mark the fields in the metadata, mark the fields that match the original code set, and then scan the marked fields. Then scan the marked fields and count the number of scanned records, the number of distinct data values, the data values recorded in the fields, and the number of times each data value appears. Then perform a homogeneity of variance test through a linear regression equation model and a credibility fitting through a logarithmic regression equation model to obtain real-time and more accurate credibility. After that, judge the relationship between the credibility and the preset credibility threshold, complete the extraction of the code set, update the original code set library, without the need for manual setting of credibility, and improve the timeliness of code set update.
[0075] Figure 1 The application scenario schematic diagram of the code set extraction method provided by the embodiment of the present application is as Figure 1 shown, including: a terminal 101 and a server 102.
[0076] Among them, the terminal 101 is used for data managers to view the code set library and display the metadata to be identified. The server 102 is used to obtain the metadata to be identified, store the corresponding fields in the metadata to the code set library according to the code set extraction method provided by this embodiment, and is also used to display the code set library on the terminal 101.
[0077] Figure 2 The flow schematic diagram of the code set extraction method provided by the embodiment of the present application. The execution subject of this embodiment can be Figure 1In the embodiment shown, the server 102 may also be related devices of other computers, and this embodiment does not make special restrictions on this.
[0078] As Figure 2 shown, the code set extraction method provided in this embodiment includes:
[0079] S201: Obtain the metadata to be recognized, where the metadata to be recognized includes multiple fields, and each field includes multiple data values.
[0080] In this embodiment, the fields included in the metadata to be recognized can be numeric fields, character fields, or text fields. For example, in this embodiment, the field is a numeric field, and the numeric field can be a short integer field, an integer field, a long integer field, or a floating-point field. The data values included in each field can be code values.
[0081] Exemplarily, Table 1 is an information table of the metadata to be recognized, and the header of this information table can include a serial number and the names of each field.
[0082]
[0083]
[0084] Table 1
[0085] As shown in Table 1, the metadata to be recognized includes field A, field B, field C, and field D, and the corresponding data values of each field such as A01, A02, BB1, C01, or D01, etc.
[0086] S202: By scanning the metadata to be recognized, mark the target fields that match the preset information.
[0087] In this embodiment, the preset information can be the code values or code value ranges in the code value table of the code set library.
[0088] Exemplarily, Table 2 is a preset information table, and this preset information table includes the preset information and the data values of the corresponding target fields.
[0089] Preset value of field A Preset value of field B A01 BB1 A02 BB2 A03 BB3 A04 BB4
[0090] Table 2
[0091] As shown in Table 2, the preset information includes A01, A02, A03, A04, BB1, BB2, BB3, and BB4. During the process of scanning the metadata to be recognized shown in Table 1, mark field A and field B that match the preset information shown in Table 2, and field C and field D will not be marked.
[0092] S203: Scan the data values of the target fields, record the correspondence between the number of scanned records and the number of distinct data values corresponding to the target fields, and store it in the log table; and after the scanning is completed, record the correspondence between the field names, data values, and the number of occurrences of the data values of each target field, and store it in the statistical table.
[0093] In this embodiment, the number of distinct data values is the number of all data values minus the number of repeatedly occurring data values. The number of scanned records is the number of data values in the scanned target fields. The log table can include a header and numerical data. Among them, the header of the log table can be the number of scanned records and the number of distinct records of each target field, and the numerical data can be natural numbers greater than 0. The statistical table can also include a header and numerical data. Among them, the header of the statistical table can be field, data value, and number of occurrences, and the numerical data of the statistical table can be the recorded numbers of the corresponding field names, specific values of the data values, and the number of occurrences of each specific value.
[0094] Exemplarily, Table 3 is the log table, and the headers of this log table are the number of scanned records, the number of distinct data values of field A, and the number of distinct data values of field B respectively.
[0095]
[0096] Table 3
[0097] Please refer to Table 1 and Table 3. As the data values of field A and field B are scanned, for each scanned data value, the number of scanned records in Table 3 increases by 1. When the number of scanned records corresponding to field A is 6, during the scanning process, A01, A02, and A03 all appear twice. Therefore, the number of distinct data values of field A in Table 3 is 3. Until the number of scanned records is 10, the data value A04 appears, and then the corresponding number of distinct data values also increases by 1 and becomes 4.
[0098] Table 4 is the updated statistical table after the scanning is completed. In this statistical table, the headers are field, data value, and number of occurrences of the data value respectively.
[0099] Field Data value Occurrence times A A01 3 A A02 3 A A03 3 A A04 8 A A05 1 A A06 1 A AA1 1 B BB1 3 B BB2 3 B BB3 3 B B04 8 B B05 1 B B06 1 B BA1 1
[0100] Table 4
[0101] As shown in Table 4, after scanning field A and field B, the number of occurrences of the data values statistically recorded during the scanning process will be stored in Table 4. For example, A01 appears 3 times and A04 appears 8 times.
[0102] S204: According to the correspondence between the number of all scanned records and the number of distinct data values of each target field, solve the linear regression equation to obtain the linear regression equation model.
[0103] In this embodiment, the least squares method can be used to solve the linear regression equation. The initial model of the linear regression equation is y = ax + b, where y is the number of distinct data values, x is the currently scanned number, and the corresponding relationship between the scanned record number and the number of distinct data values for each target field can be obtained from the log table.
[0104] Taking the log table corresponding to field A as an example, and continuing to refer to Table 3, when the currently scanned number is 10, it can be known that the corresponding x is 10 and y is 4. Similarly, many sets of x and y values can be obtained, and then the values of a and b in the initial model can be obtained according to the least squares method. Substituting the values of a and b into the initial model, the linear regression equation model can be obtained.
[0105] S205: If it is determined that the linear fitting degree of the linear regression equation model does not meet the preset conditions, then according to the corresponding relationship between the scanned record numbers and the number of distinct data values of all target fields, a logarithmic regression equation is solved to obtain a logarithmic regression equation model.
[0106] In this embodiment, the linear fitting degree is to test the obtained linear regression equation model and compare the coincidence degree between the prediction result and the corresponding relationship between the actual scanned record number and the number of distinct data values.
[0107] Specifically, in an optional embodiment of the present application, the process of determining whether the linear fitting degree of the linear regression equation model meets the preset conditions includes the following steps:
[0108] S205a: Obtain the total number of scanned records, the currently scanned number of each target field, and the number of distinct data values.
[0109] In this embodiment, the total number of scanned records is the union of the sets of data values when scanning each target field during the entire code set extraction process. The currently scanned number and the number of distinct data values of each target field can be obtained from the statistical table in step S203. Exemplarily, continuing to refer to Table 3, when the 10th value of field A is scanned, the total number of scanned records is 20, the currently scanned number is 10, and the number of distinct data values is 4.
[0110] S205b: Substitute the total number of scanned records and the currently scanned number into the linear regression equation model to obtain the observed value of the number of distinct data values of each target field when the total number of scanned records is i and the preset default level number is j, and the first average value of the number of distinct data values of all target fields when the total number of scanned records is i, where i is less than or equal to the total number of scanned records, and j is a natural number greater than 0.
[0111] In this embodiment, the level is the intensity of the factors in the linear regression equation model, and the preset default level quantity is the default value of the quantity of levels preset manually. For example, for the factor affecting the number of deduplicated data values, there is the quantity of data values of field A. If there are three levels of the data quantity of A, namely 100 data values, 10,000 data values, and 1,000,000 data values, then the preset default level quantity is said to be 3.
[0112] In this embodiment, the observed value of the number of deduplicated data values is the predicted value obtained based on the linear regression equation model. If the fitting degree of the linear regression equation is higher, then the predicted value is closer to the actual number of deduplicated data values recorded in the log table.
[0113] S205c: Determine the within-group error of each target field according to the preset default level quantity, the total number of scan records of each target field, the observed value of the number of deduplicated values, and the first average value, and determine the within-group mean square error of each target field according to the within-group error and the preset default level quantity.
[0114] In this embodiment, the calculation formula used in step S205c is as follows:
[0115]
[0116] In the formula, k is the preset default level quantity, n is the total number of scan records, X ji is the observed value of the number of deduplicated data values of each target field when the total number of scan records is i and the preset default level quantity is j, is the first average value of the number of deduplicated data values of all target fields when the total number of scan records is i, SSE is the within-group error of each target field, and MSE is the within-group mean square error of each target field.
[0117] S205d: Obtain the first overall average value of the number of deduplicated values of all target fields according to the number of deduplicated data values of each target field.
[0118] S205e: Determine the between-group error among the target fields according to the first average value, the first overall average value, the total number of scan records, and the preset default level number, and determine the between-group mean square error among the target fields according to the between-group error and the preset mode level number.
[0119] In an alternative embodiment of the present application, the calculation formula of step S205e is as follows:
[0120]
[0121] In the formula, k is the preset default level quantity, n is the total number of scan records, n i is the total number of all target fields, is the first average value of the number of distinct data values of all target fields when the total number of scanned records is i, is the average value of the number of distinct values of all target fields, SSA is the between-group error among the target fields, and MSA is the between-group mean square deviation among the target fields.
[0122] S205f: Determine the test statistic according to the within-group mean square deviation and the between-group mean square deviation.
[0123] In this embodiment, the formula used in step S205f is:
[0124]
[0125] In the formula, k is the preset default level number, SSA is the between-group error among the target fields, and F is the test statistic.
[0126] S205g: If it is determined that the test statistic is greater than the preset test statistic, it is determined that the linear fitting degree of the linear regression equation model does not meet the preset conditions; if it is determined that the test statistic is less than or equal to the preset test statistic, it is determined that the linear fitting degree of the linear regression equation model meets the preset conditions.
[0127] In this embodiment, the preset test statistic can be the critical value obtained by looking up the table, and this critical value can be F α , for example, the critical value F of F in (k - 1, nk - k) α , where k is the preset default level number, n is the total number of scanned records, and α is the confidence level set by the user. In this embodiment, α can be 95%. The specific value of the critical value F can be obtained by querying the F-test critical value table. If F is greater than F α , it is determined that the linear fitting degree of the linear regression equation model meets the preset conditions, otherwise it does not. α
[0128] The above steps S205a to S205g elaborate in detail the process of determining whether the fitting degree of the linear regression equation model meets the preset conditions. If it is determined that the linear fitting degree of the linear regression equation model does not meet the preset conditions, then according to the corresponding relationship between the number of all scanned records and the number of distinct data values of each target field, a logarithmic regression equation is solved to obtain a logarithmic regression equation model.
[0129] In this embodiment, the least squares method can be used to solve the logarithmic regression equation. The specific process of the solution is as follows: Obtain the initial model of the logarithmic regression equation as y = c ln x + d, where the parameters c and d are the parameters to be solved, x is the currently scanned number of the target field corresponding to the log table, and y is the number of distinct data values. When solving the parameters c and d, lnx can be set as the intermediate variable z, and substitute z = lnx into the initial model of the logarithmic regression equation to obtain the intermediate model y = cz + b. Finally, use the least squares method to solve the intermediate model. When solving, the corresponding relationship between the scanned record number and the number of distinct data values of each target field can be obtained from the log table.
[0130] The above steps solve the logarithmic regression equation model and the target fields for which the code set needs to be extracted. Next, the code set is extracted for the target fields.
[0131] S206: If it is determined that the credibility of the logarithmic regression equation model is greater than the preset credibility threshold, then store the target fields corresponding to the logarithmic regression equation model in the classification table of the code set library, and store the field name, data value, and the number of occurrences of the data value corresponding to the target field in the code value table of the code set library.
[0132] In this embodiment, the credibility is an index representing the fitting degree of the logarithmic regression equation model. The preset credibility threshold can be set manually in advance, or can be a credibility threshold randomly selected from a threshold range pre-stored in the server. The preset credibility threshold can be set as β, and β ∈ (0, 1).
[0133] In this embodiment, if it is determined that the credibility of the logarithmic regression equation model is greater than the preset credibility threshold, then store the target fields corresponding to the logarithmic regression equation model in the classification table of the code set library. The classification table can include the names of each field. Each field corresponds to a part of the data values in the code value table of the code set library. Please continue to refer to Table 1. There are new data values A05 and A06 in field A. At this time, store field A and the corresponding data values in the code set library to obtain a new code set library. The data values corresponding to field A in the updated code set library change from the original A01, A02, A03, A04 to A01, A02, A03, A04, A05, and A06, thus completing the extraction of the code set.
[0134] Specifically, in an alternative embodiment of the present application, the process of obtaining the credibility of the logarithmic regression equation model includes the following steps:
[0135] S206a: Obtain the total number of scanned records, the currently scanned number of each target field, and the number of distinct data values.
[0136] S206b: Substitute the total number of scanned records, the current number of scanned records for each target field, and the number of deduplicated data values into the logarithmic regression equation model to obtain the number of deduplicated data values for the target field and the second overall average of the number of deduplicated data values for all target fields when the total number of scanned records is i, where i is less than or equal to the total number of scanned records.
[0137] S206c: Determine the sum of squared residuals based on the total number of scanned records, the number of deduplicated data values for the target field when the total number of scanned records is i, and the predicted value of the number of deduplicated data values, and determine the total sum of squared deviations based on the total number of scanned records, the number of deduplicated data values for the target field when the total number of scanned records is i, and the second overall average.
[0138] S206d: Determine the credibility of the logarithmic regression equation model based on the sum of squared residuals and the total sum of squared deviations.
[0139] In an alternative embodiment of the present application, the calculation formula for determining the sum of squared residuals in step S206b is:
[0140]
[0141] In the formula, n is the total number of scanned records, y i is the number of deduplicated data values for the target field when the total number of scanned records is i, is the predicted value of the number of deduplicated data values, and SSE R is the sum of squared residuals.
[0142] The calculation formula for determining the total sum of squared deviations in step S206c is:
[0143]
[0144] In the formula, y i is the number of deduplicated data values for the target field when the total number of scanned records is i, is the second overall average, and SST is the total sum of squared deviations.
[0145] The calculation formula for determining the credibility of the logarithmic regression equation model in step S206d is:
[0146]
[0147] In the formula, SSE R is the sum of squared residuals, SST is the total sum of squared deviations, and R 2 is the credibility of the logarithmic regression equation model.
[0148] In summary, for a code set extraction method provided by an embodiment of the present application, first, by scanning the metadata to be recognized, target fields are marked, and then the target fields are scanned to obtain the corresponding relationship between the target fields where codes may exist and the number of distinct data values, and a linear regression equation model is solved. Then, the linear regression equation model is fitted to determine whether there are target fields in each target field that can be extracted into a code set. After that, a logarithmic regression equation model is solved and the logarithmic regression equation model is fitted to obtain the exact confidence level R 2 , through the confidence level R 2 is compared with a preset confidence threshold value, and the final corresponding target field is stored in the classification table of the code set library, and the corresponding data value is used as the code value table in the code set library. During the entire process of code set extraction, there is no need to manually set the confidence level, thereby improving the update timeliness of the code set
[0149] Meanwhile, by fitting the linear regression equation model, the target fields containing new data values can be quickly obtained, reducing the time spent on screening the target fields that can be selected into the code set library, and thus improving the speed of code set extraction.
[0150] Meanwhile, by fitting the logarithmic regression equation model, a real-time and more accurate confidence level is calculated, and the target fields with a confidence level greater than the preset confidence threshold value are stored in the classification table of the code set library. Otherwise, they are stored in a pre-table for subsequent manual screening and code set extraction. The accuracy of code set extraction is further improved.
[0151] In an alternative embodiment of the present application, the code set extraction method further includes the following steps:
[0152] Step a: If it is determined that the linear fitting degree of the linear regression equation model meets the preset conditions, the target field is stored in the code set preselection table, and the corresponding relationship between the name of the target field, the data value, and the number of occurrences of each data value is stored in the code set temporary statistical table.
[0153] Step b: If it is determined that there are codes in the target fields in the code set preselection table, the target fields with codes in the code set preselection table are stored in the classification table of the code set library, and the corresponding relationship between the field name, data value, and the number of occurrences of the data value of each target field recorded in the temporary statistical table is stored in the code value table of the code set library.
[0154] In this embodiment, when the linear fitting degree of the linear regression equation model meets the preset conditions, it indicates that there may be no data values containing new codes in the target fields. The code set temporary statistical table is used to temporarily store the target fields.
[0155] Exemplarily, continue to refer to Table 1. One of the data values in Field A is AA1. Although AA1 cannot be used as a code in the current situation and is stored in the code set library along with Field A, as the business progresses and is updated, the data value AA1, as a new code, will appear more frequently. At this time, the fields containing AA1 in the code set temporary statistical table will gradually increase. As the number of fields containing AA1 increases, data managers can more easily discover the new code AA1 and then store the fields containing the data value AA1 in the code set library, thus completing the extraction process of the code set.
[0156] In summary, in this embodiment, by storing the fields that cannot be stored in the code set library in the code set temporary statistical table as a source of a new code set to be extracted, the extraction and update of the code set are made faster and more timely, avoiding missing the fields that can be stored in the code set library.
[0157] In an alternative embodiment of the present application, the code set extraction method further includes the following steps:
[0158] Step d: Determine the verification rule corresponding to each target field according to the data value verification range corresponding to each target field in the code value table of the code set library and each field name, and store the verification rule in the rule library.
[0159] Step e: Obtain the target data during the data quality management process, and perform quality verification on each field in the target data according to the verification rule.
[0160] Step f: If it is determined that the field in the target data does not meet the verification rule, then store the field that does not meet the verification rule in the statistical library as a field to be extracted.
[0161] Step g: Scan the fields to be extracted. If it is determined that there is a code in the field to be extracted, then store the field to be extracted with a code in the classification table of the code set library.
[0162] In this embodiment, the verification rule is the verification standard for whether the requirements of data verification are met during the data management process. The rule library is a table, storage space, or memory for storing various data quality verification rules. Quality verification is to judge the integrity, standardization, consistency, accuracy, relevance, or uniqueness of data according to the corresponding verification rule. The statistical library can be a table for statistically storing various types of data.
[0163] Exemplarily, during the data quality management process, it is found that the new Field C does not meet the verification rule. Therefore, it is determined that the data values C01 to C04 recorded in Field C are not within the corresponding code value range in the code set library. Then, Field C is directly excluded from the beginning stage of the extraction code set, and the excluded Field C is stored in the statistical table as a field to be extracted for manual extraction of the code set.
[0164] In summary, by storing the fields that do not meet the verification rules and require manual confirmation or update in the statistics database for manual extraction of the code set, the statistics database can serve as a potential source for code set extraction, further improving the timeliness of code set extraction. When the business code is updated and the data management personnel fail to promptly learn about the new code information, the new code can be used as a potential code for manual extraction and storage in the code set library.
[0165] Figure 3 The embodiment of the present application further provides a structural schematic diagram of a code set extraction device, as Figure 3 shown. The device includes: an acquisition module 31, a scanning module 32, an operation module 33, a judgment module 34, and a code set extraction module 35.
[0166] Specifically, the acquisition module 31 is used to acquire the metadata to be recognized, where the metadata to be recognized includes multiple fields, and each field includes multiple data values.
[0167] The scanning module 32 is used to scan the metadata to be recognized and mark the target fields that match the preset information. The scanning module 32 is also used to scan the data values of the target fields, record the correspondence between the number of scanned records and the number of distinct data values corresponding to the target fields, and store them in the log table. After the scanning is completed, record the correspondence between the field names, data values, and the number of occurrences of the data values of each target field and store them in the statistical table.
[0168] The operation module 33 is used to solve the linear regression equation according to the correspondence between the number of scanned records and the number of distinct data values of each target field to obtain a linear regression equation model.
[0169] The judgment module 34 is used to determine that the linear fitting degree of the linear regression equation model does not meet the preset conditions.
[0170] The operation module 33 is also used to solve the logarithmic regression equation according to the correspondence between the number of scanned records and the number of distinct data values of each target field to obtain a logarithmic regression equation model.
[0171] The judgment module 34 is also used to determine that the credibility of the logarithmic regression equation model is greater than the preset credibility threshold.
[0172] The code set extraction module 35 is used to store the target fields corresponding to the logarithmic regression equation model in the classification table of the code set library, and store the field names, data values, and the number of occurrences of the data values corresponding to the target fields in the code value table of the code set library.
[0173] In an alternative embodiment of the present application, the acquisition module 31 is further used to acquire the total number of scanned records, the current number of scanned records of each target field, and the number of distinct data values.
[0174] The operation module 33 is specifically configured to substitute the total number of scan records and the currently scanned number into the linear regression equation model to obtain the observed value of the number of deduplicated data values of each target field when the total number of scan records is i and the preset default level number is j, and the first average value of the number of deduplicated data values of all target fields when the total number of scan records is i, where i is less than or equal to the total number of scan records, and j is a natural number greater than 0.
[0175] The operation module 33 is also specifically configured to determine the within-group error of each target field according to the preset default level number, the total number of scan records of each target field, the observed value of the number of deduplicated values, and the first average value, and determine the within-group mean square error of each target field according to the within-group error and the preset default level number. According to the number of deduplicated data values of each target field, the first overall average value of the number of deduplicated values of all target fields is obtained. According to the first average value, the first overall average value, the total number of scan records, and the preset default level number, the between-group error between target fields is determined, and the between-group mean square error between target fields is determined according to the between-group error and the preset mode level number. According to the within-group mean square error and the between-group mean square error, the test statistic is determined.
[0176] The judgment module 34 is specifically configured to, if it is determined that the test statistic is greater than the preset test statistic, determine that the linear fitting degree of the linear regression equation model does not meet the preset conditions; if it is determined that the test statistic is less than or equal to the preset test statistic, determine that the linear fitting degree of the linear regression equation model meets the preset conditions.
[0177] In an optional embodiment of the present application, the calculation formulas used by the operation module 33 to determine the within-group error and the within-group mean square error of each target field are:
[0178]
[0179] where k is the preset default level number, n is the total number of scan records, X ji is the observed value of the number of deduplicated data values of each target field when the total number of scan records is i and the preset default level number is j, is the first average value of the number of deduplicated data values of all target fields when the total number of scan records is i, SSE is the within-group error of each target field, and MSE is the within-group mean square error of each target field.
[0180] The operation module 33 uses the following calculation formulas to determine the between-group error between target fields and the between-group mean square error between target fields:
[0181]
[0182] where k is the preset default level number, n is the total number of scan records, n iis the total number of all target fields, is the first average value of the number of distinct data values of all target fields when the total number of scanned records is i, is the average value of the number of distinct values of all target fields, SSA is the between-group error among target fields, and MSA is the between-group mean square deviation among target fields.
[0183] The operation module 33 is specifically configured to determine a test statistic according to the within-group mean square deviation and the between-group mean square deviation, and the calculation formula is:
[0184]
[0185] In the formula, k is the preset default level number, SSA is the between-group error among target fields, and F is the test statistic.
[0186] In an alternative embodiment of the present application, the acquisition module 31 is further specifically configured to acquire the total number of scanned records, the current number of scanned records of each target field, and the number of distinct data values.
[0187] The operation module 33 is further specifically configured to substitute the total number of scanned records, the current number of scanned records of each target field, and the number of distinct data values into the logarithmic regression equation model to obtain the number of distinct data values of the target field and the second overall average value of the number of distinct data values of all target fields when the total number of scanned records is i, where i is less than or equal to the total number of scanned records. It is further specifically configured to determine the sum of squared residuals according to the total number of scanned records, the number of distinct data values of the target field when the total number of scanned records is i, and the predicted value of the number of distinct data values, and determine the total sum of squared deviations according to the total number of scanned records, the number of distinct data values of the target field when the total number of scanned records is i, and the second overall average value. Determine the credibility of the logarithmic regression equation model according to the sum of squared residuals and the total sum of squared deviations.
[0188] In an alternative embodiment of the present application, the calculation formula used by the operation module 33 to determine the sum of squared residuals is:
[0189]
[0190] In the formula, n is the total number of scanned records, y i is the number of distinct data values of the target field when the total number of scanned records is i, is the predicted value of the number of distinct data values, and SSE R is the sum of squared residuals.
[0191] The operation module 33 is further specifically configured to determine the total sum of squared deviations according to the total number of scanned records, the number of distinct data values of the target field when the total number of scanned records is i, and the second overall average value, and the calculation formula is:
[0192]
[0193] In the formula, y i is the number of deduplicated data values when the total number of target field scan records is i, is the second overall average value, and SST is the total sum of squared deviations;
[0194] The calculation formula used by the operation module 33 to determine the credibility of the logarithmic regression equation model is:
[0195]
[0196] In the formula, SSE R is the sum of squared residuals, SST is the total sum of squared deviations, and R is the credibility of the logarithmic regression equation model.
[0197] In an optional embodiment of the present application, when the judgment module 34 determines that the linear fitting degree of the linear regression equation model meets the preset conditions. The scanning module 32 is further specifically configured to store the target field in the preselection table of the code set, and store the corresponding relationship between the name, data value, and the number of occurrences of each data value of the target field in the temporary statistical table of the code set.
[0198] When the judgment module 34 determines that there is a code for the target field in the preselection table of the code set. The scanning module 32 is further specifically configured to store the target field with a code in the classification table of the code set library, and store the corresponding relationship between the field name, data value, and the number of occurrences of the data value of each target field recorded in the temporary statistical table in the code value table of the code set library.
[0199] In an optional embodiment of the present application, the code set extraction device further includes a quality rule generation module 36, which is used to determine the verification rule corresponding to each target field according to the data value verification range corresponding to each target field in the code value table of the code set library and each field name, and store the verification rule in the rule library.
[0200] The acquisition module 31 is further used to acquire target data during the data quality management process, and perform quality verification on each field in the target data according to the verification rules.
[0201] The judgment module 34 is further specifically configured to determine that the field in the target data does not meet the verification rule, and store the field that does not meet the verification rule in the statistical library as a field to be extracted.
[0202] The scanning module 32 is further specifically configured to scan the field to be extracted. If it is determined that the field to be extracted has a code, the field to be extracted with a code is stored in the classification table of the code set library.
[0203] Continue to refer to Figure 3, in an optional embodiment of the present application, the code set extraction device further includes a code set management module 37 and a potential code value identification module 38.
[0204] Among them, the code set management module 37 is used to manage the classification table and code value table of the code set library, associate and process the fields with mappings and the data values of the code set in the corresponding fields. The code set management module 37 is also used to provide a code set library management function to confirm and maintain the fields with existing code sets, and provide a data value management function to confirm and maintain the code values of the fields.
[0205] The potential code value identification module 38 is used to scan and count the data values that appear less frequently in the corresponding fields in the statistical library for the data manager to confirm whether they are code sets, or to scan the potential code sets and newly added data values of the fields in the incremental data records for the data manager to confirm whether they are code sets.
[0206] Figure 4 Schematic diagram of the hardware structure of a server provided by an embodiment of the present application, as Figure 4 shown, the server includes: at least one processor 41 and a memory 42.
[0207] Among them, the memory 42 is used to store computer execution instructions. For specific details, reference can be made to the relevant descriptions in the foregoing method embodiments.
[0208] Optionally, the memory 42 can be either independent or integrated with the processor 41.
[0209] When the memory 42 is independently provided, the server further includes a bus 43 for connecting the memory 42 and the processor 41.
[0210] An embodiment of the present application also provides a computer-readable storage medium, in which computer-executable instructions are stored. When the processor executes the computer-executable instructions, the code set extraction method in the foregoing method embodiments is implemented.
[0211] An embodiment of the present invention also provides a computer program product, including a computer program, which when executed by a processor, implements the code set extraction method as described above.
[0212] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the above-mentioned modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.
[0213] The modules described above as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.
[0214] In addition, each functional module in various embodiments of the present invention can be integrated in a processing unit, or each module can exist physically alone, or two or more modules can be integrated in one unit. The unit formed by the above-mentioned modules can be implemented in the form of hardware or in the form of a hardware plus software functional unit.
[0215] The integrated module implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module stored in a storage medium includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods in various embodiments of the present application.
[0216] It should be understood that the above-mentioned processor can be a central processing unit (Central Processing Unit, abbreviated as CPU), and can also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the invention can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.
[0217] The memory may include high-speed RAM memory, and may also include non-volatile storage NVM, such as at least one disk memory, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disc, etc.
[0218] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience in representation, the buses in the drawings of this application are not limited to only one bus or one type of bus.
[0219] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk or an optical disc. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0220] An exemplary storage medium is coupled to the processor, enabling the processor to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an Application Specific Integrated Circuits (ASIC). Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master control device.
[0221] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes: various media such as ROM, RAM, magnetic disks, or optical discs that can store program codes.
[0222] to illustrate the technical solutions of this application, rather than to limit them; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A method for extracting a code set, characterized in that, Including: Obtain the metadata to be recognized, where the metadata to be recognized includes multiple fields, and each field includes multiple data values; By scanning the metadata to be recognized, mark the target fields that match the preset information, where the preset information is a code value or a code value range in the code value table of the code set library; Scan the data values of the target fields, record the corresponding relationship between the number of scanned records and the number of distinct data values corresponding to the target fields, and store it in the log table; and after the scanning is completed, record the corresponding relationship between the field name, data value, and the number of occurrences of the data value of each target field, and store it in the statistical table; Solve the linear regression equation according to the corresponding relationship between the number of all scanned records and the number of distinct data values of each target field to obtain a linear regression equation model; If it is determined that the linear fitting degree of the linear regression equation model does not meet the preset conditions, then solve the logarithmic regression equation according to the corresponding relationship between the number of all scanned records and the number of distinct data values of each target field to obtain a logarithmic regression equation model; If it is determined that the credibility of the logarithmic regression equation model is greater than the preset credibility threshold, then store the target fields corresponding to the logarithmic regression equation model in the classification table of the code set library, and store the field name, data value, and the number of occurrences of the data value corresponding to the target field in the code value table of the code set library; If it is determined that the linear fitting degree of the linear regression equation model meets the preset conditions, then store the target fields in the code set preselection table, and store the corresponding relationship between the name, data value, and the number of occurrences of each data value corresponding to the target field in the code set temporary statistical table; If it is determined that there is a code for the target field in the code set preselection table, then store the target fields with codes in the code set preselection table in the classification table of the code set library, and store the corresponding relationship between the field name, data value, and the number of occurrences of the data value of each target field recorded in the temporary statistical table in the code value table of the code set library.
2. The method according to claim 1, wherein It also includes a process of judging whether the linear fitting degree of the linear regression equation model meets the preset conditions, as follows: Obtain the total number of scanned records, the current number of scanned records of each target field, and the number of distinct data values; Substitute the total number of scanned records and the current number of scanned records into the linear regression equation model to obtain the observed value of the number of distinct data values of each target field when the total number of scanned records is i and the preset default level quantity is j, and the first average value of the number of distinct data values of all target fields when the total number of scanned records is i, where i is less than or equal to the total number of scanned records, j is a natural number greater than 0, the level is the intensity of the factor affecting the linear regression equation model, and the preset default level quantity is the default value of the quantity of the level preset manually; Determine the within-group error of each target field according to the preset default level quantity, the total number of scanned records of each target field, the observed value of the number of distinct data values, and the first average value, and determine the within-group mean square deviation of each target field according to the within-group error and the preset default level quantity; Obtain the first overall average of the number of distinct data values for all target fields based on the number of distinct data values for each target field. Determine the between-group error among the target fields according to the first average, the first overall average, the total number of scan records, and the preset default level number, and determine the between-group mean square deviation among the target fields according to the between-group error and the preset model level number. Determine the test statistic according to the within-group mean square deviation and the between-group mean square deviation. If it is determined that the test statistic is greater than the preset test statistic, it is determined that the linear fitting degree of the linear regression equation model does not meet the preset conditions; if it is determined that the test statistic is less than or equal to the preset test statistic, it is determined that the linear fitting degree of the linear regression equation model meets the preset conditions.
3. The method according to claim 2, wherein The formula for determining the test statistic according to the within-group mean square deviation and the between-group mean square deviation is as follows: In the formula, k is the preset default level quantity, SSA is the between-group error among the target fields, and F is the test statistic.
4. The method according to claim 1, characterized in that, It also includes the process of obtaining the credibility of the logarithmic regression equation model, as follows: Obtain the total number of scan records, the current number of scanned records for each target field, and the number of distinct data values. Substitute the total number of scan records, the current number of scanned records for each target field, and the number of distinct data values into the logarithmic regression equation model to obtain the number of distinct data values of the target field and the second overall average of the number of distinct data values of all target fields when the total number of scan records is i, where i is less than or equal to the total number of scan records. Determine the sum of squared residuals according to the total number of scan records, the number of distinct data values of the target field when the total number of scan records is i, and the predicted value of the preset number of distinct data values, and determine the total sum of squared deviations according to the total number of scan records, the number of distinct data values of the target field when the total number of scan records is i, and the second overall average. Determine the credibility of the logarithmic regression equation model according to the sum of squared residuals and the total sum of squared deviations.
5. The method according to claim 4, wherein The formula for determining the sum of squared residuals according to the total number of scan records, the number of distinct data values of the target field when the total number of scan records is i, and the predicted value of the preset number of distinct data values, where i is less than or equal to the total number of scan records is: Where n is the total number of scan records, is the number of distinct data values of the target field when the total number of scan records is i, is the predicted value of the number of distinct data values preset, is the sum of squared residuals; The formula for determining the total sum of squared deviations according to the total number of scan records, the number of distinct data values of the target field when the total number of scan records is i, and the second overall average is: Wherein, is the number of distinct data values of the target field when the total number of scan records is i, is the second overall average value, and SST is the total sum of squared deviations; The formula for determining the credibility of the logarithmic regression equation model according to the sum of squared residuals and the total sum of squared deviations is: Wherein, is the sum of squared residuals, SST is the total sum of squared deviations, and R is the credibility of the logarithmic regression equation model.
6. The method according to any one of claims 1 to 5, characterized in that It also includes: Determine the verification rule corresponding to each target field according to the data value verification range corresponding to each target field in the code value table of the code set library and each field name, and store the verification rule in the rule library. Obtain the target data during the data quality management process, and perform quality verification on each field in the target data according to the verification rule. If it is determined that the fields in the target data do not meet the verification rules, the fields that do not meet the verification rules are stored in the statistical database as the fields to be extracted; Scan the fields to be extracted. If it is determined that there is a code in the fields to be extracted, the fields to be extracted with codes are stored in the classification table of the code set library.
7. A code set extraction device, characterized in that, Including: An acquisition module for acquiring metadata to be recognized, where the metadata to be recognized includes multiple fields, and each field includes multiple data values; A scanning module for scanning the metadata to be recognized and marking the target fields that match the preset information, where the preset information is a code value or a code value range in the code value table of the code set library; The scanning module is further configured to scan the data values of the target fields, record the corresponding relationship between the number of scanned records and the number of distinct data values corresponding to the target fields, and store it in the log table; and after the scanning is completed, record the corresponding relationship between the field names, data values, and the number of occurrences of the data values of each target field, and store it in the statistical table; An operation module for solving a linear regression equation according to the corresponding relationship between the number of scanned records and the number of distinct data values of each target field to obtain a linear regression equation model; A judgment module for determining that the linear fitting degree of the linear regression equation model does not meet the preset conditions; The operation module is further configured to solve a logarithmic regression equation according to the corresponding relationship between the number of scanned records and the number of distinct data values of each target field to obtain a logarithmic regression equation model; The judgment module is further configured to determine that the credibility of the logarithmic regression equation model is greater than a preset credibility threshold; A code set extraction module for storing the target fields corresponding to the logarithmic regression equation model in the classification table of the code set library, and storing the corresponding relationship between the field names, data values, and the number of occurrences of the data values of the target fields in the code value table of the code set library; If it is determined that the linear fitting degree of the linear regression equation model meets the preset conditions, the target fields are stored in the code set preselection table, and the corresponding relationship between the names, data values, and the number of occurrences of each data value of the target fields is stored in the code set temporary statistical table; If it is determined that there is a code in the target fields in the code set preselection table, the target fields with codes in the code set preselection table are stored in the classification table of the code set library, and the corresponding relationship between the field names, data values, and the number of occurrences of the data values of each target field recorded in the temporary statistical table is stored in the code value table of the code set library.
8. A server, characterized in that, Including: At least one processor and a memory; The memory stores computer execution instructions; The at least one processor executes the computer execution instructions stored in the memory, so that the at least one processor executes the code set extraction method according to any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer execution instructions, and when the processor executes the computer execution instructions, the code set extraction method according to any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Method for managing development jobs in development environment, equipment and program product
CN113032004A
Data field configuration method and related device
CN114357006A