A method and system for determining dimensional encoding
Through the determination of multiple matching methods and coding length levels, the problems of low accuracy and low efficiency of dimensional coding matching in the prior art are solved, and more efficient and accurate dimensional coding matching is achieved, and the dependence of manual verification is avoided.
Patent Information
- Application Number
- CN202211661775.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-12-23
AI Technical Summary
When performing dimension coding matching, the accuracy of the prior art is not high, and the dimension of the same name exists under the multi-level coding system cannot be identified, and the efficiency is low, and it relies on manual verification and encoding, occupying human resources.
A method for determining dimension encoding is provided. By obtaining the data to be matched, using various matching methods such as name matching, empirical matching and similarity matching, the initial dimension encoding is obtained, and the target dimension encoding is determined from multiple initial dimension encodings based on the encoding length and node level.
The accuracy of dimensional coding matching is improved, and the problem of the same name dimension under the multi-level coding system is avoided. Compared with manual verification and encoding, it is more efficient and does not occupy human resources.
Smart Images

Figure CN115827742B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method and system for determining dimension codes. Background Art
[0002] When technologies such as data warehouses, data mining, and data visualization are applied in the budget networking supervision system, there is a necessary prerequisite: the dimension codes in the business data table must exist. Since there are two sources of data for the budget networking supervision system: one is a text file, and the other is a database intermediate table. The data extracted in the form of a text file does not have dimension codes, only dimension names; therefore, after the business system imports the data into the target database through ELT (Extract-Transform-Load), during further analysis and processing, it is necessary to identify the dimension names and supplement the dimension codes according to national standards or local actual situations to facilitate data mining and data visualization operations.
[0003] Currently, there are mainly the following several operations for dimension code matching (i.e., supplementing dimension codes):
[0004] 1. Using the historical data existing in the database table to update the dimension codes of the currently newly imported data;
[0005] 2. Using the dimension names in the dimension table to compare with the dimension names of the data to be matched to obtain the dimension codes;
[0006] 3. Manually checking the codes.
[0007] The first two methods both compare the dimension names in the dimension table by name or serial number, so the matching accuracy is not high, and it is impossible to identify the situation of the same dimension name existing in a multi-level coding system (that is, the situation where there are the same dimension names but different dimension codes under different major categories). The third method is inefficient and consumes human resources. Summary of the Invention
[0008] To at least overcome the problems in the related art that the accuracy of dimension code matching is not high, the efficiency is low, and it is impossible to identify the same dimension name existing in a multi-level coding system, this application provides a method and system for determining dimension codes.
[0009] The solution of this application is as follows:
[0010] According to the first aspect of the embodiments of this application, a method for determining dimension codes is provided, including:
[0011] Obtaining the data to be matched;
[0012] Using at least one matching method to match and obtain the initial dimension codes of the data to be matched;
[0013] In the case where the initial dimension encodings are multiple, determine the target dimension encoding corresponding to the data to be matched from the multiple initial dimension encodings according to the encoding length of the initial dimension encoding and the preset node level corresponding to the initial dimension encoding.
[0014] Preferably, the matching methods at least include:
[0015] Name matching, experience matching, and similarity matching;
[0016] The name matching includes: obtaining the dimension name of the data to be matched;
[0017] Query the dimension encoding corresponding to the dimension name of the data to be matched in the pre-constructed dimension table;
[0018] The experience matching includes: obtaining the dimension name and dimension identifier of the data to be matched;
[0019] Query the dimension encoding corresponding to the dimension name and dimension identifier of the data to be matched in the pre-constructed experience library;
[0020] The similarity matching includes: obtaining the dimension name of the data to be matched;
[0021] Split and reconstruct the dimension name of the data to be matched to obtain a split and reconstructed name;
[0022] Query the dimension encoding corresponding to the dimension name with a similarity higher than a preset threshold to the split and reconstructed name of the data to be matched in the pre-constructed dimension table, and when there are multiple dimension encodings corresponding to the dimension names with a similarity higher than the preset threshold, select the dimension encoding corresponding to the dimension name with the highest similarity.
[0023] Preferably, determining the target dimension encoding corresponding to the data to be matched from the multiple initial dimension encodings includes:
[0024] Traverse the initial dimension encodings corresponding to the data to be matched;
[0025] Compare the encoding length of the current dimension encoding with that of the previous dimension encoding;
[0026] When the encoding length of the current dimension encoding is greater than that of the previous dimension encoding, update the parent node of the current dimension encoding to the previous dimension encoding;
[0027] When the encoding length of the current dimension encoding is equal to that of the previous dimension encoding, update the parent node of the current dimension encoding to the parent node of the previous dimension encoding;
[0028] When the encoding length of the encoding in the current dimension is less than that of the encoding in the previous dimension, traverse the chain of parent nodes of the encoding in the previous dimension, and update the parent encoding of the node whose encoding length is the same as that of the encoding in the current dimension in the chain to the parent node of the encoding in the current dimension; if there is no node with an encoding length the same as that of the encoding in the current dimension in the chain of parent nodes of the encoding in the previous dimension, then end;
[0029] Determine whether there is a left inclusion between the encoding in the current dimension and the parent node;
[0030] If there is a left inclusion, determine the encoding in the current dimension as the target dimension encoding corresponding to the data to be matched, and record the matching method of the current data to be matched.
[0031] Preferably, the method further includes:
[0032] If there is no left inclusion, clear the parent node of the encoding in the current dimension;
[0033] Determine whether the traversal of the dimension encoding of the current data to be matched is ended;
[0034] If the traversal of the dimension encoding of the current data to be matched is not ended, continue to execute the traversal of the dimension encoding of the current data to be matched;
[0035] If the traversal of the dimension encoding of the current data to be matched is ended, determine the target dimension encoding of the next data to be matched.
[0036] Preferably, using at least one matching method to obtain the initial dimension encoding of the data to be matched includes:
[0037] Determine the execution order of each matching method; the execution order is that the name matching takes precedence over the experience matching takes precedence over the similarity matching;
[0038] When the dimension encoding corresponding to the data to be matched cannot be successfully obtained after the execution of the current matching method, execute the next matching method.
[0039] Preferably, the obtaining of the data to be matched includes:
[0040] Load the dimension table; the dimension table includes the corresponding relationship between the dimension name and the dimension encoding;
[0041] Load the report; the report is the display carrier of the business data table; multiple data to be matched are stored in the business data table; at least the dimension name is included in the data to be matched;
[0042] Traverse the report, and query the data to be matched in the business data table according to the current report;
[0043] Traverse all the data to be matched corresponding to the current report.
[0044] Preferably, when loading the report and traversing the report, both are executed based on the table-level dimension.
[0045] Preferably, after traversing all the data to be matched corresponding to the current report, the method further includes:
[0046] Counting and reporting the matching result of the current report;
[0047] Updating the experience library;
[0048] Updating the dimension code corresponding to the dimension name of each data to be matched in the business data table.
[0049] Preferably, the method further includes:
[0050] When the dimension code corresponding to the data to be matched cannot be successfully obtained after the name matching is executed, determining whether to perform the experience matching according to the preset configuration parameters;
[0051] If the configuration parameter is not to perform the experience matching, skip the experience matching and directly execute the similarity matching.
[0052] According to the second aspect of the embodiments of the present application, a system for determining a dimension code is provided, including:
[0053] An acquisition module, configured to acquire data to be matched;
[0054] A matching module, configured to use at least one matching method to match and obtain an initial dimension code of the data to be matched;
[0055] A determination module, configured to, when there are multiple initial dimension codes, determine the target dimension code corresponding to the data to be matched from the multiple initial dimension codes according to the code length of the initial dimension code and the preset node level corresponding to the initial dimension code.
[0056] The technical solutions provided in this application may include the following beneficial effects: The method for determining dimension codes in this application includes: obtaining data to be matched; using at least one matching method to obtain the initial dimension codes of the data to be matched; in the case where there are multiple initial dimension codes, determining the target dimension code corresponding to the data to be matched from the multiple initial dimension codes according to the code length of the initial dimension codes and the preset node levels corresponding to the initial dimension codes. Due to the diversification of the matching methods in this application, the accuracy of obtaining the dimension codes of the data to be matched is higher than that of the prior art. Moreover, in this application, in the case where there are multiple initial dimension codes, the target dimension code corresponding to the data to be matched is determined from the multiple initial dimension codes, avoiding the situation of dimension names with the same name in a multi-level coding system. And the technical solutions in this application are more efficient than manual code checking and do not occupy human resources.
[0057] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this application. Brief Description of the Drawings
[0058] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0059] Figure 1 is a flowchart showing a method for determining dimension codes provided by an embodiment of this application;
[0060] Figure 2 is a flowchart showing a method for determining dimension codes provided by another embodiment of this application;
[0061] Figure 3 is a flowchart showing a method for determining dimension codes provided by yet another embodiment of this application;
[0062] Figure 4 is a flowchart showing a method for determining dimension codes provided by yet another embodiment of this application;
[0063] Figure 5 is a structural schematic block diagram of a system for determining dimension codes provided by an embodiment of this application.
[0064] Reference Signs: Obtaining Module - 21; Matching Module - 22; Determining Module - 23. Detailed Description of the Embodiments
[0065] Exemplary embodiments will be described in detail herein, and examples thereof are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0066] Embodiment 1
[0067] Figure 1 is a schematic flowchart of a method for determining a dimension code provided by an embodiment of the present application. Referring to Figure 1 a method for determining a dimension code includes:
[0068] S11: Obtain data to be matched;
[0069] S12: Use at least one matching method to obtain an initial dimension code of the data to be matched;
[0070] S13: When there are multiple initial dimension codes, determine the target dimension code corresponding to the data to be matched from the multiple initial dimension codes according to the coding length of the initial dimension code and the preset node level corresponding to the initial dimension code.
[0071] It should be noted that the technical solution in this embodiment is applied to scenarios where technologies such as ELT, data warehouse, data mining, and data visualization are applied in the budget networking supervision system.
[0072] ELT is used to describe the process of extracting data from the source end, transforming it, and loading it to the destination end.
[0073] A data warehouse is a strategic collection that provides all types of data support for the decision-making process at all levels of an enterprise. It is a single data storage created for analytical reporting and decision support purposes. It provides guidance for business process improvement, monitoring of time, cost, quality, and control for enterprises or departments that require business intelligence.
[0074] Data mining refers to the process of searching for information hidden in a large amount of data through algorithms.
[0075] Data mining is a technology that searches for patterns from a large amount of data by analyzing each piece of data. It mainly consists of three steps: data preparation, pattern finding, and pattern representation. Data preparation is to select the required data from relevant data sources and integrate it into a data set for data mining; pattern finding is to find the patterns contained in the data set by a certain method; pattern representation is to represent the found patterns in a way that can be understood by users as much as possible (such as visualization). The tasks of data mining include association analysis, clustering analysis, classification analysis, anomaly analysis, special group analysis, and evolution analysis, etc.
[0076] Data visualization is the scientific and technological research on the visual representation form of data. Among them, this visual representation form of data is defined as a kind of information extracted in a certain summary form, including various attributes and variables of the corresponding information units. The basic idea of data visualization technology is to represent each data item in the database as a single graphic element, and a large number of data sets constitute a data image. At the same time, the attribute values of the data are represented in the form of multi-dimensional data, and the data can be observed from different dimensions, so as to conduct a more in-depth observation and analysis of the data.
[0077] When technologies such as data mining and data visualization are applied in the budget networking supervision system, there is a necessary prerequisite: the dimension codes in the business data table must exist.
[0078] Based on this, a method for determining dimension codes is provided in this embodiment, including: obtaining the data to be matched; using at least one matching method to match and obtain the initial dimension code of the data to be matched; in the case where there are multiple initial dimension codes, according to the code length of the initial dimension code and the preset node level corresponding to the initial dimension code, determine the target dimension code corresponding to the data to be matched from multiple initial dimension codes. Due to the diversification of the matching methods in this embodiment, the accuracy of matching the dimension code of the data to be matched is higher than that of the prior art. And in this embodiment, in the case where there are multiple initial dimension codes, the target dimension code corresponding to the data to be matched is determined from multiple initial dimension codes, avoiding the situation of the same name dimensions in the multi-level coding system. And the technical solution in this embodiment has higher efficiency than manual coding verification and does not occupy human resources.
[0079] Embodiment Two
[0080] It should be noted that the matching methods at least include:
[0081] Name matching, experience matching, and similarity matching;
[0082] Name matching includes: obtaining the dimension name of the data to be matched;
[0083] Querying the dimension code corresponding to the dimension name of the data to be matched in the pre-constructed dimension table;
[0084] The experience matching includes: obtaining the dimension names and dimension identifiers of the data to be matched;
[0085] querying in a pre-constructed experience library for the dimension codes corresponding to the dimension names and dimension identifiers of the data to be matched;
[0086] The similarity matching includes: obtaining the dimension name of the data to be matched;
[0087] splitting and reconstructing the dimension name of the data to be matched to obtain a split and reconstructed name;
[0088] querying in a pre-constructed dimension table for the dimension codes corresponding to the dimension names whose similarity to the split and reconstructed name of the data to be matched is higher than a preset threshold, and when there are multiple dimension codes corresponding to the dimension names whose similarity is higher than the preset threshold, selecting the dimension code corresponding to the dimension name with the highest similarity.
[0089] Further, referring to Figure 2 , using at least one matching method to match and obtain the initial dimension code of the data to be matched, including:
[0090] determining the execution order of each matching method; the execution order is that name matching takes precedence over experience matching takes precedence over similarity matching;
[0091] when the dimension code corresponding to the data to be matched cannot be successfully matched after the execution of the current matching method, execute the next matching method.
[0092] Further, referring to Figure 2 , when the dimension code corresponding to the data to be matched cannot be successfully matched after the execution of name matching, determine whether to perform experience matching according to preset configuration parameters;
[0093] If the configuration parameter is not to perform experience matching, skip experience matching and directly execute similarity matching.
[0094] It should be noted that name matching, experience matching, and similarity matching are performed in sequence, and the matching ends as long as there is a result that meets the standard.
[0095] Among them, experience matching can be determined whether to perform experience matching through configuration parameters.
[0096] The criteria for each matching are as follows:
[0097] Name matching requires that the dimension name of the data to be matched is exactly the same as the dimension name in the dimension table, that is, using the dimension code corresponding to the dimension name in the dimension table that is exactly the same as the dimension name of the data to be matched as the dimension code obtained by matching the data to be matched;
[0098] The experience matching requires that the dimension name and dimension identifier of the data to be matched are equal to the historical matching records in the experience library;
[0099] The similarity matching returns the dimension code corresponding to the dimension name with the highest similarity to the data to be matched.
[0100] It should be noted that there are multiple dimensions in the system, and each dimension has a corresponding table. To manage these dimensions and the tables corresponding to them, a dimension identifier is assigned to each dimension to uniquely correspond to the dimension and the dimension table. After creating the business data table, the dimensions required by the business data table are fixed, and the dimension identifier corresponding to the data to be matched in the business data table has been determined.
[0101] It should be noted that the experience library is a storage structure designed to store non-standardized dimension information.
[0102] It should be noted that when performing similarity matching, the dimension name of the data to be matched is split and reconstructed. The purpose of obtaining the split and reconstructed name is to remove unreasonable words in the dimension name of the data to be matched. For example, if the dimension name of the data to be matched is "I. General Public Budget", it is split into "I", "I", "g", "e", "n", "e", "r", "a", "l", "p", "u", "b", "l", "i", "c", "b", "u", "d", "g", "e", "t", and the useless word "I." is removed and reconstructed as "g", "e", "n", "e", "r", "a", "l", "p", "u", "b", "l", "i", "c", "b", "u", "d", "g", "e", "t". Then, the similarity is calculated with the dimension names in the dimension table, and the dimension code corresponding to the dimension name with the highest similarity or a similarity higher than the preset threshold is returned as the dimension code of the current data to be matched.
[0103] Example Three
[0104] It should be noted that referring to Figure 3 , determining the target dimension code corresponding to the data to be matched from multiple initial dimension codes includes:
[0105] Traversing the initial dimension codes corresponding to the data to be matched;
[0106] Comparing the code length of the current dimension code with that of the previous dimension code;
[0107] When the code length of the current dimension code is greater than that of the previous dimension code, updating the parent node of the current dimension code to the previous dimension code;
[0108] When the code length of the current dimension code is equal to that of the previous dimension code, updating the parent node of the current dimension code to the parent node of the previous dimension code;
[0109] When the encoding length of the current dimension encoding is less than that of the previous dimension encoding, traverse the parent node chain of the previous dimension encoding, and update the parent encoding of the node whose encoding length is the same as that of the current dimension encoding in the chain to the parent node of the current dimension encoding; if there is no node with an encoding length the same as that of the current dimension encoding in the parent node chain of the previous dimension encoding, then end;
[0110] Determine whether there is a left inclusion between the current dimension encoding and the parent node;
[0111] If there is a left inclusion, then determine the current dimension encoding as the target dimension encoding corresponding to the data to be matched, and record the matching method of the current data to be matched.
[0112] If there is no left inclusion, then clear the parent node of the current dimension encoding;
[0113] Determine whether the traversal of the dimension encoding of the current data to be matched is over;
[0114] If the traversal of the dimension encoding of the current data to be matched is not over, then continue to execute the traversal of the dimension encoding of the current data to be matched;
[0115] If the traversal of the dimension encoding of the current data to be matched is over, then determine the target dimension encoding of the next data to be matched.
[0116] It should be noted that the initial dimension encoding is all the dimension encodings obtained by matching among the three matching methods. The dimension encoding is divided into category-item items, which results in the possibility that there may be sub-levels with the same dimension name under different major categories, which leads to the same dimension name, but due to different categories, their dimension encodings are different. This requires traversing the multiple initial dimension encodings that have been matched to find the corresponding upper-level encoding. (It should be noted that the data in the business data table are all tree structures similar to category-item items, with four levels of category-item items, and the encoding is similar to the following: XM (category), XM01 (item), XM0101 (sub-item), XM010101 (detail)).
[0117] It should be noted that this is only a minority case. In most cases, only one initial dimension encoding will be obtained by matching. If there is only one initial dimension encoding, then there is no need to traverse.
[0118] Examples of the three situations where the encoding length of the current dimension encoding is greater than that of the previous dimension encoding, the encoding length of the current dimension encoding is equal to that of the previous dimension encoding, and the encoding length of the current dimension encoding is less than that of the previous dimension encoding are as follows:
[0119] 1. For XM, XM01, XM0101, if the current dimension encoding is XM010102, then update the parent node of the current dimension encoding to XM0101;
[0120] 2. If the current dimension code of XM, XM01, XM0101 is XM0102, then update the parent node of the current dimension code to XM01;
[0121] 3. If the current dimension code of XM, XM01, XM0101 is XM02, then update the parent node of the current dimension code to XM.
[0122] It should be noted that if there is no node in the parent node chain of the previous dimension code whose coding length is the same as the current dimension code, that is, the traversal is empty, it means that the current dimension code is a class dimension code, and then end.
[0123] It should be noted that judging whether there is a left inclusion between the current dimension code and the parent node is to determine whether the current dimension code belongs to the large category corresponding to the dimension code of the parent node.
[0124] Example 4
[0125] It should be noted that referring to Figure 4 , obtain the data to be matched, including:
[0126] Load the dimension table; the dimension table includes the corresponding relationship between the dimension name and the dimension code;
[0127] Load the report; the report is the display carrier of the business data table; there are multiple pieces of data to be matched stored in the business data table; at least the dimension name is included in the data to be matched;
[0128] Traverse the report, and query the data to be matched in the business data table according to the current report;
[0129] Traverse all the data to be matched corresponding to the current report.
[0130] It should be noted that after traversing all the data to be matched corresponding to the current report, the method further includes:
[0131] Statistically report the matching result of the current report;
[0132] Update the experience library;
[0133] Update the dimension code corresponding to the dimension name of each data to be matched in the business data table.
[0134] It should be noted that the dimensions are divided into table-level dimensions and in-table dimensions. The data displayed in the report comes from the business data table, and for the same report in the business data table, data for multiple years, months, quarters, departments, and regions are saved. When viewing, it is necessary to specify the specific year, month, quarter, department, and region according to business needs to filter the data to be displayed. These dimensions are table-level dimensions.
[0135] The table-level dimension is used to filter the data displayed in the report. The in-table dimension is used to perform name matching, experience matching, and similarity matching to find the corresponding dimension codes. Through these dimension codes, horizontal or vertical data analysis can be carried out.
[0136] In specific practices, the table-level dimensions include: year, region, month, department or unit. When loading and traversing the report, it is executed based on the table-level dimension. The table-level dimension is used to filter out the data displayed in the report that is meaningful in business, and then matching is performed according to the data in the report to meet the business requirements.
[0137] It should be noted that the report is a carrier for displaying the data in the business data table, usually in the form of a table or a statistical chart. It displays the data in the business data table to the user in the form of a table or a statistical chart. The underlying layer of the report is the business data table. Generally, one report corresponds to one business data table, while one business data table corresponds to multiple reports. Suppose there are 10 fields in the business data table, including the dimension code field and the dimension name field. In the original data, there will definitely be corresponding data for only the dimension name, while the dimension code is empty. It is necessary to find the dimension code manually or through a program according to the dimension name. The dimension table stores the dimension information of a certain report, and there is a corresponding relationship between the dimension name and the dimension code in it. Different matching methods can be used in the dimension table to find the item code corresponding to the data to be matched during the matching process.
[0138] It should be noted that the business data table is the business data imported into the database through the data import function, such as accounting, debt, and other data.
[0139] It should be noted that the dimension table is a data table, which is the standard for the data to be matched in the report. The dimension codes and dimension names of all the data to be matched are stored in this table. The dimension table stores items similar to accounting expenditure subjects, accounting income subjects, standard department unit codes, etc. The subject and the department unit code are uniformly called dimension codes (subject code, department code, unit code) and dimension names (subject name, department name, unit name) in this embodiment for the convenience of processing. In specific practices, the dimension table stores data in key-value pairs. The dimension table has three layers. The key of the first layer is the year, and the corresponding value is all the dimension table information; the key of the second layer is the dimension table name, and the corresponding value is the specific dimension information; the key of the third layer is the dimension name, and the corresponding value is a list of dimension codes. The data in this list is stored in the form of key-value pairs, where the key is the dimension name and the value is the dimension code (such an organization is to quickly obtain all the dimension codes that meet the conditions and save calculation time when there are the same dimension names but different dimension codes).
[0140] It should be noted that the experience database is a data table that records historical matching results. Therefore, after traversing all the data to be matched corresponding to the current report, the experience database needs to be updated.
[0141] It should be noted that since the report is only a display carrier of the business data table, after obtaining the target dimension code corresponding to the data to be matched in the report, the dimension code corresponding to the dimension name of each data to be matched in the business data table needs to be updated.
[0142] It should be noted that the data to be matched in the business data table is queried according to the current report, including querying specific report data information (sorted by serial number) according to the superior business data table, year, administrative region, month, department or unit. Group queries need to be processed for monthly reports, annual reports, and department reports. The data is saved in a doubly linked list. In addition to saving the business data table information, each node in the doubly linked list also needs to add a matching method field (0 name matching, 1 experience matching, 2 similarity matching, in string type) and a parent node pointer.
[0143] Embodiment Five
[0144] Figure 5 is a structural schematic block diagram of a system for determining dimension codes provided by an embodiment of the present application. Refer to Figure 5 , a system for determining dimension codes, includes:
[0145] An acquisition module 21, configured to acquire data to be matched;
[0146] A matching module 22, configured to use at least one matching method to match and obtain an initial dimension code of the data to be matched;
[0147] A determination module 23, configured to, when there are multiple initial dimension codes, determine the target dimension code corresponding to the data to be matched from the multiple initial dimension codes according to the code length of the initial dimension codes and the preset node levels corresponding to the initial dimension codes.
[0148] It can be understood that the system for determining dimension encoding in this embodiment includes: an acquisition module 21, a matching module 22, and a determination module 23. The acquisition module 21 is used to acquire data to be matched; the matching module 22 is used to obtain the initial dimension encoding of the data to be matched by using at least one matching method; the determination module 23 is used to determine the target dimension encoding corresponding to the data to be matched from multiple initial dimension encodings according to the encoding length of the initial dimension encoding and the preset node level corresponding to the initial dimension encoding when there are multiple initial dimension encodings. Due to the diversification of the matching methods in this embodiment, the accuracy of obtaining the dimension encoding of the data to be matched is higher than that of the prior art. Moreover, in this embodiment, when there are multiple initial dimension encodings, the target dimension encoding corresponding to the data to be matched is determined from multiple initial dimension encodings, avoiding the situation of the same-name dimensions in the multi-level encoding system. And the technical solution in this embodiment has higher efficiency than manual encoding verification and does not occupy human resources.
[0149] It can be understood that the same or similar parts in the above embodiments can be referred to each other, and the content not detailed in some embodiments can be seen in the same or similar content in other embodiments.
[0150] It should be noted that in the description of this application, the terms "first", "second", etc. are only used for descriptive purposes and cannot be understood as indicating or implying relative importance. In addition, in the description of this application, unless otherwise specified, the meaning of "a plurality of" refers to at least two.
[0151] Any process or method description in the flowchart or described in other ways herein can be understood as representing a module, segment, or part of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of this application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in the reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the technical field to which the embodiments of this application belong.
[0152] It should be understood that each part of this application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following well-known technologies in the art can be used: discrete logic circuits with logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits with appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0153] Those of ordinary skill in the art can understand that all or part of the steps carried out in the method of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0154] In addition, in each of the embodiments of the present application, the functional units can be integrated in a processing module, or each unit can exist physically alone, or two or more units can be integrated in a module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the above integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.
[0155] The above-mentioned storage medium can be a read-only memory, a magnetic disk, an optical disk, etc.
[0156] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0157] Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. A method for determining a dimensional code, characterized in that, Including: Obtain the data to be matched; Using at least one matching method, obtain the initial dimension encoding of the data to be matched; When there are multiple initial dimension encodings, determine the target dimension encoding corresponding to the data to be matched from the multiple initial dimension encodings according to the encoding length of the initial dimension encoding and the preset node level corresponding to the initial dimension encoding; Determining the target dimension encoding corresponding to the data to be matched from multiple initial dimension encodings includes: Traverse the initial dimension encoding corresponding to the data to be matched; Compare the encoding length of the current dimension encoding with that of the previous dimension encoding; When the encoding length of the current dimension encoding is greater than that of the previous dimension encoding, update the parent node of the current dimension encoding to the previous dimension encoding; When the encoding length of the current dimension encoding is equal to that of the previous dimension encoding, update the parent node of the current dimension encoding to the parent node of the previous dimension encoding; When the encoding length of the current dimension encoding is less than that of the previous dimension encoding, traverse the parent node chain of the previous dimension encoding, and update the parent encoding of the first node with the same encoding length as the current dimension encoding to the parent node of the current dimension encoding; if there is no node with the same encoding length as the current dimension encoding in the parent node chain of the previous dimension encoding, then end; Determine whether there is a left inclusion between the current dimension encoding and the parent node; If there is a left inclusion, determine the current dimension encoding as the target dimension encoding corresponding to the data to be matched, and record the matching method of the current data to be matched; If there is no left inclusion, clear the parent node of the current dimension encoding; Determine whether the traversal of the dimension encoding of the current data to be matched is over; If the traversal of the dimension encoding of the current data to be matched is not over, continue to execute the traversal of the dimension encoding of the current data to be matched; If the traversal of the dimension encoding of the current data to be matched is over, determine the target dimension encoding of the next data to be matched.
2. The method according to claim 1, characterized in that, The matching method at least includes: Name matching, experience matching, and similarity matching; The name matching includes: obtaining the dimension name of the data to be matched; Query the dimension encoding corresponding to the dimension name of the data to be matched in the pre-constructed dimension table; The experience matching includes: obtaining the dimension name and dimension identifier of the data to be matched; Query the dimension encoding corresponding to the dimension name and dimension identifier of the data to be matched in the pre-constructed experience library; The similarity matching includes: obtaining the dimension name of the data to be matched; Split and reconstruct the dimension name of the data to be matched to obtain the split and reconstructed name; Query the dimension encoding corresponding to the dimension name with a similarity higher than the preset threshold to the split and reconstructed name of the data to be matched in the pre-constructed dimension table, and when there are multiple dimension encodings corresponding to the dimension names with a similarity higher than the preset threshold, select the dimension encoding corresponding to the dimension name with the highest similarity.
3. The method according to claim 2, wherein Using at least one matching method to obtain the initial dimension encoding of the data to be matched includes: Determine the execution order of each matching method; the execution order is that the name matching takes precedence over the experience matching takes precedence over the similarity matching; When the dimension code corresponding to the data to be matched cannot be successfully matched after the current matching method is executed, the next matching method is executed.
4. The method according to claim 1, wherein The obtaining of the data to be matched includes: Loading a dimension table; the dimension table includes the correspondence between dimension names and dimension codes; Loading a report; the report is a display carrier of a business data table; multiple data to be matched are stored in the business data table; at least a dimension name is included in the data to be matched; Traversing the report and querying the data to be matched in the business data table according to the current report; Traversing all the data to be matched corresponding to the current report.
5. The method according to claim 4, characterized in that, When loading the report and traversing the report, both are executed based on table-level dimensions.
6. The method according to claim 4, wherein After all the data to be matched corresponding to the current report are traversed, the method further includes: Counting and reporting the matching result of the current report; Updating the experience library; Updating the dimension code corresponding to the dimension name of each data to be matched in the business data table.
7. The method according to claim 3, characterized in that, The method further includes: When the dimension code corresponding to the data to be matched cannot be successfully matched after the name matching is executed, determining whether to perform the experience matching according to preset configuration parameters; If the configuration parameter is not to perform the experience matching, skipping the experience matching and directly executing the similarity matching.
8. A system for determining dimensional coding, characterized in that, Including: An obtaining module, configured to obtain data to be matched; A matching module, configured to use at least one matching method to match and obtain an initial dimension code of the data to be matched; A determining module, configured to, when there are multiple initial dimension codes, determine a target dimension code corresponding to the data to be matched from the multiple initial dimension codes according to the code length of the initial dimension code and the preset node level corresponding to the initial dimension code; Determining a target dimension code corresponding to the data to be matched from multiple initial dimension codes includes: Traversing the initial dimension codes corresponding to the data to be matched; Comparing the code length of the current dimension code with that of the previous dimension code; When the code length of the current dimension code is greater than that of the previous dimension code, updating the parent node of the current dimension code to the previous dimension code; When the code length of the current dimension code is equal to that of the previous dimension code, updating the parent node of the current dimension code to the parent node of the previous dimension code; When the code length of the current dimension code is less than that of the previous dimension code, traversing the parent node chain of the previous dimension code and updating the parent code of the first node with the same code length as the current dimension code in it to the parent node of the current dimension code; if there is no node with the same code length as the current dimension code in the parent node chain of the previous dimension code, end; Judging whether there is a left inclusion between the current dimension code and the parent node; If there is a left inclusion, determining the current dimension code as the target dimension code corresponding to the data to be matched and recording the matching method of the current data to be matched; If there is no left inclusion, clearing the parent node of the current dimension code; Judging whether the traversal of the dimension code of the current data to be matched is ended; If the traversal of the dimension code of the current data to be matched is not ended, continuing to execute the traversal of the dimension code of the current data to be matched; If the traversal of the dimension encoding of the currently to-be-matched data ends, determine the target dimension encoding of the next to-be-matched data.
Citation Information
Patent Citations
GIS (Geographic Information System) vector data hierarchical coding method and device based on tree-shaped hierarchical index
CN115145930A
Dependent dimensions
US20220122188A1