Metadata auditing method and apparatus, electronic device, and readable storage medium
By generating similarity parameters to automatically identify the differences between the data dictionary and the audit standards, the problem of low audit efficiency in existing technologies is solved, and efficient and accurate data dictionary auditing is achieved.
Patent Information
- Application Number
- CN202310091756.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-01-28
- Publication Date
- 2026-05-15
- Estimated Expiration
- 2043-01-28
AI Technical Summary
Existing technologies suffer from low auditing efficiency when auditing the differences between data dictionaries and current data standards in non-standard information systems, mainly because the manual identification process is greatly affected by human factors.
By acquiring multiple metadata groups and using the attribute information of the metadata to generate similarity parameters, the system can automatically identify and generate audit reports, identify parts of the data dictionary that differ significantly from the audit standards, and replace manual audits.
It improved the efficiency of data dictionary auditing, reduced human interference, and enhanced the accuracy and efficiency of auditing.
Smart Images

Figure CN116257513B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data management technology, specifically to a metadata auditing method, apparatus, electronic device, and readable storage medium. Background Technology
[0002] For existing information systems, configuring a data dictionary can standardize the management of various data within the system.
[0003] To manage different information systems serving the same business scenario, relevant technologies will unify the data formats of different information systems by establishing data standards. However, for information systems that were established before the data standards were established, the lack of data standards leads to significant differences between the data dictionaries in these information systems and the current data standards.
[0004] Currently, to ensure the stability of data flow, related technologies employ manual auditing to identify differences between the data dictionary in non-standard information systems and the existing data standards. Based on the identification results, data in non-standard information systems that do not match the data standards are corrected accordingly. However, the aforementioned manual identification process is greatly affected by human factors, resulting in low auditing efficiency. Summary of the Invention
[0005] The purpose of this disclosure is to provide a metadata auditing method, apparatus, electronic device, and readable storage medium to solve the technical problem of low auditing efficiency when auditing differences between data dictionaries and existing data standards in non-standard information systems.
[0006] In a first aspect, embodiments of this disclosure provide a metadata auditing method, including:
[0007] Multiple metadata groups are obtained, each metadata group including first metadata and second metadata, wherein the first metadata is the metadata included in the data dictionary to be audited, and the second metadata is the metadata included in the audit criteria;
[0008] Based on the attribute information of the first metadata and the attribute information of the second metadata in each metadata group, a plurality of first similarity parameters are generated that correspond one-to-one with the plurality of metadata groups, wherein the first similarity parameters are used to characterize the similarity between the first metadata and the second metadata.
[0009] Based on the plurality of first similarity parameters, a first audit report is generated. The first audit report includes a target similarity parameter and a metadata group corresponding to the target similarity parameter. The target similarity parameter is the first similarity parameter among the plurality of first similarity parameters that is less than a similarity threshold.
[0010] In one embodiment, obtaining multiple metadata groups includes:
[0011] Obtain multiple copies of the first metadata;
[0012] Semantic recognition is performed on multiple first metadata to obtain multiple metadata indication information. The multiple metadata indication information corresponds one-to-one with the multiple first metadata. The metadata indication information is used to characterize the semantics of the base station data indicated by the corresponding first metadata.
[0013] Based on the multiple metadata indication information, the multiple first metadata and the multiple second metadata included in the audit standard are matched to obtain the multiple metadata groups. The metadata indication information corresponding to the first metadata in any metadata group matches at least a portion of the corpus associated with the second metadata. Among the multiple second metadata, the corpus associated with different second metadata is different.
[0014] In one embodiment, obtaining the plurality of the first metadata includes:
[0015] Obtain multiple raw metadata entries included in the data dictionary to be audited, as well as the data call frequency of each raw metadata entry;
[0016] Based on the data call frequency of each of the original metadata, the first metadata is determined from the plurality of original metadata, wherein the data call frequency of the first metadata is greater than a preset frequency threshold.
[0017] In one embodiment, after generating the first audit report based on the plurality of first similarity parameters, the method further includes:
[0018] Obtain multiple corrected metadata that correspond one-to-one with the multiple first metadata, wherein the corrected metadata is the metadata obtained after correcting the corresponding first metadata based on the first audit report;
[0019] Based on the attribute information of each of the modified metadata and the attribute information of the corresponding second metadata, a plurality of second similarity parameters are generated that correspond one-to-one with the plurality of modified metadata.
[0020] A second audit report is generated based on the multiple second similarity parameters.
[0021] In one embodiment, generating multiple second similarity parameters corresponding one-to-one with the multiple corrected metadata based on the attribute information of each of the corrected metadata and the attribute information of the corresponding second metadata includes:
[0022] Based on the metadata priority of the modified metadata, a target metadata is determined from the plurality of modified metadata, wherein the metadata priority is used to characterize the business priority of the modified metadata in the data dictionary, and the target metadata is the modified metadata whose corresponding metadata priority is higher than the priority threshold;
[0023] Based on the attribute information of each target metadata and the attribute information of the corresponding second metadata, the plurality of second similarity parameters are generated.
[0024] In one embodiment, before determining the target metadata based on the metadata priority of the modified metadata among the plurality of modified metadata, the method further includes:
[0025] Lineage analysis is performed on the multiple modified metadata to obtain lineage analysis information;
[0026] Based on the bloodline analysis information, the metadata priority of each of the plurality of modified metadata is determined.
[0027] In one embodiment, at least one of the first metadata and the second metadata is reference metadata, and the attribute information of the reference metadata includes at least one of the data type, data length, and data precision of the target metadata.
[0028] Secondly, embodiments of this disclosure also provide a metadata auditing device, the device comprising:
[0029] The acquisition module is used to acquire multiple metadata groups, each metadata group including first metadata and second metadata, wherein the first metadata is the metadata included in the data dictionary to be audited, and the second metadata is the metadata included in the audit standards;
[0030] The first generation module is used to generate multiple first similarity parameters corresponding one-to-one with the multiple metadata groups based on the attribute information of the first metadata and the attribute information of the second metadata in each metadata group, wherein the first similarity parameters are used to characterize the similarity between the first metadata and the second metadata.
[0031] The second generation module is used to generate a first audit report based on the plurality of first similarity parameters. The first audit report includes a target similarity parameter and a metadata group corresponding to the target similarity parameter. The target similarity parameter is the first similarity parameter among the plurality of first similarity parameters that is less than a similarity threshold.
[0032] Thirdly, embodiments of this disclosure also provide an electronic device, including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the above-described metadata auditing method.
[0033] Fourthly, embodiments of this disclosure also provide a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the above-described metadata auditing method.
[0034] In this embodiment of the disclosure, metadata groups are used to characterize the data association between the data dictionary to be audited and the audit standard. Based on the attribute information of the first metadata and the attribute information of the second metadata in each metadata group, the similarity between the first metadata and the second metadata in the metadata group is determined. Then, by setting the similarity threshold, the metadata group corresponding to the target similarity parameter can be quickly identified, that is, the data part in the data dictionary that has a large difference from the audit standard can be identified, so as to replace the manual audit of the data dictionary and improve the audit efficiency of the data dictionary. Attached Figure Description
[0035] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments of this disclosure will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] Figure 1 This is a flowchart of a metadata auditing method provided in an embodiment of this disclosure;
[0037] Figure 2 This is a schematic diagram of the structure of a metadata auditing device provided in an embodiment of this disclosure;
[0038] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation
[0039] The technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this disclosure. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0040] This disclosure provides a metadata auditing method; see [link to relevant documentation]. Figure 1 , Figure 1 This is a flowchart of the metadata auditing method provided in the embodiments of this disclosure, such as... Figure 1 As shown, it includes the following steps:
[0041] Step 101: Obtain multiple metadata groups.
[0042] Each of the metadata groups includes first metadata and second metadata, wherein the first metadata is metadata included in the data dictionary to be audited, and the second metadata is metadata included in the audit criteria.
[0043] For example, the data dictionary to be audited can be the data dictionary of a base station information system established before the audit standard was published.
[0044] Step 102: Based on the attribute information of the first metadata and the attribute information of the second metadata in each metadata group, generate multiple first similarity parameters that correspond one-to-one with the multiple metadata groups.
[0045] The first similarity parameter is used to characterize the similarity between the first metadata and the second metadata.
[0046] For example, the attribute information of the first metadata includes at least one of referential attributes and data attributes. The referential attributes of the first metadata are used to describe the data that the first metadata refers to in the data dictionary. For example, when the string "base station operation data" is used as the referential attribute of a certain first metadata, it means that the first metadata is metadata used to describe various working parameters in the operation of the base station system. The data attributes of the first metadata are used to describe information such as the data type, data length, and data precision of the first metadata.
[0047] Similarly, the attribute information of the second metadata also includes referential attributes and data attributes. For their specific meanings, please refer to the above description of the attribute information of the first metadata, which will not be repeated here.
[0048] It should be noted that, in this disclosure, the first similarity parameter of a metadata group can be generated by comparing the similarity between the data attributes of the first metadata in the metadata group and the data attributes of the second metadata in the metadata group to obtain the first similarity parameter of the metadata group.
[0049] Step 103: Generate a first audit report based on the multiple first similarity parameters.
[0050] The first audit report includes a target similarity parameter and a metadata group corresponding to the target similarity parameter. The target similarity parameter is the first similarity parameter among the plurality of first similarity parameters that is less than a similarity threshold.
[0051] The output of the first audit report can intuitively display the metadata group corresponding to the target similarity parameters to the user, that is, to show the metadata part of the data dictionary that needs to be corrected. This can save the user the workload of identifying the metadata part of the data dictionary that needs to be corrected, and can greatly improve the audit efficiency of the data dictionary.
[0052] In the application, the first audit report may also include multiple first similarity parameters other than the target similarity parameter, as well as metadata groups corresponding to the other similarity parameters. In this case, the first audit report includes multiple first similarity parameters and metadata groups corresponding to each first similarity parameter. Furthermore, based on the existence of a similarity threshold, each metadata group will be labeled. The labels set include a first label for indicating the corresponding target similarity parameter and a second label for indicating the corresponding other similarity parameters, in order to facilitate subsequent data backtracking.
[0053] Specifically, after the first audit report is output, the first metadata of the metadata group corresponding to the target similarity parameter can be corrected manually, or the correction operation of the first metadata of the metadata group corresponding to the target similarity parameter can be automatically executed according to a preset correction function. This disclosure embodiment does not limit the scope of the correction.
[0054] In this embodiment of the disclosure, metadata groups are used to characterize the data association between the data dictionary to be audited and the audit standard. Based on the attribute information of the first metadata and the attribute information of the second metadata in each metadata group, the similarity between the first metadata and the second metadata in the metadata group is determined. Then, by setting the similarity threshold, the metadata group corresponding to the target similarity parameter can be quickly identified, that is, the data part in the data dictionary that has a large difference from the audit standard can be identified, so as to replace the manual audit of the data dictionary and improve the audit efficiency of the data dictionary.
[0055] It should be noted that the auditing of the data dictionary can be understood as the auditing and correction of the data rules of the data dictionary. For example, the data dictionary can define the data rule as follows: the location data indicating the geographical coordinates of the base station includes first data indicating the longitude of the base station coordinates and second data indicating the latitude of the base station coordinates, and both the first data and the second data are real numbers and retain two decimal places. The data example generated based on this data rule can be (longitude 165.24°, latitude 44.17°). In this disclosure, the correction of the first metadata based on the first audit report can be understood as the correction of the data rules in the data dictionary, such as limiting the aforementioned first data and second data to be real numbers and retaining five decimal places.
[0056] The correction operation for the data corresponding to the first metadata to be corrected in the data dictionary can be completed based on the corrected first metadata or based on the second metadata corresponding to the first metadata.
[0057] In one embodiment, at least one of the first metadata and the second metadata is reference metadata, and the attribute information of the reference metadata includes at least one of the data type, data length, and data precision of the target metadata.
[0058] In this embodiment, the attribute information of the reference metadata can be understood as the data attributes of the reference metadata.
[0059] For example, the similarity comparison of the data types of the first metadata and the second metadata can be as follows: if the data types of the first metadata and the second metadata are the same, output a first similarity value; if the data types of the first metadata and the second metadata are different, output a second similarity value, which is less than the first similarity value.
[0060] For example, the similarity comparison of the data length of the first metadata and the data length of the second metadata can be as follows: if the data length of the first metadata is less than the data length of the second metadata, a third similarity value is output; and if the data length of the first metadata is greater than or equal to the data length of the second metadata, a fourth similarity value is output, and the fourth similarity value is greater than the third similarity value.
[0061] For example, the similarity comparison of the data precision of the first metadata and the data precision of the second metadata can be as follows: if the data precision of the first metadata is lower than the data precision of the second metadata, output a fifth similarity value; and if the data precision of the first metadata is higher than or equal to the data precision of the second metadata, output a sixth similarity value, and the sixth similarity value is greater than the fifth similarity value.
[0062] When the attribute information of the reference metadata includes the data type, data length, and data precision of the target metadata, the similarity values of the first metadata and the second metadata in terms of data type, data length, and data precision can be statistically or weighted to obtain the corresponding first similarity parameter.
[0063] In one embodiment, obtaining multiple metadata groups includes:
[0064] Obtain multiple copies of the first metadata;
[0065] Semantic recognition is performed on multiple first metadata to obtain multiple metadata indication information. The multiple metadata indication information corresponds one-to-one with the multiple first metadata. The metadata indication information is used to characterize the semantics of the base station data indicated by the corresponding first metadata.
[0066] Based on the multiple metadata indication information, the multiple first metadata and the multiple second metadata included in the audit standard are matched to obtain the multiple metadata groups. The metadata indication information corresponding to the first metadata in any metadata group matches at least a portion of the corpus associated with the second metadata. Among the multiple second metadata, the corpus associated with different second metadata is different.
[0067] In this embodiment, by performing semantic recognition on each first metadata, the metadata indication information of each first metadata is determined, that is, the referential attribute of each first metadata is determined, and the matching of the first metadata and the second metadata is completed accordingly, thereby forming the aforementioned metadata group. By using the data semantics indicated by the first metadata and the corpus corresponding to the second metadata, the association between the first metadata and the second metadata is automatically identified, which can improve the generation efficiency and accuracy of the metadata group, and make the final output first audit report more reliable and timely.
[0068] The matching of the metadata indication information corresponding to the first metadata and at least a portion of the corpus associated with the second metadata can be understood as: the metadata indication information corresponding to the first metadata is the same as at least one corpus in the corpus associated with the second metadata.
[0069] The corpus associated with the second metadata can be obtained from the corpus associated with the second metadata and the preset corpus association model. By constructing the corpus, the situation of incorrect matching or missing matching of the first metadata and the second metadata can be avoided, which can make the metadata indication information more accurate and reliable.
[0070] In one embodiment, obtaining the plurality of the first metadata includes:
[0071] Obtain multiple raw metadata entries included in the data dictionary to be audited, as well as the data call frequency of each raw metadata entry;
[0072] Based on the data call frequency of each of the original metadata, the first metadata is determined from the plurality of original metadata, wherein the data call frequency of the first metadata is greater than a preset frequency threshold.
[0073] In this embodiment, the data call frequency of each original metadata in the data dictionary is used to distinguish between the primary metadata of the corresponding base station service and the non-core metadata of the corresponding data flow / record (e.g., log data, backup data, etc.). This can reduce the amount of primary metadata that needs to be processed in the audit process and make the output of the primary audit report more efficient. The data call frequency of the non-core metadata is less than or equal to the above-mentioned frequency threshold.
[0074] It should be noted that, in this disclosure, in addition to determining the first metadata from multiple original metadata based on the frequency of data calls, the first metadata can also be determined from multiple original metadata based on preset metadata tags. Other methods can also be selected to complete the operation of determining the first metadata from multiple original metadata according to actual needs. This disclosure does not limit the specific method of determining the first metadata from multiple original metadata.
[0075] The metadata tags include a first tag for indicating core business metadata and a second tag for indicating non-core business metadata. The setting of the first and second tags can be done manually or based on a preset tagging model. The tagging model can be a trained neural network model for distinguishing between core business metadata and non-core business metadata.
[0076] In one embodiment, after generating the first audit report based on the plurality of first similarity parameters, the method further includes:
[0077] Obtain multiple corrected metadata that correspond one-to-one with the multiple first metadata, wherein the corrected metadata is the metadata obtained after correcting the corresponding first metadata based on the first audit report;
[0078] Based on the attribute information of each of the modified metadata and the attribute information of the corresponding second metadata, a plurality of second similarity parameters are generated that correspond one-to-one with the plurality of modified metadata.
[0079] A second audit report is generated based on the multiple second similarity parameters.
[0080] In this embodiment, after the first audit report is generated, the user can correct the first metadata that does not match the audit standard indicated by the first audit report by manual correction or automatic correction using a preset correction function. After the correction is completed, the corrected metadata is obtained, which is the corrected data dictionary. By comparing the similarity between the corrected metadata and the second metadata, iterative audit processing of the corrected data dictionary can be achieved. This can further improve the matching degree between the data dictionary and the audit standard when applying the audit method of this disclosure.
[0081] In one embodiment, generating multiple second similarity parameters corresponding one-to-one with the multiple corrected metadata based on the attribute information of each of the corrected metadata and the attribute information of the corresponding second metadata includes:
[0082] Based on the metadata priority of the modified metadata, a target metadata is determined from the plurality of modified metadata, wherein the metadata priority is used to characterize the business priority of the modified metadata in the data dictionary, and the target metadata is the modified metadata whose corresponding metadata priority is higher than the priority threshold;
[0083] Based on the attribute information of each target metadata and the attribute information of the corresponding second metadata, the plurality of second similarity parameters are generated.
[0084] In this embodiment, compared to the method of performing full audit processing on the modified metadata, the target metadata among multiple modified metadata is determined by metadata priority, that is, the modified metadata corresponding to the business data in the modified metadata is determined, so as to reduce the number of modified metadata to be audited and make the output of the second audit report more efficient.
[0085] In one example, in addition to comparing the similarity between the target metadata and its corresponding second metadata, similarity can also be compared between the second metadata corresponding to other metadata whose service priority is higher than the priority threshold. This is to adapt to the complex relationships between various data in the base station information system. The relationships between various data in the base station information system are manifested in the following way: a change in one data rule will cause changes in other data rules. Therefore, other metadata that passed the audit in the previous audit process may fail the audit after the target metadata is corrected and becomes corrected metadata. In this case, if only the corrected metadata is audited, the output second audit report will not accurately reflect the degree of difference between the corrected data dictionary and the audit standard. Therefore, by auditing other metadata whose service priority is higher than the priority threshold, the second audit report can be made more accurate.
[0086] Among them, other metadata can be understood as the first metadata corresponding to other similarity parameters.
[0087] In one embodiment, before determining the target metadata based on the metadata priority of the modified metadata among the plurality of modified metadata, the method further includes:
[0088] Lineage analysis is performed on the multiple modified metadata to obtain lineage analysis information;
[0089] Based on the bloodline analysis information, the metadata priority of each of the plurality of modified metadata is determined.
[0090] In this embodiment, by using lineage analysis to determine the metadata priority of each modified metadata, the output second audit report can be made more accurate while ensuring audit efficiency.
[0091] For example, lineage analysis information may include an indication of the number of other modified metadata associated with each modified metadata, in which case the number of other modified metadata associated with each modified metadata can be used as the data priority of that modified metadata.
[0092] like Figure 2 As shown in the embodiments of this disclosure, a metadata auditing device 200 is also provided, the metadata auditing device 200 comprising:
[0093] The acquisition module 201 is used to acquire multiple metadata groups, each of which includes first metadata and second metadata, wherein the first metadata is metadata included in the data dictionary to be audited, and the second metadata is metadata included in the audit standards.
[0094] The first generation module 202 is used to generate a plurality of first similarity parameters corresponding one-to-one with the plurality of metadata groups based on the attribute information of the first metadata and the attribute information of the second metadata in each metadata group, wherein the first similarity parameters are used to characterize the similarity between the first metadata and the second metadata.
[0095] The second generation module 203 is used to generate a first audit report based on the plurality of first similarity parameters. The first audit report includes a target similarity parameter and a metadata group corresponding to the target similarity parameter. The target similarity parameter is a first similarity parameter among the plurality of first similarity parameters that is less than a similarity threshold.
[0096] In one embodiment, the acquisition module 201 includes:
[0097] The first acquisition unit is used to acquire multiple of the first metadata;
[0098] A semantic recognition unit is used to perform semantic recognition on multiple first metadata to obtain multiple metadata indication information. The multiple metadata indication information corresponds one-to-one with the multiple first metadata. The metadata indication information is used to characterize the semantics of the base station data indicated by the corresponding first metadata.
[0099] A matching unit is configured to match multiple first metadata and multiple second metadata included in the audit criteria based on the multiple metadata indication information to obtain multiple metadata groups. In any one of the metadata groups, the metadata indication information corresponding to the first metadata matches at least a portion of the corpus associated with the second metadata. Among the multiple second metadata, different second metadata are associated with different corpora.
[0100] In one embodiment, the first acquisition unit is specifically used for:
[0101] Obtain multiple raw metadata entries included in the data dictionary to be audited, as well as the data call frequency of each raw metadata entry;
[0102] Based on the data call frequency of each of the original metadata, the first metadata is determined from the plurality of original metadata, wherein the data call frequency of the first metadata is greater than a preset frequency threshold.
[0103] In one embodiment, the device 200 further includes:
[0104] The correction acquisition module is used to acquire multiple correction metadata that correspond one-to-one with multiple first metadata, wherein the correction metadata is metadata obtained after correcting the corresponding first metadata based on the first audit report;
[0105] The third generation module is used to generate multiple second similarity parameters that correspond one-to-one with the multiple modified metadata based on the attribute information of each of the modified metadata and the attribute information of the corresponding second metadata.
[0106] The fourth generation module is used to generate a second audit report based on the plurality of second similarity parameters.
[0107] In one embodiment, the third generation module includes:
[0108] The target determination unit is used to determine the target metadata from the plurality of modified metadata based on the metadata priority of the modified metadata, wherein the metadata priority is used to characterize the business priority of the modified metadata in the data dictionary, and the target metadata is the modified metadata whose corresponding metadata priority is higher than the priority threshold;
[0109] The generation unit is used to generate the plurality of second similarity parameters based on the attribute information of each target metadata and the attribute information of the corresponding second metadata.
[0110] In one embodiment, the device 200 further includes:
[0111] The lineage analysis module is used to perform lineage analysis on the multiple modified metadata to obtain lineage analysis information;
[0112] The priority determination module is used to determine the metadata priority of each of the plurality of modified metadata based on the bloodline analysis information.
[0113] In one embodiment, at least one of the first metadata and the second metadata is reference metadata, and the attribute information of the reference metadata includes at least one of the data type, data length, and data precision of the target metadata.
[0114] The metadata auditing device 200 provided in this embodiment can implement the various processes in the above method embodiments, and will not be described again here to avoid repetition.
[0115] Please see Figure 3 , Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this disclosure, such as... Figure 3 As shown, the electronic device includes: a processor 301, a memory 302, and a program 3021 stored in the memory 302 and executable on the processor 301.
[0116] When program 3021 is executed by processor 301, it can achieve the following: Figure 1 Any steps in the corresponding method embodiments and the achievement of the same beneficial effects will not be repeated here.
[0117] Those skilled in the art will understand that all or part of the steps of the methods described in the above embodiments can be implemented by hardware related to program instructions, and the program can be stored in a readable medium.
[0118] This disclosure also provides a readable storage medium storing a computer program, which, when executed by a processor, can perform the above-described functions. Figure 1 Any step in the corresponding method embodiment can achieve the same technical effect, and will not be repeated here to avoid repetition.
[0119] The computer-readable storage medium of this disclosure can be any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. For example, a computer-readable storage medium can be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.
[0120] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.
[0121] The program code contained on the storage medium can be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.
[0122] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, and C++, and conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or terminal. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0123] The above description represents the preferred embodiments of this disclosure. It should be noted that those skilled in the art can make various improvements and modifications without departing from the principles described herein, and these improvements and modifications should also be considered within the scope of protection of this disclosure.
Claims
1. A metadata auditing method, characterized in that, The method includes: Multiple metadata groups are obtained, each metadata group including first metadata and second metadata, wherein the first metadata is the metadata included in the data dictionary to be audited, and the second metadata is the metadata included in the audit criteria; Based on the attribute information of the first metadata and the attribute information of the second metadata in each metadata group, a plurality of first similarity parameters are generated that correspond one-to-one with the plurality of metadata groups, wherein the first similarity parameters are used to characterize the similarity between the first metadata and the second metadata. Based on the plurality of first similarity parameters, a first audit report is generated. The first audit report includes a target similarity parameter and a metadata group corresponding to the target similarity parameter. The target similarity parameter is the first similarity parameter among the plurality of first similarity parameters that is less than a similarity threshold. The acquisition of multiple metadata groups includes: Obtain multiple copies of the first metadata; Semantic recognition is performed on multiple first metadata to obtain multiple metadata indication information. The multiple metadata indication information corresponds one-to-one with the multiple first metadata. The metadata indication information is used to characterize the semantics of the base station data indicated by the corresponding first metadata. Based on the multiple metadata indication information, the multiple first metadata and the multiple second metadata included in the audit standard are matched to obtain the multiple metadata groups. The metadata indication information corresponding to the first metadata in any metadata group matches at least a portion of the corpus associated with the second metadata. Among the multiple second metadata, the corpus associated with different second metadata is different.
2. The method according to claim 1, characterized in that, The acquisition of multiple first metadata includes: Obtain multiple raw metadata entries included in the data dictionary to be audited, as well as the data call frequency of each raw metadata entry; Based on the data call frequency of each of the original metadata, the first metadata is determined from the plurality of original metadata, wherein the data call frequency of the first metadata is greater than a preset frequency threshold.
3. The method according to claim 1, characterized in that, After generating the first audit report based on the plurality of first similarity parameters, the method further includes: Obtain multiple corrected metadata that correspond one-to-one with the multiple first metadata, wherein the corrected metadata is the metadata obtained after correcting the corresponding first metadata based on the first audit report; Based on the attribute information of each of the modified metadata and the attribute information of the corresponding second metadata, a plurality of second similarity parameters are generated that correspond one-to-one with the plurality of modified metadata. A second audit report is generated based on the multiple second similarity parameters.
4. The method according to claim 3, characterized in that, Based on the attribute information of each of the corrected metadata and the attribute information of the corresponding second metadata, a plurality of second similarity parameters corresponding one-to-one with the plurality of corrected metadata are generated, including: Based on the metadata priority of the modified metadata, a target metadata is determined from the plurality of modified metadata, wherein the metadata priority is used to characterize the business priority of the modified metadata in the data dictionary, and the target metadata is the modified metadata whose corresponding metadata priority is higher than the priority threshold; Based on the attribute information of each target metadata and the attribute information of the corresponding second metadata, the plurality of second similarity parameters are generated.
5. The method according to claim 4, characterized in that, Before determining the target metadata from the plurality of corrected metadata, the method further includes: Lineage analysis is performed on the multiple modified metadata to obtain lineage analysis information; Based on the bloodline analysis information, the metadata priority of each of the plurality of modified metadata is determined.
6. The method according to claim 1, characterized in that, At least one of the first metadata and the second metadata is reference metadata, and the attribute information of the reference metadata includes at least one of the following: data type, data length, and data precision.
7. A metadata auditing device, characterized in that, The device includes: The acquisition module is used to acquire multiple metadata groups, each metadata group including first metadata and second metadata, wherein the first metadata is the metadata included in the data dictionary to be audited, and the second metadata is the metadata included in the audit standards; The first generation module is used to generate multiple first similarity parameters corresponding one-to-one with the multiple metadata groups based on the attribute information of the first metadata and the attribute information of the second metadata in each metadata group, wherein the first similarity parameters are used to characterize the similarity between the first metadata and the second metadata. The second generation module is used to generate a first audit report based on the plurality of first similarity parameters. The first audit report includes a target similarity parameter and a metadata group corresponding to the target similarity parameter. The target similarity parameter is a first similarity parameter among the plurality of first similarity parameters that is less than a similarity threshold. The acquisition module is specifically used for: Obtain multiple copies of the first metadata; Semantic recognition is performed on multiple first metadata to obtain multiple metadata indication information. The multiple metadata indication information corresponds one-to-one with the multiple first metadata. The metadata indication information is used to characterize the semantics of the base station data indicated by the corresponding first metadata. Based on the multiple metadata indication information, the multiple first metadata and the multiple second metadata included in the audit standard are matched to obtain the multiple metadata groups. The metadata indication information corresponding to the first metadata in any metadata group matches at least a portion of the corpus associated with the second metadata. Among the multiple second metadata, the corpus associated with different second metadata is different.
8. An electronic device, characterized in that, It includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the computer program, when executed by the processor, implements the steps of the metadata auditing method as described in any one of claims 1 to 6.
9. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the metadata auditing method as described in any one of claims 1 to 6.