Data standard recommendation method and device

By hierarchically dividing and vectorizing the standard information, a standard vector library is constructed, which solves the problem of time-consuming and labor-consuming manual judgment in traditional data governance methods, and realizes efficient and accurate data standard recommendations.

CN120336503APending Publication Date: 2025-07-18SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510414685.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

Traditional data governance methods rely on manual judgment and manual operations, which are time-consuming and labor-intensive and prone to errors. With the increase in data volume and technology development, it is difficult to select and apply appropriate data standards.

Method used

By hierarchically dividing the standard information in the preset standard pool, a standard vector library is built, and the data to be managed from the standard vector library is matched based on the hierarchy and priority, and the corresponding standard information is recommended.

Benefits of technology

It improves the accuracy and efficiency of data standard recommendations, reduces labor costs, and improves the speed and accuracy of data processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336503A_ABST
    Figure CN120336503A_ABST
Patent Text Reader

Abstract

The invention provides a data standard recommendation method and device. The accuracy of data standard recommendation can be improved. The data standard recommendation method comprises the steps of obtaining standard information and to-be-treated data in a preset standard pool; wherein the standard information comprises a standard English name field, a standard Chinese name field, a standard description field and a data type field; dividing the standard information in the preset standard pool according to a preset hierarchy to obtain a hierarchical structure and a hierarchical priority; vectorizing the standard information divided into hierarchical structures and hierarchical priorities, and constructing a standard vector library; based on a hierarchical structure and a hierarchical priority, performing hierarchical matching on the to-be-treated data from the standard vector library to obtain a recommendation result; wherein the recommendation result represents standard information correspondingly associated with the recommended to-be-governed data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and particularly to a data standard recommendation method and apparatus. Background Art

[0002] With the in-depth development of digital transformation, enterprises and organizations are facing increasing data management and governance requirements. Data governance is a key process to ensure data quality, availability, and security. It not only helps to unify the definition and usage of data but also promotes data sharing and integration between different systems. In this process, an important challenge is how to effectively identify and apply appropriate data standards. Traditional data governance methods often rely on manual judgment and manual operations, which are not only time-consuming and laborious but also error-prone. With the growth of data volume and the development of technology, new data types are constantly emerging, which further increases the difficulty of selecting and applying appropriate data standards. Summary of the Invention

[0003] To solve the above technical problems, the present invention is proposed. Embodiments of the present invention provide a data standard recommendation method and apparatus, which can improve the accuracy of data standard recommendation.

[0004] According to one aspect of the present invention, there is provided a data standard recommendation method, including: obtaining standard information and data to be governed in a preset standard pool; wherein the standard information includes a standard English name field, a standard Chinese name field, a standard description field, and a data type field; dividing the standard information in the preset standard pool according to a preset hierarchy to obtain a hierarchical structure and a hierarchical priority; vectorizing the standard information with the divided hierarchical structure and hierarchical priority to construct a standard vector library; based on the hierarchical structure and hierarchical priority, hierarchically matching the data to be governed from the standard vector library to obtain a recommendation result; wherein the recommendation result represents the standard information associated with the recommended data to be governed.

[0005] In one embodiment, dividing the standard information in the preset standard pool according to a preset hierarchy to obtain a hierarchical structure and a hierarchical priority includes: based on the data type field in the standard information and the preset hierarchy, dividing the standard information into pure character type, pure numeric type, character-numeric mixed type, and time type as the first level and setting it as the first priority; wherein the first priority means that when filtering the data to be governed, the data to be governed is preferentially filtered according to the data type, and the standard information associated with the recommended data to be governed is recommended among the standard information of the same data type.

[0006] In one embodiment, the standard information in the preset standard pool is divided according to a preset hierarchy to obtain a hierarchy structure and a hierarchy priority. It further includes: determining the consistent words and background words of the standard Chinese name field of each standard information based on the standard type after the first-level division; wherein, the consistent words represent the common attributes of the standard information, and the background words represent the characteristic background description of the consistent words; setting the consistent words as the second priority based on the preset hierarchy; setting the background words as the third priority based on the preset hierarchy.

[0007] In one embodiment, determining the consistent words and background words of the standard Chinese name field of each standard information based on the standard type after the first-level division includes: respectively extracting the standard Chinese name field and the data type field of the standard information based on the standard type after the first-level division; performing word segmentation processing on the standard Chinese name field to obtain the word segmentation results of multiple standard information; for each word segmentation result, obtaining the standard information set containing the word segmentation result; counting the data type fields in which the word segmentation result appears in the standard information set, and the number of times each data type field appears; calculating the proportion of the data type field with the most occurrences in the standard information set; when the proportion is greater than a preset threshold, using the data type field with the most occurrences as the consistent word of the word segmentation result.

[0008] In one embodiment, the data standard recommendation method further includes: comparing the word segmentation result with a preset word segmentation threshold to obtain a comparison result; when the comparison result is less than the preset word segmentation threshold, defining the word segmentation result as a background word; wherein, the background word represents the characteristic background description of the consistent word.

[0009] In one embodiment, vectorizing the standard information with the divided hierarchy structure and hierarchy priority to construct a standard vector library includes: respectively vectorizing the standard description field, the consistent word, and the background word based on the standard description field, the data type field, the consistent word, and the background word; obtaining three standard vector libraries according to the vectorized standard description field, the consistent word, and the background word; wherein, the first standard vector library corresponds to the standard description field, the second standard vector library corresponds to the consistent word, and the third standard vector library corresponds to the background word.

[0010] In one embodiment, before hierarchically matching the data to be governed from the standard vector library based on the hierarchy structure and hierarchy priority to obtain a recommendation result, the data standard recommendation method includes: performing word segmentation processing and vectorization processing on the standard Chinese name field of the data to be governed; converting the data type of the data to be governed into a pure character type, a pure numeric type, a character-numeric mixed type, or a time type, denoted as the field type of the data to be governed.

[0011] In one embodiment, based on the hierarchy and hierarchical priorities, the data to be governed is hierarchically matched with the standard vector library to obtain a recommendation result, including: based on the field type of the data to be governed, hierarchically matching with the standard vector library according to the first priority to obtain a set of standard information of the same type.

[0012] In one embodiment, based on the hierarchy and hierarchical priorities, hierarchically matching the data to be governed with the standard vector library to obtain a recommendation result further includes: in the set of standard information of the same type, according to the consistency words of the second priority, performing vector retrieval on the word segmentation result of the data field to be governed in the second standard vector library corresponding to the consistency words to obtain a set of hit standards; for the set of hit standards, according to the background words of the third priority, performing vector retrieval in the third standard vector library corresponding to the background words to obtain a recommended standard list; sorting the recommended standard list according to the similarity to obtain the recommendation result.

[0013] According to another aspect of the present invention, there is provided a data standard recommendation device, including: an acquisition module for acquiring standard information and data to be governed in a preset standard pool; wherein, the standard information includes a standard English name field, a standard Chinese name field, a standard description field, and a data type field; a division module for dividing the standard information in the preset standard pool according to a preset level to obtain a hierarchy and hierarchical priorities; a construction module for vectorizing the standard information with the divided hierarchy and hierarchical priorities to construct a standard vector library; a recommendation module for hierarchically matching the data to be governed with the standard vector library based on the hierarchy and hierarchical priorities to obtain a recommendation result; wherein, the recommendation result represents the standard information associated with the recommended data to be governed.

[0014] The data standard recommendation method and device provided by the present invention can quickly locate a specific range by hierarchically dividing the standard information in the preset standard pool, reduce unnecessary full-scale scans, and make the data structure clearer and more orderly. And the hierarchical index can significantly accelerate the query speed and reduce the amount of data to be processed for data recommendation. Therefore, a reasonable hierarchical division of the standard information in the standard pool can greatly improve the efficiency and accuracy of data processing. Vectorizing the standard information can enhance the relevance between standard information based on semantics, greatly improving the retrieval quality and retrieval efficiency. Finally, the data to be governed is associated with the corresponding standard information based on the constructed hierarchical standard vector library, which can reduce the labor cost during data governance and improve the data governance efficiency. Description of the Drawings

[0015] The above and other purposes, features and advantages of the present invention will become more apparent by describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0016] Figure 1 It is a flowchart of a data standard recommendation method provided by an exemplary embodiment of the present invention.

[0017] Figure 2 It is a structural diagram of a data standard recommendation device provided by an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0018] Below, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described here.

[0019] Traditional data governance methods often rely on human judgment and manual operations, which is not only time-consuming and labor-intensive, but also prone to errors. With the growth of data volume and the development of technology, new data types continue to emerge. It takes a lot of manpower to accurately associate data fields with governance standards, which is also prone to errors. To solve the above problems, Figure 1 is a flow chart of a data standard recommendation method provided by an exemplary embodiment of the present invention. Figure 1 For example, first obtain the standard information and data to be managed in the preset standard pool (see Figure 1 S100); wherein the standard information includes a standard English name field, a standard Chinese name field, a standard description field, and a data type field. Secondly, the standard information in the preset standard pool is divided according to the preset level to obtain a hierarchical structure and hierarchical priority (see Figure 1 Then, the standard information of the hierarchical structure and hierarchical priority is vectorized to construct a standard vector library (see Figure 1 Finally, based on the hierarchical structure and hierarchical priority, the data to be managed is hierarchically matched from the standard vector library to obtain the recommendation results (see Figure 1 The recommendation result indicates the standard information corresponding to the recommended data to be managed.

[0020] Combined with the following Figure 1 , a more detailed introduction to the data standard recommendation method provided in the embodiment of the present application is given.

[0021] In S100, obtain the standard information in the preset standard pool and the data to be governed.

[0022] Among them, the standard information includes a standard English name field, a standard Chinese name field, a standard description field, and a data type field. The standard information of data refers to a set of rules, conventions, and guidelines used to ensure the consistency, accuracy, and interchangeability of data. Each standard defines a governance method for the corresponding field. For example, for the ID card standard, it stipulates that the ID card must have 18 digits and only contain numbers and letters. The standard pool refers to all the stored standards, and each standard contains at least information such as a standard English name, a standard Chinese name, a standard description, and a data format.

[0023] In some embodiments, an example of the preset standard pool can be as shown in Table 1:

[0024] Table 1

[0025]

[0026] The data to be governed refers to the data fields that need to be associated with specific standards during the data governance process, and it mainly contains the following information: field name, field Chinese description, and data format. When new data sources are accessed (such as adding new user behavior log fields), business rules change (such as the original "customer level" field needs to adapt to new compliance standards), or the technical architecture is adjusted (such as fields migrating from local to the cloud and new security rules need to be defined), data to be governed may be generated.

[0027] In some embodiments, an example of the data to be governed can be as shown in Table 2:

[0028] Table 2

[0029]

[0030] In S200, divide the standard information in the preset standard pool according to a preset hierarchy to obtain a hierarchical structure and a hierarchical priority.

[0031] In some embodiments, first, all the standard information in the standard pool can be extracted as: a standard English name field, a standard Chinese name field, a standard description field, and a data type field. Based on the data type field in the standard information and the preset hierarchy, the standard information is divided into pure character type, pure numeric type, character-numeric mixed type, and time type, which is used as the first level and set as the first priority; among them, the first priority means that when filtering the data to be governed, the data to be governed is filtered according to the data type first, and in the standard information of the same data type, the standard information corresponding to the associated data to be governed is recommended. In the subsequent standard recommendation process, the standards are filtered according to the data type first, and the recommendation is only carried out among the standards of the same data type, effectively improving the accuracy and efficiency.

[0032] In some embodiments, to continue hierarchical division, based on the standard types after the first-level division, the consistent words and background words of the standard Chinese name fields of each standard information can be determined; wherein, the consistent words represent the common attributes of the standard information, and the background words represent the characteristic background descriptions of the consistent words; based on a preset hierarchy, the consistent words are set as the second priority; based on a preset hierarchy, the background words are set as the third priority.

[0033] As a possible implementation, the consistent words refer to the common attributes of the standard. For example, the name of the legal representative and the name of the beneficiary are both names in specific scenarios. The background words are the specific background descriptions of the consistent words. For example, the legal representative and the beneficiary can distinguish names and can also distinguish ages. That is to say, the scope of the consistent words is larger than that of the background words, and the background words are used to further distinguish the background words.

[0034] Therefore, in order to determine the consistent words, in some embodiments, based on the standard types after the first-level division, the standard Chinese name fields and data type fields of the standard information can be extracted respectively; the standard Chinese name fields are subjected to word segmentation processing to obtain the word segmentation results of multiple standard information; for each word segmentation result, the standard information set containing the word segmentation result is obtained; the data type fields in which the word segmentation results appear in the standard information set and the number of times each data type field appears are counted; the proportion of the data type field with the most occurrences in the standard information set is calculated; when the proportion is greater than a preset threshold, the data type field with the most occurrences is used as the consistent word of the word segmentation result.

[0035] Word segmentation (Tokenization) is a key step in natural language processing (NLP) and text data preprocessing. Word segmentation refers to the process of splitting continuous text (such as sentences, paragraphs) into meaningful words or tokens. Since the data to be governed may be relatively complex and has certain differences from the standard information, to improve the accuracy of matching, the standard Chinese name fields of the standard information can be subjected to word segmentation processing. If the text is not segmented, the search engine may not be able to accurately match the user's query. Word segmentation processing can improve the accuracy of text retrieval and enhance semantic understanding. Summarize the word segmentation results of all standard information, and respectively count the data type consistency of each word, where the data type consistency means: the proportion of the data type with the most occurrences in all the standards containing the segmented word in this part of the standards. The higher the proportion, the more general the segmented word is and can be used to classify data of a specific data type.

[0036] Next, after dividing the consistent words at the second level, the background words can be further divided. The word segmentation results are compared with the preset word segmentation threshold to obtain a comparison result. When the comparison result is less than the preset word segmentation threshold, the word segmentation result is defined as a background word. Among them, the background word represents the characteristic background description of the consistent word. That is to say, by traversing the word segmentation results of each standard, it is determined whether each word belongs to a consistent word or a background word through the threshold. Those greater than the threshold indicate that the field containing the word only points to the same type of data and has strong classification ability, and are determined as consistent words, while those less than the threshold are defined as background words.

[0037] In some embodiments, two new variables are added to each field so far: consistent words and background words, which respectively represent the background and common capabilities of the standard information. Table 3 shows the structure of a kind of standard information.

[0038] Table 3

[0039]

[0040] In some embodiments, Table 4 shows the hierarchical structure and hierarchical priority of a preset standard pool.

[0041] Table 4

[0042]

[0043] Continue to refer to Figure 1 , in S300, the standard information with the divided hierarchical structure and hierarchical priority is vectorized to construct a standard vector library.

[0044] In some embodiments, based on the standard description field, data type field, consistent words and background words, the standard description field, consistent words and background words are respectively vectorized. According to the vectorized standard description field, consistent words and background words, three standard vector libraries are obtained. Among them, the first standard vector library corresponds to the standard description field, the second standard vector library corresponds to the consistent words, and the third standard vector library corresponds to the background words.

[0045] As a possible implementation manner, an Embedding model can be used to vectorize the standard description field, consistent words and background words respectively. The Embedding model maps text, images or other data into a low-dimensional dense vector space, enabling it to be better understood and calculated by machines. Embedding is a technology that converts discrete data (such as words, sentences, images) into continuous vectors, making entities with similar semantics closer in the vector space. Therefore, through the Embedding model, unstructured data (text, images) can be converted into computable vectors, improving the relevance between semantics.

[0046] In order to facilitate subsequent matching of the correlation degree between the data to be governed and the standard information, in some embodiments, the data to be governed can also be subjected to word segmentation processing and vectorization. For example, the standard Chinese name field of the data to be governed is subjected to word segmentation processing and vectorization processing, and the data type of the data to be governed is converted into a pure character type, a pure number type, a character-number mixed type, or a time type, which is recorded as the field type of the data to be governed.

[0047] Next, in S400, based on the hierarchical structure and hierarchical priority, the data to be governed is hierarchically matched from the standard vector library to obtain a recommendation result.

[0048] The recommendation result represents the standard information associated with the recommended data to be governed.

[0049] After the data to be governed is subjected to word segmentation processing and vectorization, for the fields of the data to be governed, first, according to the first priority (data type), hierarchical matching is performed from the standard vector library, and the standards of the same data type are retrieved from the standard pool to obtain a set of standard information of the same type. Next, in the set of standard information of the same type, according to the consistency words of the second priority, vector retrieval is performed on the word segmentation result of the data to be governed fields in the second standard vector library corresponding to the consistency words to obtain the set of standards after hitting; finally, for the set of standards after hitting, according to the background words of the third priority, vector retrieval is performed in the third standard vector library corresponding to the background words to obtain a recommended standard list; the recommended standard list is sorted according to the similarity to obtain a recommendation result. That is to say, the matching result of each priority is filtered based on the result of the previous priority, and finally the recommendation result is obtained. Among all the remaining standard information, the standard information with the highest recommendation similarity is recommended to complete the matching of the standard information of the data to be governed and classify the data to be governed according to the standard.

[0050] Figure 2 It is a schematic structural diagram of a data standard recommendation device provided by an exemplary embodiment of the present invention, as Figure 2 shown, the data standard recommendation device 2 includes: an acquisition module 21 for acquiring the standard information and the data to be governed in a preset standard pool; wherein, the standard information includes a standard English name field, a standard Chinese name field, a standard description field, and a data type field; a division module 22 for dividing the standard information in the preset standard pool according to a preset level to obtain a hierarchical structure and a hierarchical priority; a construction module 23 for vectorizing the standard information with the divided hierarchical structure and hierarchical priority to construct a standard vector library; a recommendation module 24 for hierarchically matching the data to be governed from the standard vector library based on the hierarchical structure and hierarchical priority to obtain a recommendation result; wherein, the recommendation result represents the standard information associated with the recommended data to be governed.

[0051] The data standard recommendation device provided by the present invention can quickly locate a specific range by hierarchically dividing the standard information in the preset standard pool, reducing unnecessary full-scale scans and making the data structure clearer and more orderly. Moreover, the hierarchical index can significantly accelerate the query speed and reduce the amount of data to be processed for data recommendation. Therefore, reasonable hierarchical division of the standard information in the standard pool can greatly improve the efficiency and accuracy of data processing. Vectorizing the standard information can enhance the relevance between standard information based on semantics, greatly improving the retrieval quality and efficiency. Finally, the data to be governed can be associated with the corresponding standard information based on the constructed hierarchical standard vector library, reducing the labor cost during data governance and improving the data governance efficiency.

[0052] In one embodiment, the division module 22 can be configured to: divide the standard information into pure character type, pure numeric type, character-numeric mixed type, and time type based on the data type field and preset hierarchy in the standard information, as the first hierarchy and set as the first priority; wherein, the first priority indicates that when filtering the data to be governed, the data to be governed is preferentially filtered according to the data type, and the standard information corresponding to the data to be governed is recommended among the standard information of the same data type.

[0053] In one embodiment, the division module 22 can also be configured to: determine the consistent words and background words of the standard Chinese name field of each standard information based on the standard types after dividing the first hierarchy; wherein, the consistent words represent the common attributes of the standard information, and the background words represent the characteristic background descriptions of the consistent words; set the consistent words as the second priority based on the preset hierarchy; set the background words as the third priority based on the preset hierarchy.

[0054] In one embodiment, the division module 22 can also be configured to: respectively extract the standard Chinese name field and data type field of the standard information based on the standard types after dividing the first hierarchy; perform word segmentation on the standard Chinese name field to obtain the word segmentation results of multiple standard information; for each word segmentation result, obtain the standard information set containing the word segmentation result; count the data type fields that appear in the standard information set for the word segmentation result, and the number of times each data type field appears; calculate the proportion of the data type field with the most occurrences in the standard information set; when the proportion is greater than the preset threshold, use the data type field with the most occurrences as the consistent word of the word segmentation result.

[0055] In one embodiment, the data standard recommendation device 2 can be configured to: compare the word segmentation result with the preset word segmentation threshold to obtain a comparison result; when the comparison result is less than the preset word segmentation threshold, define the word segmentation result as a background word; wherein, the background word represents the characteristic background description of the consistent word.

[0056] In one embodiment, the construction module 23 may be configured to: vectorize the standard description field, the consistency words, and the background words respectively based on the standard description field, the data type field, the consistency words, and the background words; obtain three standard vector libraries according to the vectorized standard description field, consistency words, and background words; wherein, the first standard vector library corresponds to the standard description field, the second standard vector library corresponds to the consistency words, and the third standard vector library corresponds to the background words.

[0057] In one embodiment, the data standard recommendation device 2 may be configured to: perform word segmentation processing and vectorization processing on the standard Chinese name field of the data to be governed; convert the data type of the data to be governed into a pure character type, a pure numeric type, a character-numeric mixed type, or a time type, which is denoted as the field type of the data to be governed.

[0058] In one embodiment, the recommendation module 24 may be configured to: based on the field type of the data to be governed, hierarchically match from the standard vector library according to the first priority to obtain a set of standard information of the same type.

[0059] In one embodiment, the recommendation module 24 may also be configured to: in the set of standard information of the same type, perform vector retrieval on the word segmentation result of the data field to be governed in the second standard vector library corresponding to the consistency words according to the consistency words of the second priority to obtain a set of standards after hitting; for the set of standards after hitting, perform vector retrieval in the third standard vector library corresponding to the background words according to the background words of the third priority to obtain a recommended standard list; sort the recommended standard list according to the similarity to obtain a recommendation result.

[0060] An embodiment of the present invention provides a data standard recommendation device. The device embodiment can be implemented by software, or by hardware, or by a combination of software and hardware. From the hardware level, in addition to the CPU, memory, network interface, and non-volatile memory, the device where the device is located in the embodiment usually may also include other hardware, such as a forwarding chip responsible for processing packets, and so on. Taking software implementation as an example, as a logically meaningful device, it is formed by the CPU of the device where it is located reading the corresponding computer program instructions in the non-volatile memory into the memory and running.

[0061] According to another aspect of the present invention, a computer-readable storage medium is provided, and the storage medium stores a computer program, and the computer program is used to execute the data standard recommendation method of any one of the above embodiments.

[0062] In addition to the above methods and devices, embodiments of the present invention may also be a computer program product, which includes computer program instructions, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the data standard recommendation method according to various embodiments of the present invention described above.

[0063] According to another aspect of the present invention, there is provided an electronic device, which includes: a processor; a memory for storing processor-executable instructions; and the processor for executing the data standard recommendation method according to any of the above embodiments.

[0064] In addition, an embodiment of the present invention may also be a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the data standard recommendation method according to various embodiments of the present invention described above.

[0065] The above are only the preferred embodiments of the present invention, and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for recommending data standards, characterized in that, Including: Obtain the standard information and the data to be governed in the preset standard pool; wherein, the standard information includes a standard English name field, a standard Chinese name field, a standard description field, and a data type field; Divide the standard information in the preset standard pool according to a preset hierarchy to obtain a hierarchical structure and a hierarchical priority; Vectorize the standard information with the divided hierarchical structure and hierarchical priority to construct a standard vector library; Based on the hierarchical structure and hierarchical priority, hierarchically match the data to be governed from the standard vector library to obtain a recommendation result; wherein, the recommendation result represents the standard information associated with the recommended data to be governed.

2. The data standard recommendation method according to claim 1, wherein Dividing the standard information in the preset standard pool according to a preset hierarchy to obtain a hierarchical structure and a hierarchical priority includes: Based on the data type field and the preset hierarchy in the standard information, divide the standard information into pure character type, pure numeric type, character-numeric mixed type, and time type, as the first level and set it as the first priority; wherein, the first priority means that when filtering the data to be governed, the data to be governed is filtered according to the data type first, and the standard information associated with the data to be governed is recommended among the standard information of the same data type.

3. The data standard recommendation method according to claim 2, wherein Dividing the standard information in the preset standard pool according to a preset hierarchy to obtain a hierarchical structure and a hierarchical priority further includes: Based on the standard type after dividing the first level, determine the consistent words and background words of the standard Chinese name field of each standard information; wherein, the consistent words represent the common attributes of the standard information, and the background words represent the characteristic background description of the consistent words; Based on the preset hierarchy, set the consistent words as the second priority; Based on the preset hierarchy, set the background words as the third priority.

4. The data standard recommendation method according to claim 3, wherein Determining the consistent words and background words of the standard Chinese name field of each standard information based on the standard type after dividing the first level includes: Based on the standard type after dividing the first level, respectively extract the standard Chinese name field and the data type field of the standard information; Perform word segmentation on the standard Chinese name field to obtain the word segmentation results of multiple standard information; For each word segmentation result, obtain the standard information set containing the word segmentation result; Count the data type fields in which the word segmentation result appears in the standard information set, and the number of times each data type field appears; Calculate the proportion of the data type field with the most occurrences in the standard information set; When the proportion is greater than the preset threshold, use the data type field with the most occurrences as the consistent word of the word segmentation result.

5. The data standard recommendation method according to claim 4, wherein The data standard recommendation method further includes: Compare the word segmentation result with a preset word segmentation threshold to obtain a comparison result; When the comparison result is less than the preset word segmentation threshold, define the word segmentation result as a background word; wherein, the background word represents the characteristic background description of the consistent word.

6. The data standard recommendation method according to claim 3, wherein Vectorizing the standard information with the divided hierarchical structure and hierarchical priority to construct a standard vector library includes: Based on the standard description field, data type field, consistent words, and background words, vectorize the standard description field, consistent words, and background words respectively; According to the vectorized standard description fields, consistency words, and background words, three standard vector libraries are obtained; among them, the first standard vector library corresponds to the standard description fields, the second standard vector library corresponds to the consistency words, and the third standard vector library corresponds to the background words.

7. The data standard recommendation method according to claim 6, wherein Before obtaining the recommendation result by hierarchically matching the data to be governed from the standard vector library based on the hierarchical structure and hierarchical priority, the data standard recommendation method includes: Perform word segmentation processing and vectorization processing on the standard Chinese name field of the data to be governed. Convert the data type of the data to be governed into pure character type, pure numeric type, character-numeric mixed type, or time type, which is denoted as the field type of the data to be governed.

8. The data standard recommendation method according to claim 7, wherein Based on the hierarchical structure and hierarchical priority, hierarchically match the data to be governed from the standard vector library to obtain the recommendation result, including: Based on the field type of the data to be governed, hierarchically match from the standard vector library according to the first priority to obtain a set of standard information of the same type.

9. The data standard recommendation method according to claim 8, wherein Based on the hierarchical structure and hierarchical priority, hierarchically match the data to be governed from the standard vector library to obtain the recommendation result, which also includes: In the set of standard information of the same type, perform vector retrieval on the word segmentation result of the data field to be governed in the second standard vector library corresponding to the consistency words according to the second priority of the consistency words to obtain the set of standards after hitting. For the set of standards after hitting, perform vector retrieval in the third standard vector library corresponding to the background words according to the third priority of the background words to obtain the recommended standard list. Sort the recommended standard list according to the similarity to obtain the recommendation result.

10. A data standard recommendation device, characterized in that, Including: An acquisition module for acquiring the standard information and the data to be governed in the preset standard pool; among them, the standard information includes the standard English name field, the standard Chinese name field, the standard description field, and the data type field. A division module for dividing the standard information in the preset standard pool according to the preset level to obtain the hierarchical structure and hierarchical priority. A construction module for vectorizing the standard information with the hierarchical structure and hierarchical priority divided to construct a standard vector library. A recommendation module for hierarchically matching the data to be governed from the standard vector library based on the hierarchical structure and hierarchical priority to obtain the recommendation result; among them, the recommendation result represents the standard information associated with the recommended data to be governed.

Citation Information

Patent Citations

  • Data matching method and device and electronic equipment

    CN114153962A

  • Electronic medical record fine item extraction method and system based on semantic information

    CN114239582A