A data element construction method, a terminal device, and a computer storage medium

By acquiring data items to be standardized, analyzing the tagging term types using a preset data element model, and adding data terms, standardized data elements are generated, solving the problem of insufficient existing data element resources and realizing rapid and standardized data element construction.

CN115858827BActive Publication Date: 2026-03-27ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-17
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing data element resources are insufficient to meet business needs, and temporarily constructed data elements are unlikely to meet specification requirements.

Method used

By acquiring data items to be standardized, analyzing the type of tagged words using a preset data element model, obtaining the type of untagged words, and adding data words according to the type of untagged words, standardized data elements are generated.

Benefits of technology

It enables efficient, standardized, and automated data element construction, quickly replenishing data element resources to meet business needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115858827B_ABST
    Figure CN115858827B_ABST
Patent Text Reader

Abstract

The application provides a data element construction method, a terminal device and a computer storage medium. A to-be-standardized data item is acquired. A to-be-standardized word type in the to-be-standardized data item is analyzed by using a preset data element model. An un-to-be-standardized word type is acquired according to the to-be-standardized word type. A data word is added according to the un-to-be-standardized word type. A standardized data element is generated by using the to-be-standardized data item and the added data word. In this way, the newly-added data element can be efficiently, normatively and automatically constructed. When the existing data element cannot meet the requirement, the function helps the user to quickly and normatively add the data element.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a data element construction method, a terminal device, and a computer storage medium. Background Technology

[0002] With the popularization and development of internet technology, massive amounts of data are constantly emerging from our lives. The development of big data and artificial intelligence technologies, based on distributed data storage and computing, has provided the foundation and application scenarios for the use of massive amounts of data. To enable users to more easily extract value from massive amounts of data, technologies such as data organization and processing, and data asset management have also received widespread attention. Among these, automated data governance technologies, such as automatic data element benchmarking, have played a significant role in improving the quality and reducing the cost of data governance.

[0003] The concept of data elements plays a very important role in the fields of data governance and data standardization. However, there are the following pain points in using data elements: 1. Existing data element resources are insufficient to meet business needs; 2. Temporarily constructed data elements are difficult to meet the requirements of specifications. Summary of the Invention

[0004] To address the aforementioned technical problems, this application proposes a data element construction method, a terminal device, and a computer storage medium.

[0005] To address the aforementioned technical problems, this application proposes a data element construction method, comprising:

[0006] Obtain the data items to be standardized;

[0007] The type of tagging term in the data item to be standardized is analyzed using a preset data element model;

[0008] Obtain the untargeted word type based on the targeted word type;

[0009] Data terms were not added according to the specified terminology type;

[0010] Standardized data elements are generated using the data items to be standardized and the added data terms.

[0011] The step of generating standardized data elements using the data items to be standardized and the added data terms includes:

[0012] According to the type of the target word of the data item to be standardized, obtain the data words to be standardized output by the preset data element model;

[0013] The standardized data element is generated by combining the data words to be standardized with the added data words.

[0014] The obtaining of the to-be-standardized data word output by the preset data element model comprises:

[0015] The obtaining of the several candidate data words of the to-be-standardized data word of the target word type and the confidence thereof output by the preset data element model comprises:

[0016] The obtaining of the candidate data word with the highest confidence and the confidence higher than the preset confidence threshold in the candidate data word corresponding to each target word type as the to-be-standardized data word.

[0017] The adding of the data word according to the non-target word type comprises:

[0018] The output of the non-target word type comprises:

[0019] The adding of the data word in the construction instruction to the standard data element corpus in response to the construction instruction input by the user based on the non-target word type.

[0020] The word type comprises an object class word, a representation word and a characteristic word; and the preset data element model comprises an object class word model, a characteristic word model and a representation word model.

[0021] The analysis of the target word type in the to-be-standardized data item by using the preset data element model comprises:

[0022] The input of the to-be-standardized data item into the object class word model, the obtaining of the target matching result of the object class word model, and the determination of the target word type comprising the object class word if the target matching result is successful.

[0023] The input of the to-be-standardized data item into the characteristic word model, the obtaining of the target matching result of the characteristic word model, and the determination of the target word type comprising the characteristic word if the target matching result is successful.

[0024] The input of the to-be-standardized data item into the representation word model, the obtaining of the target matching result of the representation word model, and the determination of the target word type comprising the representation word if the target matching result is successful.

[0025] The data element construction method further comprises:

[0026] The obtaining of the object class word corpus set from the standard data element corpus;

[0027] The extraction of several positive samples and several negative samples from the object class word corpus set to form an object class word training sample set;

[0028] The training of the object class word model by using the object class word training sample set;

[0029] The positive sample includes source data item information consistent with the object class word, and the negative sample includes source data item information inconsistent with the object class word.

[0030] The object class word corpus set includes a source application system name corpus, a Chinese data table name corpus, and a Chinese field name corpus.

[0031] The object class word model is trained by using the object class word training sample set.

[0032] The feature vector of the object class word training sample set is extracted, and the feature vector includes a source application system name word vector similarity, a Chinese data table name word vector similarity, and a Chinese field name word vector similarity.

[0033] The object class word model is trained by using the feature vector.

[0034] The object class word corpus set includes N object class word corpus groups, and each object class word corpus group includes a source application system name corpus, a Chinese data table name corpus, and a Chinese field name corpus.

[0035] The object class word model is trained by using the feature vector.

[0036] The feature vector of each object class word corpus group is obtained.

[0037] The feature vectors of the N object class word corpus groups are calculated according to a preset weight to obtain the confidence.

[0038] The object class word model is trained by using the confidence.

[0039] To solve the above technical problem, the application further provides a terminal device, which includes a memory and a processor coupled with the memory.

[0040] The memory is configured to store program data, and the processor is configured to execute the program data to implement the data element construction method.

[0041] To solve the above technical problem, the application further provides a computer storage medium, which is configured to store program data, and the program data is configured to implement the data element construction method when executed by a computer.

[0042] Compared with the prior art, the beneficial effects of the present application are that the terminal device acquires a to-be-standardized data item; analyzes a pair of standard word types in the to-be-standardized data item by using a preset data element model; acquires an unpaired standard word type according to the pair of standard word types; adds a data word according to the unpaired standard word type; and generates a standardized data element by using the to-be-standardized data item and the added data word. In this way, the newly-added data element can be efficiently, normatively and automatically constructed, and when the existing data element cannot meet the requirements, the function helps the user to quickly and normatively add the data element. BRIEF DESCRIPTION OF DRAWINGS

[0043] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0044] Among them:

[0045] Figure 1 is a flowchart of the first embodiment of the data element construction method provided by the present application;

[0046] Figure 2 is a flowchart of the second embodiment of the data element construction method provided by the present application;

[0047] Figure 3 is a schematic diagram of the standard data element corpus provided by the present application;

[0048] Figure 4 is a training process schematic diagram of the preset data element model provided by the present application;

[0049] Figure 5 is a flowchart of the third embodiment of the data element construction method provided by the present application;

[0050] Figure 6 is a sub-step flowchart of step S32 in the third embodiment of the data element construction method provided by the present application;

[0051] Figure 7 is a sub-step flowchart of step S14 in the first embodiment of the data element construction method provided by the present application;

[0052] Figure 8 is a sub-step flowchart of step S15 in the first embodiment of the data element construction method provided by the present application;

[0053] Figure 9 is a flowchart of the fourth embodiment of the data element construction method provided by the present application;

[0054] Figure 10 Figure 1 is a structural schematic diagram of an embodiment of a terminal device provided by the present application;

[0055] Figure 11 Figure 2 is a structural schematic diagram of an embodiment of a computer storage medium provided by the present application. DETAILED DESCRIPTION

[0056] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0057] The terms "first", "second", "third", "fourth" and the like (if any) in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in other than the order illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a list of steps or units does not necessarily limit those steps or units to the clearly listed ones, but can include other steps or units that are not clearly listed or inherent to those processes, methods, products or devices.

[0058] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present application.

[0059] First, the professional terms of the present application are introduced:

[0060] Data element: according to the definition of the national standard document "GB / T 18391.1-2009 Information Technology Metadata Registration System (MDR)", a data element (abbreviated as DE) is a data unit that specifies its definition, identification, representation and allowed values. In a certain context, a data element is usually used to construct an information unit of a specific concept semantics that is semantically correct, independent and unambiguous. A data element includes an object class word, a property word and a representation word.

[0061] Object class word: refers to the concept, abstract concept or collection of real world transactions that can be clearly identified and have the same rules for characteristics and behaviors.

[0062] Characteristic word: refers to the characteristics shared by all members of an object.

[0063] Representation word: representation is a description of the way data elements are expressed. Any change in any part of the various components of representation will result in a different representation, for example, the height of a person is measured in "centimeters" or "meters" as the unit of measurement, which are two different representations of the height of the person. The representation of data elements can be marked with some terms with representation meaning, such as name, code, amount, quantity, date, percentage, etc.

[0064] For details, see Figure 1 , Figure 1 is the flowchart of the first embodiment of the data element construction method provided by the present application.

[0065] The data element construction method of the present application is applied to a terminal device, wherein the terminal device of the present application can be a server, a local terminal, or a system cooperated by a server and a local terminal. Accordingly, each part of the terminal device, such as each unit, sub-unit, module, sub-module, can be all set in the server, all set in the local terminal, or set in the server and the local terminal respectively.

[0066] Further, the above-mentioned server can be hardware or software. When the server is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software or software modules, such as software or software modules for providing a distributed server, or as a single software or software module, which is not limited here.

[0067] As shown in Figure 1 , the specific steps are as follows:

[0068] Step S11: obtaining a data item to be standardized.

[0069] Specifically, the terminal device obtains a data item to be standardized from a business system. The business system can be a public security system, a bank system, an education system, etc. The data item to be standardized is a data item, a data list, etc. that needs to be standardized in the above-mentioned system, such as the naming of business name and business content.

[0070] In an embodiment of the present application, the data item to be standardized is a data item that cannot be standardized by the standard data element corpus. In other embodiments of the present application, the data item to be standardized can also be any data item that needs to be constructed and added to the standard corpus.

[0071] For example, in the field of public security, the standard data element corpus refers to the national standard document "GAT 543.X-2011 Public Security Data Elements". In the field of news publishing, the standard data element corpus refers to the national standard document "GB / T 40989-2021 News Publishing Knowledge Service Knowledge Object Identifier (KOI)".

[0072] Step S12: Analyzing the type of the reference word in the data item to be standardized using the preset data element model.

[0073] Specifically, the terminal device obtains the preset data element model through neural network training, and analyzes the type of the reference word in the data item to be standardized using the preset data element model.

[0074] An embodiment of the present application is proposed for training the preset data element model. For details, please refer to steps S21-S23 and steps S31-S32.

[0075] Further, in an embodiment of the present application, the word type includes object class words, representation words, and characteristic words. The preset data element model includes an object class word model, a characteristic word model, and a representation word model. The type of the reference word can be further determined by the above data element models, as follows:

[0076] The terminal device inputs the data item to be standardized into the object class word model to obtain the reference result of the object class word model. If the reference result is successful, it is determined that the type of the reference word includes object class words.

[0077] The terminal device inputs the data item to be standardized into the characteristic word model to obtain the reference result of the characteristic word model. If the reference result is successful, it is determined that the type of the reference word includes characteristic words.

[0078] The terminal device inputs the data item to be standardized into the representation word model to obtain the reference result of the representation word model. If the reference result is successful, it is determined that the type of the reference word includes representation words.

[0079] After inputting the data item to be standardized into the above three data element models, the reference word type of the data item to be standardized can be obtained by combining the reference results of the three data element models.

[0080] It should be noted that the terminal device can process steps S21-S23 simultaneously, that is, in an embodiment of the present application, the object class word model, the characteristic word model and the representation word model can be processed as a whole. After the terminal device inputs the to-be-standardized data item into the preset data element model, the preset data element model can be processed by the object class word model, the characteristic word model and the representation word model in the preset data element model respectively.

[0081] An embodiment of the present application is provided for obtaining a preset data element model. The method for training the preset data element in the embodiment can be applied to the creation process of any model. It should be noted that the training method of the characteristic word model and the training method of the representation word model are the same as the training method of the object class word model. For details, please refer to Figure 3 and Figure 4 Referring to Figure 2 , Figure 2 is a flowchart of a second embodiment of the data element construction method provided by the present application; Figure 3 is a schematic diagram of a standard data element corpus provided by the present application; Figure 4 is a schematic diagram of the training process of a preset data element model provided by the present application.

[0082] As Figure 2 shown, the specific steps are as follows:

[0083] Step S21: obtaining an object class word corpus set from a standard data element corpus.

[0084] The source of the standard data element corpus is a standard data element list of a corresponding business field, for example, in the field of public security, the standard data element corpus refers to the national standard document “GAT 543.X-2011 Public Security Data Element”, and in the field of news publishing, the standard data element corpus refers to the national standard document “GB / T 40989-2021 News Publishing Knowledge Service Knowledge Object Identifier (KOI)”.

[0085] Specifically, as Figure 3 shown, the terminal device obtains the object class word corpus set, the characteristic word corpus set and the representation word corpus set from the standard data element corpus by text preprocessing techniques such as word segmentation and stop word removal, and each corpus set contains a plurality of corpus groups.

[0086] In an embodiment of the present application, before obtaining the object class word corpus set from the standard data element corpus in step S21, the data information that has completed accurate standardization corresponding to each standard data element is sorted out.

[0087] Step S22: extracting a plurality of positive samples and a plurality of negative samples from the object class word corpus set to form an object class word training sample set.

[0088] Specifically, asFigure 4 As shown, the terminal device randomly extracts a plurality of positive samples and a plurality of negative samples from the accurate target data item information completed in the object class word corpus as the object class word training sample set, the characteristic word sample set, and the representation word sample set, which are used to train the object class word model, the characteristic word model, and the representation word model.

[0089] First, the selected sample set in the completed target data item for building the standard data element word vector library is used as the training set of each standard data element model.

[0090] For each object class word model, the positive sample of the training sample set is selected from the source data item information with the same object class word as the target data element, the negative sample of the training set is selected from the source data item information with the different object class word as the target data element, and the sample quantity is consistent with the positive sample through random sampling.

[0091] For each characteristic word model, the positive sample of the training sample set is selected from the source data item information with the same characteristic word as the target data element, the negative sample of the training set is selected from the source data item information with the different characteristic word as the target data element, and the sample quantity is consistent with the positive sample through random sampling.

[0092] For each representation word model, the positive sample of the training sample set is selected from the source data item information with the same representation word as the target data element, the negative sample of the training set is selected from the source data item information with the different representation word as the target data element, and the sample quantity is consistent with the positive sample through random sampling.

[0093] The object class word model, the characteristic word model, and the representation word model can be a logistic regression model or any other neural network training model.

[0094] Step S23: training the object class word model using the object class word training sample set.

[0095] Specifically, the terminal device trains the object class word model using the object class word training sample set, so that the trained object class word model can be used for the same word type data element.

[0096] Through the model training process of steps S21-S23, the data element model for similarity recognition of the object class word in the standardized data item can be obtained.

[0097] In an embodiment of the present application, the object class word training sample set is used for training. It should be noted that the training process of the characteristic word model and the representation word model is the same as that of the object class word model. For details, please refer to Figure 5 , Figure 5 is the flowchart of the third embodiment of the data element construction method provided by the present application.

[0098] Step S31: Extracting the feature vector of the object class word training sample set.

[0099] Wherein, each content segmentation corpus in each object class word set, characteristic word set and representation word set corresponds to each content segmentation corpus, and the terminal device constructs a word vector library through a text-to-word vector technology such as TF-IDF, word2vector algorithm, and each corpus finally obtains a word vector library.

[0100] Wherein, the feature vector includes source application system name word vector similarity, Chinese data table name word vector similarity and Chinese field name word vector similarity.

[0101] For each object class word model, the feature vector includes source application system name word text similarity, Chinese data table name corpus text similarity and Chinese field name corpus text similarity of the sample set and the corresponding object class word.

[0102] In other embodiments of the present application, for each characteristic word model, the feature vector includes source application system name text similarity, Chinese data table name corpus text similarity and Chinese field name corpus text similarity of the sample set and the corresponding characteristic word.

[0103] In other embodiments of the present application, for each representation word model, the feature vector includes data item content segmentation vector similarity of the sample set and the corresponding representation word, and data type one-hot vector encoding, wherein the vector similarity can be obtained by a series of methods such as cosine similarity.

[0104] Step S32: Training the object class word model using the feature vector.

[0105] Specifically, the terminal device trains the object class word model using the above feature vector. In other embodiments, the terminal device can train the characteristic word model and the representation word model simultaneously using the feature vector.

[0106] Wherein, the object class corpus set includes N object class corpus groups, and each object class corpus group includes a source application system name corpus, a Chinese data table name corpus and a Chinese field name corpus.

[0107] Through steps S31-S32, the accuracy of model training is further improved.

[0108] In order to further realize training the object class word model using the feature vector, the present application proposes steps S321 and S322 as sub-steps of step S32, please refer to Figure 6 , Figure 6is a sub-step flowchart of step S32 in the third embodiment of the data element construction method provided in the present application.

[0109] As shown in Figure 6 , the specific steps are as follows:

[0110] Step S321: Obtain the feature vector of each object class word corpus group.

[0111] Specifically, the terminal device constructs a word vector library through some text-to-word vector conversion technologies, such as TF-IDF, word2vector, etc. Each corpus finally gets a word vector library.

[0112] Step S322: Calculate the confidence according to the preset weight of the feature vectors of the N object class word corpus groups.

[0113] Specifically, please continue to refer to Figure 4 , the terminal device calculates the confidence Y of the feature vectors of the N object class word corpus groups according to the preset weight.

[0114] The preset weight is obtained by continuous training and optimization, or can be set by the staff through prior information.

[0115] The specific calculation formula is as follows: Y=a1X1+a2X2+a3X3+...+anXn n X n

[0116] Wherein, X1, X2...Xn represents the value of each feature vector of each object class word corpus group, a1, a2...an represents the weight of each feature vector of each object class word corpus group, Y represents the finally calculated confidence, the value range is [0, 1].

[0117] Step S323: Train the object class word model using the confidence.

[0118] Specifically, the terminal device calculates the loss value by the difference between the predicted confidence and the standard confidence, trains the object class word model using the loss value, and continuously optimizes the object class word model to make the predicted confidence continuously approach the standard confidence. After achieving the expected effect, the training of the object class word model can be completed. The training process of the characteristic class word model and the representation class word model is the same as that of the above object class word model, which will not be described in detail here.

[0119] Through steps S321-S322, the accuracy of model training is further improved.

[0120] Step S13: Obtain the non-target word type according to the target word type.

[0121] Specifically, the terminal device acquires the analysis result of the preset data element model, and acquires the non-successful keyword type according to the keyword type.

[0122] The keyword type includes an object class keyword, a representation keyword, and a feature keyword type. The non-successful keyword type is a keyword type that does not hit under the preset rule after the data item to be matched is input into the preset data element model.

[0123] For example, as shown in Table 1, the keyword type of the data item to be matched is the object class keyword and the representation keyword, and the non-successful keyword type is the feature keyword.

[0124] Hit rule Rule content Object class word Citizen Representation word Number

[0125] Table 1

[0126] In an embodiment of the present application, whether the keyword type is successfully matched is determined by the confidence and the preset confidence threshold. If the matching is unsuccessful, the corresponding keyword type is the non-successful keyword type.

[0127]

[0128]

[0129] Table 2

[0130] Further, as shown in Table 2, the keyword type is determined by the confidence and the preset confidence threshold. If the matching result is that all data items are completely matched, the object class keyword, the feature keyword, and the representation keyword returned are used to construct a new data element. If it is other conditions, step S14 is continued to be executed.

[0131] Step S14: Adding data keywords according to the non-successful keyword type.

[0132] Specifically, the terminal device acquires the analysis result of the preset data element model, and adds the non-successful keyword type. For example, please continue to refer to Table 1, that is, the case of the third item in Table 2. The non-successful keyword type is the feature keyword type, and the corresponding feature keyword is added to form a complete standardized data element.

[0133] Further, in order to realize the above-mentioned adding data keywords according to the non-successful keyword type and constructing a complete data element, the present application proposes steps S141-S142 as the sub-steps of step S14. Please refer to Figure 7 , Figure 7 is a sub-step flowchart of step S14 in the first embodiment of the data element construction method provided by the present application.

[0134] As shown in Figure 7 , the specific steps are as follows:

[0135] Step S141: outputting the untagged word type.

[0136] Specifically, the terminal device acquires the output result of the preset data element model, parses the output result, and outputs the untagged word type. The terminal device visualizes the untagged word type to guide the staff to add the word type of the data word.

[0137] Step S142: adding the data word in the construction instruction to the standard data element corpus in response to the construction instruction input by the user based on the untagged word type.

[0138] Specifically, the terminal device responds to the construction instruction input by the user, wherein the construction instruction input by the user is the data content input for the untagged word type. The terminal device adds the data word in the data content to the standard data element corpus as a data source for the next data element construction.

[0139] Through steps S141-S142, the existing data element resources are supplemented according to the output result, and the data element is constructed by reusing the object class word, the characteristic word, or the representation word of the tagged word type according to different situations, and the result is more efficient and accurate.

[0140] Step S15: generating a standardized data element by using the to-be-standardized data item and the added data word.

[0141] Specifically, the terminal device acquires the output result of the preset data element model, parses the output result, and outputs the untagged word type. The terminal device visualizes the untagged word type to guide the staff to add the word type of the data word.

[0142] Further, the terminal device can re-incorporate the newly generated standardized data element into the standardized data element corpus and update the standardized data element corpus, so that the standardized data element corpus can be directly reused in the next tagging process without the need to repeatedly construct the data element.

[0143] Through steps S11-S15, the newly added data element can be efficiently, normatively, and automatically constructed. When the existing data element cannot meet the requirements, the function helps the user to quickly and normatively add the data element.

[0144] For details, please refer to Figure 8 , Figure 8 is a sub-step flowchart of step S15 in the first embodiment of the data element construction method provided in the present application.

[0145] As shown in Figure 8 , the specific steps are as follows:

[0146] Step S151: obtaining the to-be-standardized data word output by the preset data element model according to the to-be-standardized data item's counterpart word type.

[0147] Specifically, the terminal device obtains the to-be-standardized data word output by the preset data element model according to the to-be-standardized data item's counterpart word type.

[0148] Step S152: combining the to-be-standardized data word and the added data word to generate a standardized data element.

[0149] Specifically, the terminal device combines the to-be-standardized data word and the added data word to generate a standardized data element according to the output result in Table 2.

[0150] Further, the generated standardized data element can be stored in a standard data element corpus, and can be directly called in the next data counterpart standardization, so that the increase of data element resources can in turn improve the use efficiency of data elements and strengthen the management and governance effect of data elements.

[0151] For example, the "ID number" data item can be based on the object class word "citizen" and the representation word "number" to construct a data element, which belongs to the result 3 in Table 2. Based on this condition, only one characteristic word "identity" needs to be added according to the construction requirement of the standard file based on the characteristic word and the method of step S142 to describe the characteristic word of the data element corresponding to the data item. Then a new data element Chinese name "citizen ID number" is finally obtained and stored in the standard data element corpus.

[0152] The application also provides an embodiment as a judgment standard, which judges whether the counterpart word type in the to-be-standardized data item is successfully standardized by comparing the confidence of the counterpart word type calculated by the preset data element model with the size of the preset confidence threshold. For details, please refer to Figure 9 , Figure 9 is a flowchart of the fourth embodiment of the data element construction method provided by the application.

[0153] As shown in Figure 9 , the specific steps are as follows:

[0154] Step S41: obtaining a plurality of candidate data words of the counterpart word type and their confidence output by the preset data element model.

[0155] Specifically, the terminal device obtains a plurality of candidate data words of the same counterpart word type output by the preset data element model, and the confidence of successful standardization of each candidate data word.

[0156] Through the above model, the feature word model can be used to judge the confidence of a data item belonging to the feature word. The representation word model can be used to judge the confidence of a data item belonging to the representation word. The object word model can be used to judge the confidence of a data item belonging to the object word.

[0157] Step S42: Obtain the candidate data word corresponding to each pair of mark word types with the highest confidence and higher than the preset confidence threshold as the to-be-standardized data word.

[0158] Specifically, the terminal device arranges the candidate data words corresponding to each pair of mark word types in the preset data element model in order of confidence, and selects the candidate data word with a confidence higher than the preset confidence threshold as the to-be-standardized data word.

[0159] For example, as shown in Table 3, after inputting the to-be-standardized data item into the preset data element model, the candidate data words with a confidence lower than 0.7 can be filtered out because the preset confidence is 0.7; then, one candidate data word with the highest confidence is selected from the candidate data words with a confidence higher than 0.7 as the to-be-standardized data word. For example, according to the result of Table 3, the terminal device selects "citizen" and "number" as the to-be-standardized data words.

[0160] Hit rule Rule content Confidence Object class word Citizen 0.981 Representation word Number 0.92 Representation word Encoding 0.82 Representation word Number 0.71

[0161] Table 3

[0162] Through steps S41-S42, through confidence judgment, part of the object class words, feature words or representation words can be reused to construct data elements according to different situations, and the result is more efficient and accurate.

[0163] To implement the above data element construction method, the present application further provides a terminal device, which is specifically described in Figure 10 , Figure 10 is a structural schematic diagram of an embodiment of the terminal device provided by the present application.

[0164] The terminal device 400 of the embodiment includes a processor 41, a memory 42, an input / output device 43 and a bus 44.

[0165] The processor 41, the memory 42 and the input / output device 43 are respectively connected with the bus 44, the memory 42 stores program data, and the processor 41 is used to execute the program data to realize the data element construction method described in the above embodiments.

[0166] In this embodiment, processor 41 can also be referred to as a CPU (Central Processing Unit). Processor 41 may be an integrated circuit chip with signal processing capabilities. Processor 41 can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor, or processor 41 can be any conventional processor.

[0167] This application also provides a computer storage medium; please refer to the following: Figure 11 , Figure 11 This is a schematic diagram of a computer storage medium 500 according to an embodiment of the present application. The computer storage medium 500 stores a computer program 51, which, when executed by a processor, is used to implement the data element construction method of the above embodiment.

[0168] When the embodiments of this application are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0169] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. A method for constructing data elements, characterized in that, The data element construction method includes: Obtain the data items to be standardized; Analyze the types of matching words in the data items to be standardized using a preset data element model; Obtain the types of non-matching words according to the types of matching words; Add data words according to the types of non-matching words; Generate standardized data elements using the data items to be standardized and the added data words; The adding data words according to the types of non-matching words includes: In response to the construction instruction input by the user based on the types of non-matching words, add the data words in the construction instruction to the standard data element corpus; The generating standardized data elements using the data items to be standardized and the added data words includes: According to the types of matching words of the data items to be standardized, obtain the data words to be standardized output by the preset data element model; Combine the data words to be standardized with the added data words to generate the standardized data elements.

2. The data element construction method according to claim 1, wherein The obtaining the data words to be standardized output by the preset data element model includes: Obtain a number of candidate data words and their confidence levels of the types of matching words output by the preset data element model; Obtain the candidate data word with the highest confidence level and a confidence level higher than the preset confidence threshold among the candidate data words corresponding to each type of matching word as the data word to be standardized.

3. The data element construction method according to claim 1, wherein The adding data words according to the types of non-matching words includes: Output the types of non-matching words; In response to the construction instruction input by the user based on the types of non-matching words, add the data words in the construction instruction to the standard data element corpus.

4. The data element construction method according to claim 1, wherein The types of words include object class words, representation words, and characteristic words; the preset data element model includes an object class word model, a characteristic word model, and a representation word model; The analyzing the types of matching words in the data items to be standardized using the preset data element model includes: Input the data items to be standardized into the object class word model, obtain the matching result of the object class word model, and if the matching result is successful, determine that the types of matching words include object class words; Input the data items to be standardized into the characteristic word model, obtain the matching result of the characteristic word model, and if the matching result is successful, determine that the types of matching words include characteristic words; Input the data items to be standardized into the representation word model, obtain the matching result of the representation word model, and if the matching result is successful, determine that the types of matching words include representation words.

5. The data element construction method according to claim 4, wherein The data element construction method further includes: Obtain the object class word corpus set from the standard data element corpus; Extract a number of positive samples and a number of negative samples from the object class word corpus set to form an object class word training sample set; Train the object class word model using the object class word training sample set; The positive samples include source data item information whose benchmark data elements are consistent with the class term of this object, and the negative samples include source data item information whose benchmark data elements are inconsistent with the class term of this object.

6. The data element construction method according to claim 5, characterized in that, The object-type word corpus collection includes a source application system name corpus, a Chinese data table name corpus, and a Chinese field name corpus; The step of training the object-class word model using the object-class word training sample set includes: Extract the feature vectors from the training sample set of the object class words, wherein the feature vectors include the similarity of word vectors of source application system names, the similarity of word vectors of Chinese data table names, and the similarity of word vectors of Chinese field names; The object-classification word model is trained using the feature vectors.

7. The data element construction method according to claim 6, characterized in that, The object-type word corpus set includes N object-type word corpus groups, and each object-type word corpus group includes a source application system name corpus, a Chinese data table name corpus, and a Chinese field name corpus. The step of training the object-class word model using the feature vectors includes: Obtain the feature vector of each object class word corpus group; Calculate the confidence level of the feature vectors of N object class word corpus groups according to preset weights; The object class word model is trained using the confidence score.

8. A terminal device, characterized in that, The terminal device includes a memory and a processor coupled to the memory; The memory is used to store program data, and the processor is used to execute the program data to implement the data element construction method as described in any one of claims 1 to 7.

9. A computer storage medium, characterized in that, The computer storage medium is used to store program data, which, when executed by the computer, is used to implement the data element construction method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Benchmarking method and system for data items, files and databases

    CN110196834A

  • Data benchmarking method and device and storage device

    CN110795482A