Processing method for tubular standard data
By obtaining the descriptive information of data standard items and using field matching and similarity matching conditions, the problem of low efficiency in standardization data processing is solved, and fast and accurate data standardization operations are achieved.
Patent Information
- Application Number
- CN202510863363.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-10-17
AI Technical Summary
Existing technologies are inefficient in processing standardized data, especially when faced with multiple associations, making it difficult to quickly and accurately match non-standardized field names with data standards.
By obtaining the description information of the data standard items and using field matching conditions and similarity matching conditions, the target data standard items are determined and data standardization operations are performed based on them. Specific methods include string similarity calculation, word segmentation hit rate analysis, and the application of text prediction models to improve the matching success rate.
It realizes the automatic and accurate processing of standard-compliant data, improves the matching success rate, reduces computing resource consumption, and improves processing efficiency.
Smart Images

Figure CN120805905A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of information security, in particular to a processing method for standardization data. BACKGROUND
[0002] In the era of big data, data standardization and standardization work are crucial to ensure data consistency and improve data sharing capabilities. However, when dealing with large-scale data sets, there is often a problem of how to quickly and accurately match non-standardized field names with existing data standards. This process is not only complex, but also time-consuming, especially when there are multiple association relationships between field names and data standards.
[0003] Existing data standard recommendation techniques mainly rely on two methods: keyword matching technology and recommendation algorithm based on shallow neural network. However, keyword matching relies too much on the surface consistency of field names and standard names, ignoring deep information such as field descriptions, purposes, and standard aliases and definitions. The limitations of training data and the lack of context understanding may result in inaccurate word vectors generated by the model, that is, there is a technical problem of low processing efficiency of standardization data in the prior art.
[0004] In view of the inaccurate and low-efficiency processing of standardization data in the related art, no effective solution has been proposed so far. SUMMARY
[0005] The main purpose of the present application is to provide a processing method for standardization data to solve the problem of low processing efficiency of standardization data in the related art.
[0006] In order to achieve the above purpose, according to one aspect of the present application, a processing method for standardization data is provided. The method comprises: obtaining data standard description information corresponding to a plurality of data standard items respectively; in the case that the fields of the standardization data and the fields in the data standard description information of the first data standard item satisfy the field matching condition, determining that the first data standard item is the target data standard item, wherein the field matching condition is that the matching result between the fields is greater than a first threshold; in the case that the standardization data and the data standard description information of the second data standard item satisfy the similarity matching condition, determining that the second data standard item is the target data standard item, wherein the similarity matching condition is determined according to the word segmentation hit rate between the standardization data and the data standard description information; based on the data standard description information of the target data standard item, performing a data standardization operation on the standardization data.
[0007] To achieve the above object, according to another aspect of the present application, a processing device for cross-standard data is provided. The device comprises: an acquisition unit, which acquires data standard description information corresponding to a plurality of data standard items respectively; a first determination unit, which determines a first data standard item as a target data standard item in a case where a field of the cross-standard data and a field in the data standard description information of the first data standard item satisfy a field matching condition, wherein the field matching condition is that a matching result between the fields is greater than a first threshold; a second determination unit, which determines a second data standard item as the target data standard item in a case where the cross-standard data and the data standard description information of the second data standard item satisfy a similarity matching condition, wherein the similarity matching condition is determined according to a word segmentation hit rate between the cross-standard data and the data standard description information; and a processing unit, which performs a data standardization operation on the cross-standard data based on the data standard description information of the target data standard item.
[0008] Optionally, the first determination unit comprises a matching module, which is configured to determine that the cross-standard data and the first data standard item satisfy the field matching condition in a case where a name field of the cross-standard data is the same as a standard name field indicated by the data standard description information of the first data standard item, and determine the first data standard item as the target data standard item; and determine that the cross-standard data and the first data standard item satisfy the field matching condition in a case where at least one reference field in the name field of the cross-standard data is the same as at least one standard name field indicated by the data standard description information of the first data standard item, and determine a key field in the name field of the cross-standard data, and determine the target data standard item in the at least one first data standard item according to the key field.
[0009] Optionally, the second determination unit comprises a calculation module, which is configured to determine a cross-standard data word segmentation set matched with the cross-standard data and a plurality of data standard word segmentation sets respectively matched with the plurality of data standard items; and calculate a word segmentation hit rate between the cross-standard data word segmentation set and each data standard word segmentation set in the plurality of data standard word segmentation sets in sequence based on a weight value corresponding to each word segmentation by using a similarity algorithm.
[0010] Optionally, the matching module is further configured to determine that a field at a target position in the name field of the cross-standard data is a key field, and determine that a first data standard item corresponding to a standard name field matched with the key field is the target data standard item; and determine that a field corresponding to a part of speech with a number of repetitions greater than a third threshold in the name field of the cross-standard data is a key field, and determine that a first data standard item corresponding to a standard name field matched with the key field is the target data standard item.
[0011] Optionally, the second determination unit is further configured to determine a data standard item with a word segmentation hit rate greater than a second threshold as the second data standard item; and determine the second data standard item corresponding to the highest word segmentation hit rate as the target data standard item.
[0012] Optionally, the second determining unit is further configured to determine a feature of the standard data matched with the standard data; the text prediction model outputs a prediction probability of a data standard item matched with the standard data based on the feature of the standard data; and the data standard item with the prediction probability greater than a fourth threshold is determined as the target data standard item.
[0013] Optionally, the first determining unit comprises a selection module configured to determine a target dictionary based on the plurality of data standard items; and construct a graph structure matched with the standard data based on the target dictionary, wherein each node on the graph structure represents a vocabulary, and edges between the nodes are used to indicate that the vocabularies have a combination relationship; and determine a target path based on the graph structure, and splice the vocabularies on the target path to obtain the field of the standard data.
[0014] In the embodiments of the present application, the data standard description information corresponding to the plurality of data standard items is obtained; in the case that the field of the standard data and the field in the data standard description information of the first data standard item satisfy a field matching condition, the first data standard item is determined as the target data standard item, wherein the field matching condition is that the matching result between the fields is greater than a first threshold, and the fast matching is directly performed based on the fields; in the case that the standard data and the data standard description information of the second data standard item satisfy a similarity matching condition, the second data standard item is determined as the target data standard item, wherein the similarity matching condition is determined according to the segmentation hit rate between the standard data and the data standard description information, which can effectively process the synonyms and variant words, and improve the success rate of matching; the purpose of performing the data standardization operation on the standard data based on the data standard description information of the target data standard item is achieved, thereby realizing the technical effect of automatically and accurately processing the standard data, and further solving the technical problem of low processing efficiency of the standard data. BRIEF DESCRIPTION OF DRAWINGS
[0015] The accompanying drawings, which form a part of the present application, are intended to provide further understanding of the present application, and are used to interpret the illustrative embodiments of the present application and their descriptions, and do not constitute improper limitations to the present application. In the drawings:
[0016] Figure 1 Fig. 1 shows a hardware structure block diagram of a computer terminal for implementing the processing method of the standard data;
[0017] Figure 2 Fig. 2 is a flowchart of the processing method of the standard data according to an embodiment of the present application;
[0018] Figure 3 Fig. 3 is a flowchart of the processing method of the standard data according to an embodiment of the present application;
[0019] Figure 4is a schematic diagram of a processing device for the data of the national standard provided by the embodiment of the application;
[0020] Figure 5 is a structural block diagram of an electronic device according to the embodiment of the application. DETAILED DESCRIPTION
[0021] In order to make the personnel in the technical field better understand the scheme of the present application, the technical scheme in the embodiments of the present application will be clearly and completely described below in combination with the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor should belong to the scope of protection of the present application.
[0022] It should be noted that the terms "first", "second" and the like in the specification and claims of the present application and the above-described drawings are used to distinguish similar objects, and do not necessarily indicate a specific order or a chronological sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0023] Embodiment 1
[0024] According to the embodiments of the present application, a method embodiment for processing the data of the national standard is also provided. It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that herein.
[0025] The method embodiment provided by the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 A hardware structural block diagram of a computer terminal (or mobile device) for implementing the method for processing the data of the national standard is shown. As shown in Figure 1As shown, the computer terminal 10 (or mobile device) can include one or more processors 102 (the processor 102 can include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it can also include a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which can be included as one of the ports of the BUS bus), a network interface, a power supply and / or a camera. Those skilled in the art can understand that Figure 1 The structure shown is only schematic, which does not limit the structure of the above-mentioned electronic device. For example, the computer terminal 10 can include more or fewer components than those shown in Figure 1 or have a different configuration than that shown in Figure 1 .
[0026] It should be noted that the one or more processors 102 and / or other data processing circuits described above can be referred to herein as "data processing circuits" in general. The data processing circuit can be embodied in whole or in part as software, hardware, firmware or any combination thereof. In addition, the data processing circuit can be a single independent processing module, or any one of the other elements incorporated into the computer terminal 10 (or mobile device) in whole or in part. As referred to in the embodiments of the present application, the data processing circuit serves as a processor to control (for example, selection of a variable resistance terminal path connected to an interface).
[0027] The memory 104 can be used to store software programs and modules of application software, such as program instructions / data storage means corresponding to the processing method of the fiducial data in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned processing method of the fiducial data. The memory 104 can include a high-speed random access memory, and can also include a non-volatile memory such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some examples, the memory 104 can further include a memory remotely disposed with respect to the processor 102, which can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0028] The transmission device 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission device 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission device 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.
[0029] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 (or mobile device).
[0030] Under the above operating environment, this application provides Figure 2 The method of processing the standardization data is shown in FIG. Figure 2 This is a flowchart of the method for processing standardization data according to Example 1 of the present application.
[0031] Step S101, obtaining data standard description information corresponding to a plurality of data standard items;
[0032] Step S102: If a field of the standard-compliant data and a field in the data standard description information of the first data standard item meet a field matching condition, determining the first data standard item as a target data standard item, wherein the field matching condition is that a matching result between the fields is greater than a first threshold;
[0033] Step S103: if the standard-compliant data and the data standard description information of the second data standard item meet a similarity matching condition, determining the second data standard item as the target data standard item, wherein the similarity matching condition is determined based on a word segmentation hit rate between the standard-compliant data and the data standard description information;
[0034] Step S104: performing data standardization operations on the standard-compliant data based on the data standard description information of the target data standard item.
[0035] As an optional implementation, in the above step S101, data standard description information corresponding to multiple data standard items is obtained; the above data standard items can be, for example, a collection of standard data elements or fields defined within an organization or industry, and each item contains a detailed definition and usage rules for a specific data element, usually covering the Chinese name, alias, definition, data type, data length and other attributes of the field; the above data standard description information is a complete description of each data standard item, including but not limited to the standard Chinese name, standard alias, standard definition, data type, data length, business rules, etc.
[0036] Further, in step S102, in a case where the field of the data standard data and the field in the data standard description information of the first data standard item satisfy a field matching condition, the first data standard item is determined as the target data standard item, wherein the field matching condition is that a matching result between the fields is greater than a first threshold value.
[0037] In an optional embodiment, step S102 is described, for example, the data standard data is "order payment status", and the Chinese name, alias and the like of the field of all related data standard items in the data standard library are compared, for example, the field alias list of the data standard item "transaction status" may contain "payment status"; the matching result of the "order payment status" field and the field alias "payment status" of the "transaction status" data standard item is calculated, which can be obtained by simple string similarity calculation, such as edit distance, Levenshtein distance or regular expression matching method; assuming that in the above example, the matching score of "order payment status" and "payment status" field is 0.85 obtained by comparison algorithm, if the first threshold value is set to 0.8, the score exceeds the first threshold value, indicating that "order payment status" and "transaction status" data standard item have high matching degree; further, the "transaction status" data standard item is determined as the "target data standard item" to guide the subsequent data standardization operation.
[0038] Optionally, the above-mentioned field matching condition can be a string complete match: such as "order payment status" and the field name "order payment status" in the data standard item are completely consistent; alias matching: such as "order payment status" and the alias "payment status" in the data standard item are matched; keyword matching: the system identifies the keywords "payment" and "status" in the field "order payment status", and then searches the data standard description information to contain the same keywords to achieve indirect matching. The setting of the above-mentioned first threshold value can be determined according to experience or through machine learning, which is not specifically limited here.
[0039] The user information and data involved in the present application are information and data authorized by the user. Specifically, the user information and data (including but not limited to data standard data and data standard) involved in the present application are information and data authorized by the user or sufficiently authorized.
[0040] Optionally, in the above-mentioned step S103, in a case where the data standard data and the data standard description information of the second data standard item satisfy a similarity matching condition, the second data standard item is determined as the target data standard item, wherein the similarity matching condition is determined according to the word segmentation hit rate between the data standard data and the data standard description information. Optionally, the above-mentioned word segmentation hit rate is determined by using the Jaccard similarity coefficient formula, word vector similarity calculation and the like, which is not specifically limited here.
[0041] The step S103 is described in an optional embodiment, taking the "promotion effect evaluation" field as an example, which is used to evaluate the performance of the promotion activity, such as sales, customer feedback, etc. There is no direct matching field name in the data standard library, but there is a data standard item "sales activity performance" which describes the sales performance, market reaction and other information. Through Chinese word segmentation and stop word filtering, "promotion effect evaluation" is converted into the vocabulary set A={promotion, effect, evaluation}, and "sales activity performance" is converted into the vocabulary set B={sales, activity, performance}, their intersection is {sales, activity}, assuming that the company sets the similarity matching threshold T to be 0.3, and the Jaccard similarity coefficient calculated by the formula is 0.4 (i.e. 2 / 5), which meets the matching condition, and the "promotion effect evaluation" field will be standardized according to the specification requirements of the "sales activity performance" data standard.
[0042] As an optional embodiment, in the above step S104, the data standardization operation is performed on the data based on the data standard description information of the target data standard item. Specifically, it can refer to the process of adjusting the non-standard format of the data to be consistent with the data format, type and length and other attributes specified by the selected target data standard item (first or second data standard item), for example, the target data standard item requires the data type to be "date and time", and the original field stores the date in the form of a string, which needs to be converted to date and time format; mapping the data values in the fields of the data to ensure that they meet the enumeration values, codes or categories defined in the target data standard item, for example, mapping "male" and "female" in the "gender" field to "M" and "F" in the standard definition, respectively, without specific limitation here.
[0043] Through the above embodiment, the data standard description information corresponding to a plurality of data standard items is obtained; in the case that the fields of the data meet the field matching condition in the data standard description information of the first data standard item, the first data standard item is determined as the target data standard item, wherein the field matching condition is that the matching result between the fields is greater than the first threshold, and the fast matching is directly based on the fields; in the case that the data meets the similarity matching condition with the data standard description information of the second data standard item, the second data standard item is determined as the target data standard item, wherein the similarity matching condition is determined according to the word segmentation hit rate between the data and the data standard description information, which can effectively handle synonyms and variant words and improve the success rate of matching; the purpose of performing data standardization operation on the data based on the data standard description information of the target data standard item is achieved, thereby realizing the technical effect of automatically and accurately processing the data, and further solving the technical problem of low processing efficiency of the data.
[0044] Optionally, in the method for processing the data of the standard mark provided in the embodiments of the present application, in the case where the field of the data of the standard mark and the field in the data standard description information of the first data standard item meet the field matching condition, the first data standard item is determined as the target data standard item, and the method further comprises:
[0045] S1, in the case where the name field of the data of the standard mark is the same as the standard name field indicated by the data standard description information of the first data standard item, it is determined that the data of the standard mark and the first data standard item meet the field matching condition, and the first data standard item is determined as the target data standard item.
[0046] S2, in the case where at least one reference field in the name field of the data of the standard mark is the same as the standard name field indicated by the data standard description information of at least one first data standard item, it is determined that the data of the standard mark and the first data standard item meet the field matching condition, a key field in the name field of the data of the standard mark is determined, and the target data standard item is determined in the at least one first data standard item according to the key field.
[0047] In the step S1, in the case where the name field of the data of the standard mark is the same as the standard name field indicated by the data standard description information of the first data standard item, it is determined that the data of the standard mark and the first data standard item meet the field matching condition, and the first data standard item is determined as the target data standard item.
[0048] As an optional embodiment, for example, the standard name field of the “commodity unique identifier” standard item in the data standard library is explicitly indicated as “commodity number”, which is the same as the field name of the data of the standard mark that needs to be processed, the field is matched with the “commodity unique identifier” standard item, and the data of the standard mark is processed according to the data requirement of the “commodity unique identifier” standard item subsequently. The above-mentioned instant matching strategy simplifies the search process, avoids traversal and similarity calculation of all data standard items, and thus significantly improves the processing speed.
[0049] Further in the step S2, in the case where at least one reference field in the name field of the data of the standard mark is the same as the standard name field indicated by the data standard description information of at least one first data standard item, it is determined that the data of the standard mark and the first data standard item meet the field matching condition, a key field in the name field of the data of the standard mark is determined, and the target data standard item is determined in the at least one first data standard item according to the key field.
[0050] As an optional implementation, for example, the Chinese name of the field of the data of the standard can be divided into the word group of "channel / transaction / date", that is, the above-mentioned reference fields are "channel", "transaction", and "date", and the last word "date" is further selected as the basic keyword of the standard word, that is, the above-mentioned key field, and then the next operation is performed to filter out the wrong recommended standard "channel type" caused by the modifier "channel".
[0051] Through the above technical features, for those field names that are not completely consistent with the standard name but partially match, by identifying the reference field in the name field (that is, the word partially coinciding with the standard name field), the search range can be quickly narrowed down, and the data standard items with higher relevance are focused on. Further, based on the key field, the target standard item can be accurately located, blind search in all data standard items is avoided, a large amount of computing resources is saved, and the matching efficiency and accuracy are improved.
[0052] Optionally, in the method for processing the data of the standard provided in the embodiments of the present application, before the second data standard item is determined as the target data standard item, in the case that the data of the standard and the data standard description information of the second data standard item satisfy the similarity matching condition, the method further comprises:
[0053] S1, determining a standard data word set matched with the data of the standard and a plurality of data standard word sets respectively matched with the plurality of data standard items;
[0054] S2, calculating the word hit rate between the standard data word set and each data standard word set in the plurality of data standard word sets in sequence based on the weight value corresponding to each word by using a similarity algorithm.
[0055] Optionally, in the above steps S1-S2, the standard data word set and the data standard word set are generated by word segmentation, and for the words in the standard data word set and each data standard word set, a corresponding weight value is given, which can be determined based on the common degree of the words, the position of the words in the description, and the relevance to the business field or scene, and the like. For example, if the field description of the data of the standard is "record the time of each user login", the weight of core words such as "user", "login", and "time" can be higher than that of other words; and for the data standard item description, if "user behavior", "login event", and "timestamp" are mentioned, the corresponding word weight will also be increased accordingly.
[0056] Further, based on the word set and the corresponding weight value, a similarity algorithm (such as Jaccard similarity, cosine similarity, edit distance, etc.) is used to calculate the word hit rate between the standard data word set and each data standard word set, and the Jaccard similarity formula is as follows:
[0057]
[0058] wherein A and B are the standard data field set and one of the data standard field set respectively.
[0059] Specifically, assuming that the standard data field description is "record the order amount of online shopping of users", the tokenization tool is used for tokenization, and the standard data field set {record, user, online shopping, order amount} is obtained. The description information of each data standard item is processed in the same way: the description of the first data standard item is "total consumption of users", and the tokenization set is {user, consumption, total amount}; the description of the second data standard item is "online sales", and the tokenization set is {online, sales, amount}.
[0060] Then, the tokenization hit rate between the standard data field set and the two data standard field sets is calculated. Assuming that the weights of the words "user", "online shopping" and "order amount" are 0.3, 0.4 and 0.6 respectively, in the first data standard item, "user" and "total amount" hit, and "consumption" and "order" have some relevance in business logic, but do not hit directly in the dictionary; in the second data standard item, "online shopping" and "amount" hit, and "sales" and "order" also have relevance in business.
[0061] The calculation result is as follows: the tokenization hit rate with the first data standard item tokenization set is 0.5 (considering the high weights of "user" and "total amount" and the relevance in business). The tokenization hit rate with the second data standard item tokenization set is 0.6 ("online shopping" and "amount" have high weights, plus the business logic score of "sales" and "order"). Further, the "online sales" data standard item can be determined as the target data standard item of the "record the order amount of online shopping of users" standard data field. User feedback and historical matching data can also be combined to automatically adjust the word weight, so as to more quickly identify key information in future standardization and improve the matching speed.
[0062] Through the above embodiments described in the present application, by tokenization and weight assignment, the calculation focuses on the words that are most likely to match, avoiding indiscriminate comparison of all standard items, and significantly improving the processing efficiency. At the same time, the similarity algorithm based on word weight can more accurately evaluate the semantic similarity between the standard data and the standard items, reducing the possibility of false matching.
[0063] Optionally, in the method for processing the standardization data provided in the embodiments of the present application, the key field in the name field of the standardization data is determined, and the target data standard item is determined in at least one first data standard item according to the key field, including at least one of the following:
[0064] In the first mode, a field at a target position in the name field of the data is determined as a key field, and a first data standard item corresponding to a standard name field matched with the key field is determined as a target data standard item.
[0065] In the second mode, a field corresponding to a part of speech with a repetition number greater than a third threshold value in the name field of the data is determined as a key field, and a first data standard item corresponding to a standard name field matched with the key field is determined as a target data standard item.
[0066] Optionally, in the first mode, a field at a target position in the name field of the data is determined, for example, in the "customer account information" of a record, the key field "account type" can be located at the beginning or end of the record, and further in the data standard library, a standard name field matched with the key field "account type" is searched, where the "standard name field" refers to a field in a data standard item used to identify and describe a standard, for example, "account type" in a data standard item "account classification code", if the matching is successful, "account classification code" is determined as the target data standard item of the "account type" field, and the standardization operation is performed on the "account type" field, including data type conversion, data format unification, enumeration value mapping, etc.
[0067] The first mode quickly locks the target data standard item by locating the key field in the name field, avoiding detailed comparison of the entire field name and its description, and the key field is usually the most representative and distinguishing part of the field name, which is used as the focus of matching, thereby greatly reducing the search time and consumption of computing resources.
[0068] Optionally, in the second mode, first, the name field of the data is subjected to natural language processing, and a part of speech tagging technology is used to identify the grammatical role of each word, for example, in the field "diabetes diagnosis date and hypertension medication record", "diabetes" and "hypertension" can be marked as disease name (DiseaseName), and "diagnosis date" and "medication record" can be marked as date (Date) and description (Description), respectively. Based on the part of speech tagging, the repetition number of each part of speech in the field is counted, for example, if the repetition number of "DiseaseName" in the above field is the largest, reaching more than 3 (assuming that the third threshold value is 3), the data standard item matched with the disease name (DiseaseName) is determined as the target data standard item, and the above process is not limited.
[0069] The second mode focuses on the statistics of the number of word forms, that is, the word form frequently appearing in the field name can represent an important category or attribute, and the core vocabulary is identified through statistical analysis, thereby reducing unnecessary field comparison, and especially in the case of processing complex field names and multiple data standards, the data standard item with high matching degree can be quickly identified, and the processing efficiency is greatly improved.
[0070] It should be further noted that the first mode and the second mode can be combined to form a more powerful search framework. For example, the first mode is used to quickly locate the possible matching item, and the second mode is used to further optimize and confirm the final matching result through word form analysis. This multi-mode fusion can balance speed and accuracy, and provides a comprehensive solution for processing inefficient data.
[0071] Optionally, in the method for processing the data provided in the embodiments of the present application, the second data standard item is determined as the target data standard item, comprising:
[0072] S1, determining the data standard item with a word segmentation hit rate greater than a second threshold as a second data standard item;
[0073] S2, determining the second data standard item corresponding to the highest word segmentation hit rate as the target data standard item.
[0074] In the above steps S1-S2, based on the word segmentation hit rate, the data standard item with a hit rate higher than the second threshold is determined as the candidate second data standard item list, and then the standard item with the highest word segmentation hit rate is selected from the second data standard item as the target data standard item.
[0075] Optionally, in the method for processing the data provided in the embodiments of the present application, in the case that the data standard description information of the data and the second data standard item satisfies the similarity matching condition, after determining the second data standard item as the target data standard item, the method further comprises:
[0076] S1, determining the data standard item matching the data;
[0077] S2, the text prediction model outputs a prediction probability of the data standard item matching the data based on the data characteristics;
[0078] S3, determining the data standard item with a prediction probability greater than a fourth threshold as the target data standard item.
[0079] In the step S1, the features of the data include, but are not limited to, field name, field description, data type, data source, and business scenario, etc. For example, the data with the field name "transaction amount" and the field description "the final payment amount of the customer for purchasing goods" can have the features of "transaction", "amount", "numeric", and "finance". It should be noted that the process can further mine semantic features, encode the text by using a word vector model (such as Word2Vec or BERT) to capture the relationship between words and the meaning of the context, and introduce an entity recognition technology to automatically find the entity categories mentioned in the field description, such as "customer", "goods", or "payment".
[0080] In the steps S2-S3, the input of the text prediction model is the vector representation of the data features, and the output is the prediction probability of each predefined data standard item. When the prediction probability of multiple data standard items is higher than the fourth threshold, an additional decision mechanism can be adopted to select the final target data standard item, for example, selecting the standard item with the highest prediction probability, or considering the prediction probability of multiple standard items by weighted average, and selecting the item with the highest comprehensive score.
[0081] Through the above technical features described in the present application, the text prediction model learns the mapping relationship between the data standard items and the data features in the training process, and can further output the matching probability of a series of data standard items according to the input data features, thereby improving the processing efficiency of the data.
[0082] Optionally, in the method for processing the data provided in the embodiments of the present application, before determining that the first data standard item is the target data standard item, the method further comprises:
[0083] S1, determining a target dictionary based on the plurality of data standard items;
[0084] S2, constructing a graph structure matched with the data based on the target dictionary, wherein each node on the graph structure represents a word, and the edges between the nodes are used to indicate that the words have a combination relationship;
[0085] S3, determining a target path based on the graph structure, and splicing the words on the target path to obtain the field of the data.
[0086] Optionally, in the step S1, the plurality of standard item names are analyzed to extract the basic words constituting the field name, such as "account", "balance", "transaction", "date", etc., and the target dictionary is constructed.
[0087] In step S2, for each field to be labeled, a word graph is generated by using a word segmentation tool to scan the target dictionary. The nodes of the graph represent the words in the dictionary, and the edges between the nodes represent the possible combination relationships between the words. For example, for the field "transaction amount date", there can be multiple edges such as "transaction"→"amount" and "amount"→"date". After the word graph is constructed, it can be parsed to identify all possible combination sequences of the words in the field.
[0088] In step S3, a dynamic programming algorithm is applied to find the optimal path from the root node to the leaf node in the DAG. The optimal path can be the one with the highest probability of word combination along the path, which is determined by the word frequency and the context. The algorithm considers all possible paths and selects the one with the maximum total probability as the target path. Further along the target path, the node words are spliced to form the recommended data field name. For example, for the word graph of "transaction amount date", if the optimal path is "transaction"→"amount", then the recommended field name is "transaction amount".
[0089] Through the above embodiments described in the present application, the determination of the target path in the graph structure can use algorithms such as the shortest path algorithm or the probability-based path selection algorithm. According to the connection strength between the nodes and the business importance of the words, a path from the initial word node of the input field to the final suggested field name can be quickly found. The algorithm collects the node words along this path and splices them in order to form a complete field name. This not only ensures the rationality of the field name, but also greatly improves the generation speed and matching speed, thereby improving the processing efficiency of the labeled data.
[0090] The following is a complete description of the present scheme with a flowchart Figure 3 as an example:
[0091] Start the recommendation process. In step S302, determine whether the labeled field is the same as the standard name or alias. If the names are exactly the same, execute S302-1 and recommend the standard.
[0092] If they are not exactly the same, determine in step S304 whether the labeled field name contains the standard name or alias. If there is a containing relationship, execute S306 to perform word segmentation on the labeled field and obtain the last word as the basic keyword. In step S308, match the data standard containing the basic keyword as a candidate. In step S310, calculate the Jaccard similarity coefficient between the labeled field and the candidate data standard.
[0093] Specifically, the above process is the matching process in the case that the field of the standard is strongly related to the data standard. The first type is that the field name of the standard is the same as the standard name. The data standard list is traversed once. If there is a matching item that meets the condition, the data standard is directly recommended. The second type is that the field name of the standard contains the standard name or the standard alias. Since a field name can contain multiple standard names, the error recommendation caused by the adjectival words needs to be filtered. The Chinese name of the field of the standard is segmented by using a Chinese segmentation algorithm and the stop words are filtered. For example, the Chinese name of the field "channel transaction date" can be divided into the word group "channel / transaction / date". The next operation is performed to filter out the error recommended standard "channel type" caused by the adjectival word "channel".
[0094] For the obtained basic keyword, the data standard whose name or alias contains the keyword is screened in the data standard information library. The similarity algorithm described in the related art is used to judge the obtained data standard. The data standard with the highest similarity is recommended in the standard higher than the set threshold value (the threshold value is determined based on the recommendation accuracy in the test data. Generally, it can be considered that the recommendation higher than the threshold value has higher reliability). If the similarity between all screened data standards and the field of the standard is lower than the threshold value, the method is not used for recommendation. The process enters the recommendation process of the weak correlation between the field of the standard and the data standard.
[0095] It should be noted that the Chinese segmentation algorithm used in the present scheme is to realize efficient word graph scanning according to the prefix dictionary, generate a directed acyclic graph composed of all possible word segmentation conditions in a sentence, and then find the maximum probability by dynamic programming to obtain the maximum segmentation combination based on the word frequency. Therefore, in order to ensure the accuracy of the division result for the recommendation model, a related dictionary needs to be added to the existing data standard name.
[0096] If the name or alias is not contained, S304-1 is executed, and the Jaccard similarity coefficient of the field of the standard and the standard in the data standard list is calculated in sequence. It is judged in S312 whether the maximum value of the similarity coefficient exceeds the threshold value. If the threshold value is exceeded, S312-1 is executed, and the data standard with the maximum similarity coefficient is recommended.
[0097] Specifically, the above process is the matching process in the case of weak correlation between the field and the data standard. In the case of weak correlation between the field and the data standard, the case of medium correlation between the field and the data standard is considered. In this method, the data standard list is traversed, and the similarity between the field and the data standard is calculated in sequence. The data standard with the highest similarity in the data standard with a similarity higher than the recommended threshold is recommended. The similarity is calculated by generating two word sets from the Chinese name and field description of the field and the Chinese name, alias and standard definition of the data standard through a word segmentation algorithm, then filtering the stop words in the two sets, and finally calculating the Jaccard similarity coefficient of the two sets. The recommended threshold is mainly determined according to the similarity of the successfully recommended items in the sample data. The threshold is selected as a reference for recommending with high reliability when the similarity is higher than the threshold. If the similarity with all data standard items is lower than the threshold after calculation, the method is abandoned.
[0098] If it does not exceed the threshold, S314 is performed to segment the field Chinese name and field description information and filter the stop words; S316, the segmented result is vectorized through a pre-trained word2vec model; S318, a matrix composed of multiple word vectors is input into a TextCNN network (short text classification model) to obtain the probability value falling on each standard; it is judged S320 whether the maximum probability is greater than the set threshold? If it is greater than the set threshold, S322 is performed to recommend the standard with the maximum probability;
[0099] Specifically, the above process is the matching process in the case of weak correlation between the field and the data standard. In the case of weak correlation between the field and the data standard, the case of medium correlation between the field and the data standard is considered. In this method, the data standard list is traversed, and the similarity between the field and the data standard is calculated in sequence. The data standard with the highest similarity in the data standard with a similarity higher than the recommended threshold is recommended. The similarity is calculated by generating two word sets from the Chinese name and field description of the field and the Chinese name, alias and standard definition of the data standard through a word segmentation algorithm, then filtering the stop words in the two sets, and finally calculating the Jaccard similarity coefficient of the two sets. The recommended threshold is mainly determined according to the similarity of the successfully recommended items in the sample data. The threshold is selected as a reference for recommending with high reliability when the similarity is higher than the threshold. If the similarity with all data standard items is lower than the threshold after calculation, the method is abandoned.
[0100] Firstly, the input information composed of field Chinese name and field description needs to be processed in the input layer. The input text is divided into several words by word segmentation, and the word vectors are generated by the pre-trained Word2Vec model after filtering stop words. The vectors can capture the semantic relationship between words. Then the matrix composed of multiple word vectors is convolved by the convolution kernel, and the feature vector is obtained through a layer of max pooling. Finally, the feature vector is mapped to the label domain through the full connection layer, and the probability belonging to each class is obtained through the softmax layer. The classification with the maximum probability is taken as the recommended standard for recommendation. A threshold is also set for the result probability. When the maximum probability is still lower than the threshold, the recommendation result is set as no suitable standard. It can be understood that TextCNN processes one-dimensional text data, and the convolution kernel only slides in one-dimensional space (vertically). Therefore, the convolution kernel has the same length as the word vector, and the feature extraction is completed from top to bottom.
[0101] In addition, the short text classification algorithm of the foregoing technical solution only describes the process of directly completing classification through data standards and various aspects of the standard field description. In fact, the classification process is also detachable. Data standards are a systematic set of content, which has classification topics, large categories, sub-categories, and small category labels from large to small, which can assist in judgment. The above technical solution can be divided into several classifications from large to small and performed sequentially. It includes constructing a global dictionary containing all data standard item names, aliases, definitions, and other text information, as well as corresponding classification labels (such as topics, large categories, sub-categories, and small categories). Through the trained classifier, the top-level classification area most relevant to the standard data is selected by predicting the high-level classification labels (such as topics or large categories) in the global dictionary. Further, in the selected top-level classification area, TextCNN or other short text classification algorithms are used again to predict the secondary classification (such as sub-categories or small categories), paying more attention to the semantic similarity between the standard data and the data standard items at a finer granularity, thereby further narrowing the range of candidate standard items. Finally, in the subdivided field, short text classification algorithms are applied again to select the highest probability item as the final recommended standard by comparing the matching probabilities of each candidate standard and the standard data.
[0102] The method for processing the data of the same standard provided in the embodiments of the present application comprises the following steps: obtaining data standard description information corresponding to a plurality of data standard items; determining a first data standard item as a target data standard item in the case that a field of the data of the same standard and a field in the data standard description information of the first data standard item satisfy a field matching condition, wherein the field matching condition is that a matching result between the fields is greater than a first threshold value, and the matching is directly and quickly performed based on the fields; and determining a second data standard item as the target data standard item in the case that the data of the same standard and the data standard description information of the second data standard item satisfy a similarity matching condition, wherein the similarity matching condition is determined according to a word segmentation hit rate between the data of the same standard and the data standard description information, and the synonym and the variant word can be effectively processed, and the success rate of the matching is improved. The purpose of performing a data standardization operation on the data of the same standard based on the data standard description information of the target data standard item is achieved, and the technical effect of automatically and accurately processing the data of the same standard is achieved, and the technical problem of low processing efficiency of the data of the same standard is solved.
[0103] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a group of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described herein can be executed in an order different from that shown herein.
[0104] Embodiment 2
[0105] The embodiments of the present application also provide a processing device for processing the data of the same standard. It should be noted that the processing device for processing the data of the same standard in the embodiments of the present application can be used to execute the method for processing the data of the same standard provided in the embodiments of the present application. The processing device for processing the data of the same standard provided in the embodiments of the present application is introduced as follows.
[0106] According to the embodiments of the present application, a device for implementing the above-mentioned method for processing the data of the same standard is also provided, as shown in Figure 4 The device comprises an obtaining unit 402 configured to obtain data standard description information corresponding to a plurality of data standard items;
[0107] A first determining unit 404 is configured to determine a first data standard item as a target data standard item in the case that a field of the data of the same standard and a field in the data standard description information of the first data standard item satisfy a field matching condition, wherein the field matching condition is that a matching result between the fields is greater than a first threshold value;
[0108] A second determining unit 406 is configured to determine a second data standard item as the target data standard item in the case that the data of the same standard and the data standard description information of the second data standard item satisfy a similarity matching condition, wherein the similarity matching condition is determined according to a word segmentation hit rate between the data of the same standard and the data standard description information;
[0109] The processing unit 408 performs a data standardization operation on the standard-penetrating data based on the data standard description information of the target data standard item.
[0110] The processing device for standard-penetrating data provided by the embodiments of the present application acquires data standard description information corresponding to a plurality of data standard items respectively, determines a first data standard item as a target data standard item in a case where a field of the standard-penetrating data and a field in the data standard description information of the first data standard item satisfy a field matching condition, wherein the field matching condition is that a matching result between the fields is greater than a first threshold value, and directly performs fast matching based on the fields; determines a second data standard item as the target data standard item in a case where the standard-penetrating data and the data standard description information of the second data standard item satisfy a similarity matching condition, wherein the similarity matching condition is determined according to a word segmentation hit rate between the standard-penetrating data and the data standard description information, which can effectively process synonyms and variant words and improve the success rate of matching. The purpose of performing a data standardization operation on the standard-penetrating data based on the data standard description information of the target data standard item is achieved, thereby realizing the technical effect of automatically and accurately processing the standard-penetrating data, and further solving the technical problem of low processing efficiency of the standard-penetrating data.
[0111] Optionally, in the processing device for standard-penetrating data provided by the embodiments of the present application, the first determining unit includes a matching module, which is configured to determine that the standard-penetrating data and the first data standard item satisfy the field matching condition and determine the first data standard item as the target data standard item in a case where a name field of the standard-penetrating data is the same as a standard name field indicated by the data standard description information of the first data standard item; and determine a key field in the name field of the standard-penetrating data, and determine the target data standard item in the at least one first data standard item according to the key field in a case where at least one reference field in the name field of the standard-penetrating data is the same as at least one standard name field indicated by the data standard description information of the first data standard item.
[0112] Optionally, the second determining unit includes a calculation module, which is configured to determine a standard-penetrating data word segmentation set matched with the standard-penetrating data and a plurality of data standard word segmentation sets respectively matched with a plurality of data standard items; and calculate a word segmentation hit rate between the standard-penetrating data word segmentation set and each data standard word segmentation set in the plurality of data standard word segmentation sets in sequence based on a weight value corresponding to each word segmentation through a similarity algorithm.
[0113] Optionally, the above-mentioned matching module is also used to determine that the field at the target position in the name field of the standardization data is a key field, and determine that the first data standard item corresponding to the standard name field that matches the key field is the target data standard item; determine that the field corresponding to the part of speech whose number of repetitions in the name field of the standardization data is greater than a third threshold is a key field, and determine that the first data standard item corresponding to the standard name field that matches the key field is the target data standard item.
[0114] Optionally, the second determining unit is further configured to determine a data standard item having a word segmentation hit rate greater than a second threshold as a second data standard item; and determine the second data standard item corresponding to the highest word segmentation hit rate as the target data standard item.
[0115] Optionally, the above-mentioned second determination unit is also used to determine the standardization data features that match the standardization data; the text prediction model outputs the prediction probability of the data standard item that matches the standardization data based on the standardization data features; and the data standard item whose prediction probability is greater than the fourth threshold is determined as the target data standard item.
[0116] Optionally, the above-mentioned first determination unit includes: a selection module for determining a target dictionary based on multiple data standard items; constructing a graph structure that matches the standardization data according to the target dictionary, wherein each node on the graph structure represents a word, and the edges between the nodes are used to indicate that there is a combination relationship between the words; determining the target path based on the graph structure, and splicing the words on the target path to obtain the fields of the standardization data.
[0117] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the contents disclosed in the above embodiment 1. It should be noted that the above modules or units can be hardware components or software components stored in a memory (e.g., memory 104) and processed by one or more processors (e.g., processors 102a, 102b, ..., 102n). The above modules can also be run as part of the device in the computer terminal 10 provided in the first embodiment.
[0118] Example 3
[0119] An embodiment of the present application may provide an electronic device, Figure 5 This is a structural block diagram of an electronic device according to an embodiment of the present application. Figure 5 As shown, the electronic device may include: one or more ( Figure 5 Only one is shown) processor 502, memory 504, storage controller, and peripheral interface, wherein the peripheral interface is connected to the radio frequency module, audio module and display.
[0120] The memory can be used to store software programs and modules, such as program instructions / modules corresponding to the method and device in the embodiments of the present application. The processor executes various functions and data processing by running the software programs and modules stored in the memory, that is, implements the above method. The memory can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory can further include a memory remotely arranged with respect to the processor, which can be connected to the terminal through a network. Examples of the above network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.
[0121] The processor can call information and application programs stored in the memory through the transmission device to perform the following steps: obtaining data standard description information corresponding to a plurality of data standard items respectively; in the case that the field of the data standard description information of the first data standard item and the field of the data standard description information of the first data standard item satisfy the field matching condition, determining that the first data standard item is the target data standard item, wherein the field matching condition is that the matching result between the fields is greater than a first threshold; in the case that the data standard description information of the second data standard item and the data standard description information of the second data standard item satisfy the similarity matching condition, determining that the second data standard item is the target data standard item, wherein the similarity matching condition is determined according to the word segmentation hit rate between the data standard description information and the data standard description information; performing a data standardization operation on the data based on the data standard description information of the target data standard item.
[0122] The processor can also call information and application programs stored in the memory through the transmission device to perform the following steps: in the case that the name field of the data standard description information of the first data standard item and the name field of the data standard description information of the first data standard item are the same, determining that the data standard description information of the first data standard item satisfies the field matching condition, and determining that the first data standard item is the target data standard item; in the case that at least one reference field in the name field of the data standard description information of the first data standard item is the same as at least one standard name field indicated by the data standard description information of the first data standard item, determining that the data standard description information of the first data standard item satisfies the field matching condition, determining a key field in the name field of the data standard description information, and determining the target data standard item in the at least one first data standard item according to the key field.
[0123] The processor can also call information and application programs stored in the memory through the transmission device to perform the following steps: determining a data standard description information set matched with the data standard description information and a plurality of data standard description information sets matched with a plurality of data standard items respectively; calculating the word segmentation hit rate between the data standard description information set and each data standard description information set in the plurality of data standard description information sets in sequence based on the weight value corresponding to each word segmentation through a similarity algorithm.
[0124] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: determining that a field of the target position in the name field of the through-standard data is a key field, and determining that a first data standard item corresponding to the standard name field matched with the key field is a target data standard item; determining that a field corresponding to a part of speech with a part of speech repetition number greater than a third threshold in the name field of the through-standard data is a key field, and determining that a first data standard item corresponding to the standard name field matched with the key field is a target data standard item.
[0125] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: determining that a data standard item with a tokenization hit rate greater than a second threshold is a second data standard item; and determining that the second data standard item corresponding to the highest tokenization hit rate is the target data standard item.
[0126] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: determining a through-standard data feature matched with the through-standard data; the text prediction model outputting a prediction probability of a data standard item matched with the through-standard data based on the through-standard data feature; and determining that the data standard item with the prediction probability greater than a fourth threshold is the target data standard item.
[0127] The processor can further call information and application programs stored in the memory through the transmission device to perform the following steps: determining a target dictionary based on a plurality of data standard items; constructing a graph structure matched with the through-standard data according to the target dictionary, wherein each node on the graph structure represents a vocabulary, and an edge between the nodes is used to indicate that the vocabularies have a combination relationship; determining a target path based on the graph structure, and splicing the vocabularies on the target path to obtain a field of the through-standard data.
[0128] By adopting the embodiments of the present application, data standard description information corresponding to a plurality of data standard items is obtained; in a case where a field of the through-standard data and a field in the data standard description information of the first data standard item satisfy a field matching condition, the first data standard item is determined as the target data standard item, wherein the field matching condition is that a matching result between the fields is greater than a first threshold, and the field is directly matched quickly; in a case where the through-standard data and the data standard description information of the second data standard item satisfy a similarity matching condition, the second data standard item is determined as the target data standard item, wherein the similarity matching condition is determined according to a tokenization hit rate between the through-standard data and the data standard description information, which can effectively process synonymous words and variant words and improve the success rate of matching; the purpose of performing a data standardization operation on the through-standard data based on the data standard description information of the target data standard item is achieved, thereby realizing the technical effect of automatically and accurately processing the through-standard data, and further solving the technical problem of low processing efficiency of the through-standard data.
[0129] Those skilled in the art can understand that Figure 5 The structure shown is only schematic, and the electronic device can also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a palm computer, a Mobile Internet Device (MID), a PAD, or the like. Figure 5 This does not limit the structure of the electronic device described above. For example, the electronic device can further include more or fewer components (such as a network interface, a display device, etc.) than those shown in the embodiment, or have a different configuration from that shown in the embodiment. Figure 5 Figure 5
[0130] Those skilled in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing the hardware related to the terminal device by a program, and the program can be stored in a computer readable storage medium, which can include a flash disk, a Read-Only Memory (ROM), a Random Access Memory (RAM), a magnetic disk or an optical disk, etc.
[0131] Embodiment 4
[0132] The embodiments of the present application also provide a storage medium. Optionally, in the embodiment, the storage medium can be used to save the program code executed by the processing method of the transfix data provided in the embodiment one.
[0133] Optionally, in the embodiment, the storage medium can be located in any one of the computer terminals in the computer terminal group in the computer network, or in any one of the mobile terminals in the mobile terminal group.
[0134] The present application also provides a computer program product adapted to execute the steps of the processing method of the transfix data when executed on a data processing device.
[0135] The above-mentioned serial numbers of the embodiments of the present application are only for description, and do not represent the advantages or disadvantages of the embodiments.
[0136] In the above embodiments of the present application, the description of each embodiment has its own focus, and the parts not described in detail in a certain embodiment can be referred to the related description of other embodiments.
[0137] In several embodiments provided in the present application, it should be understood that the disclosed technology can be implemented by other ways. Among them, the above-described device embodiments are only schematic, for example, the division of the units is only a logical function division, and actual implementation can have another division manner, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed units can be indirect coupling or communication connection through some interfaces, units or modules, and can be electrical or other forms.
[0138] The units described as separate components can or can not be physically separate, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on a plurality of network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0139] In addition, the functional units in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.
[0140] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on this understanding, the technical solutions of the present application essentially or the part of the prior art that contributes to the technical solutions or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The foregoing storage medium includes: a U disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store program codes.
[0141] The above is only the preferred embodiment of the present application, and it should be pointed out that for ordinary skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, and these improvements and refinements should be considered as the protection scope of the present application.
Claims
1. A method for processing standardization data, characterized in that: include: Obtaining data standard description information corresponding to multiple data standard items; If a field of the standard-compliant data and a field in the data standard description information of the first data standard item meet a field matching condition, determining the first data standard item as a target data standard item, wherein the field matching condition is that a matching result between the fields is greater than a first threshold; If the standard-compliant data and the data standard description information of the second data standard item meet a similarity matching condition, determining the second data standard item as the target data standard item, wherein the similarity matching condition is determined based on a word segmentation hit rate between the standard-compliant data and the data standard description information; Based on the data standard description information of the target data standard item, a data standardization operation is performed on the standard-compliant data.
2. The method according to claim 1, characterized in that The determining that the first data standard item is a target data standard item when a field of the standard-compliant data and a field in the data standard description information of the first data standard item meet a field matching condition further includes: If the name field of the standard-compliant data is identical to the standard name field indicated by the data standard description information of the first data standard item, determining that the field matching condition is satisfied between the standard-compliant data and the first data standard item, and determining that the first data standard item is the target data standard item; When there is at least one reference field in the name field of the standard-compliant data that is identical to the standard name field indicated by the data standard description information of at least one of the first data standard items, it is determined that the field matching condition is satisfied between the standard-compliant data and the first data standard item, the key field in the name field of the standard-compliant data is determined, and the target data standard item is determined in at least one of the first data standard items based on the key field.
3. The method according to claim 1, characterized in that When the data standard description information of the standard-compliant data and the second data standard item meets a similarity matching condition, before determining the second data standard item as the target data standard item, the method further includes: Determining a set of standardized data segmentation words that matches the standardized data and a plurality of data standard segmentation sets that respectively match a plurality of the data standard items; The segmentation hit rate between the standard data segmentation set and each of the plurality of data standard segmentation sets is calculated in sequence based on the weight value corresponding to each segmentation set by a similarity algorithm.
4. The method according to claim 2, characterized in that The determining of the key field in the name field of the standardization data, and determining the target data standard item in at least one of the first data standard items according to the key field, includes at least one of the following: Determine the field at the target position in the name field of the standardization data as the key field, and determine the first data standard item corresponding to the standard name field matching the key field as the target data standard item; Determine that the field corresponding to the part of speech in the name field of the standardized data whose part of speech is repeated more than a third threshold is the key field, and determine that the first data standard item corresponding to the standard name field that matches the key field is the target data standard item.
5. The method according to claim 3, characterized in that The determining the second data standard item as the target data standard item includes: The data standard item whose word segmentation hit rate is greater than a second threshold is determined as the second data standard item; and the second data standard item corresponding to the highest word segmentation hit rate is determined as the target data standard item.
6. The method according to claim 1, characterized in that When the data standard description information of the standard-compliant data and the second data standard item meets a similarity matching condition, after determining that the second data standard item is the target data standard item, the method further includes: Determining the standardization data features that match the standardization data; The text prediction model outputs a prediction probability of the data standard item matching the standard data based on the standard data feature; The data standard item whose predicted probability is greater than a fourth threshold is determined as the target data standard item.
7. The method according to claim 1, characterized in that Before determining the first data standard item as the target data standard item, the method further includes: determining a target dictionary based on a plurality of said data standard items; Constructing a graph structure that matches the standardization data according to the target dictionary, wherein each node on the graph structure represents a word, and the edges between the nodes are used to indicate that there is a combination relationship between the words; A target path is determined based on the graph structure, and the words on the target path are concatenated to obtain the fields of the standardization data.
8. A device for processing standardization data, characterized in that: include: An acquisition unit, which acquires data standard description information corresponding to a plurality of data standard items; a first determining unit, configured to determine that the first data standard item is a target data standard item if a field of the standard-compliant data and a field in the data standard description information of the first data standard item satisfy a field matching condition, wherein the field matching condition is that a matching result between the fields is greater than a first threshold; a second determining unit, configured to determine that the second data standard item is the target data standard item if the standard-compliant data and the data standard description information of the second data standard item satisfy a similarity matching condition, wherein the similarity matching condition is determined based on a word segmentation hit rate between the standard-compliant data and the data standard description information; A processing unit performs a data standardization operation on the standard-compliant data based on the data standard description information of the target data standard item.
9. A computer-readable storage medium, characterized in that The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the computer-readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
10. An electronic device, characterized in that: include: a memory storing an executable program; A processor, configured to run the program, wherein the program executes the method according to any one of claims 1 to 7 when running.
11. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.