A data element construction method and device, electronic equipment and storage medium
Patent Information
- Application Number
- CN202211396821.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-09
- Publication Date
- 2026-09-15
- Estimated Expiration
- 2042-11-09
AI Technical Summary
[0005]然而,采用数据元素的自动对标方法,通常都是预先构建数据元素,或者,在既有数据元素的基础上构建新的数据元素,并且,由于传统的数据元素构建方法效率较低,故而,极大地影响了数据元素的构建效率
[0038] In the data element construction method provided in this application embodiment, each data item involved in the target business scenario and its respective data item name are obtained. From the keyword unit set of each data item name, a subset of keyword units that meet the preset keyword unit conditions for each data item name are selected. Then, based on the keyword units contained in each of the obtained keyword unit subsets, data elements corresponding to each data item are generated. When it is determined that each obtained data element meets the preset data element matching conditions, each data element is saved to the preset target data element set.
Smart Images

Figure CN115827927B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a method, apparatus, electronic device and storage medium for constructing data elements. Background Technology
[0002] With the popularization and development of Internet technology, massive amounts of data are constantly emerging from daily life. Furthermore, big data technology and artificial intelligence technology, based on distributed data storage and computing, provide the basic conditions and application scenarios for the use of massive amounts of data.
[0003] Therefore, in order to extract target data that meets the corresponding business needs from massive amounts of data, it is necessary to manage the massive amounts of data. Among them, the automatic data element benchmarking technology plays a good role in improving the quality of data management and reducing the cost of data management.
[0004] For example, in a data quality assessment scenario, in order to assess the data quality of the data to be assessed, a subset of the data to be assessed is usually obtained first. Then, the subset of the data to be assessed is matched with a preset set of data elements to determine the target data elements corresponding to the field information in the subset of the data to be assessed. Further, based on the correspondence between the preset data elements and the verification rules, the target verification rules corresponding to the target data elements are determined. Then, based on the target verification rules, the subset data pass rate of the subset of the data to be assessed is determined.
[0005] However, the automatic data element benchmarking method usually involves pre-constructing data elements or building new data elements based on existing data elements. Furthermore, since the traditional data element construction method is inefficient, it greatly affects the construction efficiency of data elements.
[0006] Therefore, the construction efficiency of data elements is relatively low in the existing technology. Summary of the Invention
[0007] This application provides a data element construction method, apparatus, electronic device, and storage medium to improve the efficiency of data element construction.
[0008] In a first aspect, embodiments of this application provide a method for constructing data elements, the method comprising:
[0009] Obtain each data item involved in the target business scenario and its respective data item name, and filter out the keyword unit subsets of each data item name that meet the preset keyword unit conditions from the keyword unit set of each data item name.
[0010] Based on the keyword units contained in each of the obtained subsets of keyword units, data elements corresponding to each data item are generated; wherein, each data element represents: at least one keyword unit of the corresponding data item;
[0011] When it is determined that each obtained data element meets the preset data element matching conditions, each data element is saved to the preset target data element set.
[0012] Secondly, embodiments of this application also provide a data element construction apparatus, the apparatus comprising:
[0013] The acquisition module is used to acquire each data item involved in the target business scenario and its respective data item name, and to filter out the keyword unit subset of each data item name that meets the preset keyword unit conditions from the keyword unit set of each data item name.
[0014] The generation module is used to generate data elements corresponding to each data item based on the keyword units contained in each subset of obtained keyword units; wherein each data element represents: at least one keyword unit of the corresponding data item;
[0015] The filtering module is used to save each data element to a preset target data element set when it is determined that each obtained data element meets the preset data element matching conditions.
[0016] In one possible embodiment, when selecting a subset of keyword units for each data item name that satisfy preset keyword unit conditions from the keyword unit sets for each data item name, the acquisition module is specifically used for:
[0017] Perform the following operations for each data item name:
[0018] Divide a data item name into a set of keyword units for that data item name, and determine the part-of-speech of each keyword unit contained in the set of keyword units.
[0019] From each keyword unit, at least one keyword unit with the part of speech of noun is selected, and the set of nouns consisting of at least one keyword unit is taken as the subset of keyword units that meet the keyword unit conditions.
[0020] In one possible embodiment, when generating data elements corresponding to each data item based on the keyword units contained in each of the obtained subsets of keyword units, the generation module is specifically used for:
[0021] Based on the keyword units contained in each subset of keyword units and the number of keyword units, the keyword unit combination corresponding to each data item is obtained;
[0022] When matching data items in the matching test set with a set number of test pairings, obtain the matching success probability of each keyword unit combination, and select at least one keyword unit combination from each keyword unit combination according to the preset matching success probability selection conditions.
[0023] Based on at least one keyword unit combination, generate the corresponding data elements for each data item.
[0024] In one possible embodiment, when obtaining the keyword unit combination corresponding to each data item based on the keyword units contained in each subset of keyword units and the number of keyword units therein, the generation module is specifically used for:
[0025] For each subset of keyword units, perform the following operations:
[0026] Determine the number of keyword units contained in a subset of keyword units;
[0027] If the number of keyword units is not greater than the preset threshold for the number of units, then the sequence of keyword units contained in a subset of keyword units will be used as the keyword unit combination of the corresponding data item.
[0028] If the number of keyword units exceeds the threshold, then select the keyword unit sequence that meets the threshold from the keyword units contained in a subset of keyword units, and use the keyword unit sequence as the keyword unit combination of the corresponding data item.
[0029] In one possible embodiment, when generating data elements corresponding to each data item based on at least one keyword unit combination, the generation module is specifically used for:
[0030] If at least one keyword unit combination satisfies the preset data element composition conditions, then the keyword unit combination can be used as the data element of the corresponding data item.
[0031] In one possible embodiment, when it is determined that each of the obtained data elements satisfies a preset data element matching condition, the filtering module is specifically used for:
[0032] When matching data items with a set target number of matching times in the target business scenario, obtain the matching success probability of each data element, and obtain the corresponding total matching success probability based on the obtained matching success probabilities.
[0033] When the total probability of successful pairing is greater than the preset probability threshold for successful pairing, each data element is determined to meet the data element matching condition.
[0034] Thirdly, embodiments of this application also propose an electronic device including a processor and a memory, wherein the memory stores program code that, when executed by the processor, causes the processor to perform the steps of the data element construction method described in the first aspect.
[0035] Fourthly, embodiments of this application also propose a computer-readable storage medium comprising program code, which, when executed on an electronic device, causes the electronic device to perform the steps of the data element construction method described in the first aspect.
[0036] Fifthly, embodiments of this application also provide a computer program product, which, when invoked by a computer, causes the computer to execute the data element construction method steps as described in the first aspect.
[0037] The beneficial effects of this application are as follows:
[0038] In the data element construction method provided in this application embodiment, each data item involved in the target business scenario and its respective data item name are obtained. From the keyword unit set of each data item name, a subset of keyword units that meet the preset keyword unit conditions for each data item name are selected. Then, based on the keyword units contained in each of the obtained keyword unit subsets, data elements corresponding to each data item are generated. When it is determined that each obtained data element meets the preset data element matching conditions, each data element is saved to the preset target data element set.
[0039] This approach generates data elements corresponding to each data item based on the keyword units contained in each subset of obtained keyword units, thus achieving automatic construction of data elements. Furthermore, when it is determined that each obtained data element meets the preset data element matching conditions, each data element is saved to the preset target data element set. This ensures the probability of successful pairing of the constructed data elements with data items in the target business scenario, improving the accuracy of data element construction. As such, it avoids the technical drawbacks of existing technologies, such as the need to pre-construct data elements or construct new data elements based on existing data elements, and the low efficiency of traditional data element construction methods. Therefore, it improves the efficiency of data element construction.
[0040] Furthermore, other features and advantages of this application will be set forth in the following description and will be apparent in part from the description, or may be learned by practicing the application. The objectives and other advantages of this application may be realized and obtained by means of the structures particularly pointed out in the written description, claims, and drawings. Attached Figure Description
[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:
[0042] Figure 1 This is an optional schematic diagram of the system architecture applicable to the embodiments of this application;
[0043] Figure 2 A schematic diagram illustrating the implementation process of a data element construction method provided in this application embodiment;
[0044] Figure 3 This application provides an illustration of an application scenario for obtaining a subset of keyword units.
[0045] Figure 4 This is a schematic flowchart of a method for generating data elements provided in an embodiment of this application;
[0046] Figure 5 A method based on the embodiments of this application is provided. Figure 2 A logical diagram;
[0047] Figure 6 This application provides a schematic diagram of a specific application process for constructing data elements.
[0048] Figure 7 This is a schematic diagram of the structure of a data element construction device provided in an embodiment of this application;
[0049] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0050] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of this application will be clearly and completely described below with reference to the accompanying drawings of the embodiments of this application. Obviously, the described embodiments are only some embodiments of the technical solutions of this application, and not all embodiments. Based on the embodiments recorded in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the technical solutions of this application.
[0051] It should be noted that in the description of this application, "multiple" is understood as "at least two". "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. A connected to B can represent: A and B directly connected, or A and B connected through C. Furthermore, in the description of this application, terms such as "first" and "second" are used only for descriptive purposes and should not be construed as indicating or implying relative importance or order.
[0052] Furthermore, the data collection, dissemination, and use in the technical solution of this application all comply with the requirements of relevant national laws and regulations.
[0053] To facilitate understanding by those skilled in the art, some of the nouns and terms involved in the embodiments of this application will be briefly described and explained as follows:
[0054] (1) Data Element (DE): also known as a data unit, is a data unit whose definition, identification, representation and allowed values are described by a set of attributes. In a certain context, it is usually used to construct a semantically correct, independent and unambiguous information unit of a specific concept. Data element can be understood as the basic unit of data. In particular, a data model is a whole structure composed of several related data elements in a certain order.
[0055] (2) Forward Maximum Matching (FMM) algorithm: Forward means taking words from left to right. The maximum length of the word taken is the length of the long words in the dictionary. Each time, one word is subtracted from the right until there is a single word in the dictionary or only one word remains.
[0056] (3) Backward Maximum Matching (BMM) algorithm: also known as the reverse maximum matching algorithm, its basic principle is similar to the FMM algorithm, except that the word segmentation order is changed to right to left. Generally, a maximum word length segment is selected from the beginning of a string. If the sequence is less than the maximum word length, the entire sequence is selected. Each time, one word is subtracted from the left until there is a single word in the dictionary or only one word remains.
[0057] (4) N-shortest path method: This is an important algorithm used for word segmentation. Its basic idea is to give a string to be processed, find all possible words in the dictionary, construct a directed acyclic graph of the string, and calculate the shortest N paths from the beginning to the end. Since paths of equal length are allowed to be parallel, the final result set will be greater than or equal to N.
[0058] (5) Conditional Random Field (CRF): It is a discriminative probability model that is often used to label or analyze sequence data.
[0059] (6) Hidden Markov Model (HMM): A statistical model used to describe a Markov process with hidden unknown parameters. The challenge lies in determining the hidden parameters from the observable parameters and then using these parameters for further analysis, such as pattern recognition. It should be noted that in this paper, it can be used for word segmentation and part-of-speech tagging.
[0060] (7) Long-tail keywords: refers to combination keywords that are not target keywords but are related to target keywords and can bring search traffic. Among them, "long tail" has two characteristics: fine and long. Fine means that the proportion of long tail is small, and long means that although the proportion is small, the number is large.
[0061] It should be noted that the above terminology naming method is only an example, and the embodiments of this application do not limit the naming method of the above terms.
[0062] Furthermore, based on the above explanations of terms and related terminology, the design concept of the embodiments of this application will be briefly introduced below:
[0063] With the popularization and development of Internet technology, massive amounts of data are constantly emerging from daily life. Furthermore, big data technology and artificial intelligence technology, based on distributed data storage and computing, provide the basic conditions and application scenarios for the use of massive amounts of data.
[0064] Therefore, in order to extract target data that meets the corresponding business needs from massive amounts of data, data management is required. However, when carrying out data governance or management work in some industries that do not yet have industry-standard data elements, how to quickly, efficiently, and specifically create data elements suitable for business scenarios, and how to efficiently and automatically construct raw data, are urgent problems to be solved.
[0065] In view of this, in order to improve the construction efficiency of data elements and solve the problem of constructing new data elements based on existing data elements, a data element construction method is proposed in this embodiment of the application. The method includes: obtaining each data item involved in the target business scenario and its respective data item name; and selecting a subset of keyword units that meet the preset keyword unit conditions from the keyword unit set of each data item name; then, generating data elements corresponding to each data item based on the keyword units contained in each of the obtained keyword unit subsets, wherein each data element represents at least one keyword unit of the corresponding data item; and finally, when it is determined that each obtained data element meets the preset data element matching conditions, saving each data element to the preset target data element set.
[0066] In particular, the preferred embodiments of this application will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit this application. Furthermore, the embodiments of this application and the features in the embodiments can be combined with each other without conflict.
[0067] See Figure 1 The diagram illustrates a system architecture applicable to an embodiment of this application. This system architecture includes a target terminal (101a, 101b) and a server 102. The target terminal (101a, 101b) and the server 102 can interact via a communication network. The communication network can employ wireless communication or wired communication methods.
[0068] For example, the target terminal (101a, 101b) can access the network and communicate with the server 102 through cellular mobile communication technology, wherein the cellular mobile communication technology includes, for example, 5th generation mobile network (5G) technology.
[0069] Optionally, the target terminal (101a, 101b) can access the network and communicate with the server 102 via short-range wireless communication, wherein the short-range wireless communication method includes, for example, Wireless Fidelity (Wi-Fi) technology.
[0070] This application embodiment does not impose any limitation on the number of communication devices involved in the above system architecture, such as Figure 1 As shown, only the target terminal (101a, 101b) and server 102 are described as examples. The following is a brief introduction to each of the above devices and their respective functions.
[0071] The target terminal (101a, 101b) is a device that can provide voice and / or data connectivity to a user, and can be a device that supports wired and / or wireless connection methods.
[0072] For example, the target terminals (101a, 101b) include, but are not limited to: mobile phones, tablets, laptops, handheld computers, mobile internet devices (MID), wearable devices, virtual reality (VR) devices, augmented reality (AR) devices, wireless terminal devices in industrial control, wireless terminal devices in autonomous driving, wireless terminal devices in smart grids, wireless terminal devices in transportation safety, wireless terminal devices in smart cities, or wireless terminal devices in smart homes, etc.
[0073] Furthermore, the target terminals (101a, 101b) may have related clients installed. These clients can be software, such as applications (APPs), browsers, short video apps, web pages, or mini-programs. In this embodiment, the target terminals (101a, 101b) can be used to send the various data items involved in the target business scenario and their respective data item names to the server 102.
[0074] Server 102 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and big data and artificial intelligence platforms.
[0075] It is worth mentioning that, in this embodiment of the application, the server 102 is used to obtain each data item involved in the target business scenario and its respective data item name, and to filter out the keyword unit subsets of each data item name that meet the preset keyword unit conditions from the keyword unit set of each data item name, and then generate the data elements corresponding to each data item based on the keyword units contained in each obtained keyword unit subset, so that when it is determined that each obtained data element meets the preset data element matching conditions, each data element is saved to the preset target data element set.
[0076] The data element construction method provided by the exemplary embodiments of this application will be described below in conjunction with the above system architecture and with reference to the accompanying drawings. It should be noted that the above system architecture is only shown for the purpose of understanding the spirit and principles of this application, and the embodiments of this application are not limited in any way.
[0077] See Figure 2 The diagram shown is an implementation flowchart of a data element construction method provided in this application embodiment. Taking a server as an example, the specific implementation flow of this method is as follows:
[0078] S201: Obtain each data item involved in the target business scenario and its respective data item name, and filter out the keyword unit subsets of each data item name that meet the preset keyword unit conditions from the keyword unit set of each data item name.
[0079] It should be noted that each keyword unit set can contain one or more keyword units of one part of speech. For example, if the data item name of a certain data item is "device number", then the corresponding keyword unit set is {"device", "number"}. The preset keyword unit condition characterizes whether the part of speech of the keyword units that can be classified into a subset of keyword units meets the preset part of speech requirements. The part of speech requirements can be any one of various parts of speech such as verb, noun, adjective, etc.
[0080] For example, if the preset keyword unit condition is that the part of speech of the keyword unit in the keyword unit subset is noun, then the two keyword units in the above keyword unit set {"device", "number"} both satisfy the preset keyword unit condition. Therefore, the keyword unit set {"device", "number"} can be used as the keyword unit subset with the data item name "device number", that is, {"device", "number"}.
[0081] Furthermore, in this embodiment of the application, no restrictions are placed on the language type of the data item name; it can be Chinese, English, or other types of languages.
[0082] S202: Based on the keyword units contained in each of the obtained keyword unit subsets, generate the data elements corresponding to each data item.
[0083] Each data element represents at least one keyword unit of the corresponding data item.
[0084] Optionally, before the server generates the corresponding data elements for each data item based on the keyword units contained in each keyword unit subset, if it detects that there are duplicate keyword unit subsets in the obtained keyword unit subsets, the keyword unit subsets can be filtered and deduplicated, thereby reducing the demand on system resources (e.g., memory).
[0085] For example, taking data item 1 named "Equipment Source Number" and data item 2 named "Equipment Placement Number" as examples, if the preset keyword unit condition is that the part of speech of the keyword unit in the keyword unit subset is noun, then the corresponding keyword unit subsets are as follows: Keyword unit subset 1 {Equipment, Number} and Keyword unit subset 2 {Equipment, Number}. Obviously, Keyword unit subset 1 {Equipment, Number} and Keyword unit subset 2 {Equipment, Number}, therefore, only the keyword unit subset, i.e. {Equipment, Number}, needs to be saved once, and data item 1 and data item 2 and keyword unit subset {Equipment, Number} need to be recorded.
[0086] S203: When it is determined that each obtained data element meets the preset data element matching conditions, each data element is saved to the preset target data element set.
[0087] It should be noted that the preset data element matching conditions represent that each obtained data element meets the corresponding business requirements, such as achieving a certain probability of successful pairing of data items.
[0088] In a preferred embodiment, when performing step S201, in the scenario where the server creates corresponding data elements for the target business scenario, it can obtain each data item involved in the target business scenario and its respective data item name from the data item information set received from the target terminal. Then, based on the part-of-speech of each keyword unit in the keyword unit set of each data item name, it can filter out at least one keyword unit that meets the preset part-of-speech selection conditions from the keyword units contained in the keyword unit set of each data item name, thereby obtaining a subset of keyword units for each data item name that meets the preset keyword unit conditions.
[0089] Furthermore, when the server selects a subset of keyword units for each data item name that meets the preset keyword unit conditions from the keyword unit sets for each data item name, refer to... Figure 3As shown, for each data item name, the following operations are performed: A data item name, Data.Item.Name, is partitioned to obtain a set of keyword units, Keyword.Unit.Set, for Data.Item.Name. The part-of-speech (NOS) of each keyword unit (e.g., Keyword.Unit1, Keyword.Unit2, and Keyword.Unit3) contained in Keyword.Unit.Set is determined (in order: noun n, verb v, and noun n). Then, from each keyword unit (i.e., Keyword.Unit1, Keyword.Unit2, and Keyword.Unit3), at least one keyword unit (i.e., Keyword.Unit1 and Keyword.Unit3) with the part-of-speech of noun n is selected. The noun set Noun.Set formed by at least one keyword unit (i.e., Keyword.Unit1 and Keyword.Unit3) is taken as the subset of keyword units Keyword.Cell.Subset that satisfies the keyword unit condition Keyword.Unit.Condition.
[0090] In a preferred embodiment, see [reference] Figure 4 As shown, during step S202, after the server obtains each subset of keyword units, it can generate the corresponding data elements for each data item based on the keyword units contained in each subset. The specific steps are as follows:
[0091] S2021: Based on the keyword units contained in each keyword unit subset and the number of keyword units, obtain the keyword unit combination corresponding to each data item.
[0092] Optionally, when performing step S2021, the server performs the following operations for each keyword unit subset: determining the number of keyword units contained in a keyword unit subset; if the number of keyword units is not greater than a preset unit number threshold, then the keyword unit sequence contained in a keyword unit subset is used as the keyword unit combination of the corresponding data item; if the number of keyword units is greater than the unit number threshold, then the keyword unit sequence that meets the unit number threshold is selected from the keyword units contained in a keyword unit subset, and the keyword unit sequence is used as the keyword unit combination of the corresponding data item.
[0093] For example, assuming the preset unit number threshold is 3, if the number of keyword units in a subset of keyword units is 2, it is easy to see that the number of keyword units is 2, which is less than the preset unit number threshold of 3. Therefore, the sequence of keyword units in a subset of keyword units can be directly used as the keyword unit combination of the corresponding data item. If the number of keyword units in a subset of keyword units is 4, it is easy to see that the number of keyword units is 4, which is greater than the preset unit number threshold of 3. Therefore, from the keyword units in a subset of keyword units, the sequence of keyword units that meets the preset unit number threshold is selected, and the sequence of keyword units is used as the keyword unit combination of the corresponding data item.
[0094] It should be noted that the relative order of each keyword unit in the keyword unit sequence remains unchanged. Therefore, if the number of keyword units in a keyword unit subset is 4 and the preset unit number threshold is 3, the above method can be used to select keyword unit sequences that meet the preset unit number threshold from the keyword units in a keyword unit subset. There are 3 possible keyword unit sequences, and furthermore, there are also 3 possible combinations of keyword units for the corresponding data items.
[0095] S2022: When matching data items in the matching test set with a set number of test pairings, obtain the matching success probability of each keyword unit combination, and select at least one keyword unit combination from each keyword unit combination according to the preset matching success probability selection conditions.
[0096] Specifically, during step S2022, after obtaining the keyword unit combination corresponding to each data item, the server can perform data item pairing with a set number of test pairings in the preset matching test set, obtain the number of successful pairings for each keyword unit combination, and then obtain the success probability of each keyword unit combination. Finally, according to the preset selection criteria for success probability, at least one keyword unit combination is selected from each keyword unit combination, thereby ensuring that a small number of keyword unit combinations can occupy a large success probability of pairing. This avoids the "long tail" characteristic of the success probability of keyword unit combinations to a certain extent, thereby reducing the demand on system resources.
[0097] Optionally, the selection criteria for the above-mentioned matching success probability can be as follows: arrange the matching success probabilities of each keyword unit combination in descending order, and then select the keyword unit combination with the highest matching success probability as the first keyword unit combination to be selected, and then select keyword unit combinations with relatively high matching success probabilities in turn, until the sum of the matching success probabilities of at least one selected keyword unit combination is greater than the preset cumulative success probability threshold.
[0098] For example, assuming a preset cumulative success probability threshold of 85%, there are 8 keyword unit combinations corresponding to each data item. After performing a set number of matching tests (e.g., 10,000 times) on the data items in the preset matching test set, the number of successful matches and the success probability of the above 8 keyword unit combinations are shown in Table 1:
[0099] Table 1
[0100] Keyword.unit.Com1 3485 34.85% Keyword.unit.Com2 2641 26.41% Keyword.unit.Com3 1780 17.80% Keyword.unit.Com4 1258 12.58% Keyword.unit.Com5 421 4.21% Keyword.unit.Com6 255 2.55% Keyword.unit.Com7 94 0.94% Keyword.unit.Com8 66 0.66%
[0101] Based on the table above, the number of successful pairings and the probability of successful pairing for each keyword unit combination are arranged in descending order of number and probability. Among the above 8 keyword unit combinations, the cumulative success rate of the first 4 keyword unit combinations is 91.64%, which is greater than the preset cumulative success rate threshold of 85%, and the cumulative success rate of the first 3 keyword unit combinations is 79.06%. Therefore, at least one keyword unit combination is selected from the above 8 keyword unit combinations as follows: Keyword.unit.Com1, Keyword.unit.Com2, Keyword.unit.Com3, and Keyword.unit.Com4.
[0102] It is worth noting that, based on the method and steps described in S2022, a high success rate of matching can be achieved by combining a small number of extracted keyword units. This means that a large number of keyword unit combinations with low matching frequency at the tail end of the cumulative distribution that follows the "long tail characteristic" are temporarily set aside, thus greatly satisfying business needs.
[0103] S2023: Based on at least one keyword unit combination, generate the corresponding data elements for each data item.
[0104] It should be noted that a data element generally consists of three parts: object class, property, and representation. The object class is a collection of things in the real world or in an abstract concept. It has clear boundaries and meanings, and its properties and behaviors follow the same rules and can be identified. The property is a certain characteristic shared by all individuals of the object class, which is the basis for distinguishing the object from other members. The representation is a combination of value range, data type, and representation method. When necessary, it also includes information such as unit of measurement and character set.
[0105] Given the composition structure of data elements, keyword unit combinations can be used as a reference for constructing data elements. If at least one of the keyword unit combinations meets the preset data element composition conditions, then the keyword unit combination can be used as the data element of the corresponding data item.
[0106] The preset data element composition conditions include, but are not limited to, one of the following rule patterns:
[0107] 1. The pattern for combining corresponding keyword units is "object class + characteristic word + representation word", such as "citizen identity number";
[0108] 2. The pattern of corresponding keyword unit combination is as follows: when the object class word is clear, it conforms to "characteristic word + descriptive word", such as "gender + name", omitting the clear object class word "person".
[0109] Furthermore, based on the data elements obtained from the above steps, and in conjunction with industry standards and the opinions of business experts, it can be finally confirmed whether they should be used as standard data elements in the future.
[0110] In a preferred embodiment, when performing step S203, the server needs to determine whether each obtained data element meets the preset data element matching conditions during the process of saving each obtained data element to the preset target data element set. Only when it is determined that each obtained data element meets the preset data element matching conditions will the obtained data element be saved to the preset target data element set.
[0111] For example, when the server obtains the pairing success probability of each data element in the target business scenario for the target number of pairings, it obtains the corresponding total pairing success probability based on the obtained pairing success probability. Thus, when the total pairing success probability is greater than the preset pairing success probability threshold, it can be determined that each data element meets the above-mentioned preset data element matching conditions.
[0112] It should be noted that the preset matching success probability threshold can be set based on experience, and is usually not less than the cumulative success probability threshold in step S2022 above.
[0113] Obviously, based on the above step S203, the constructed data elements have been verified, that is, to ensure that the constructed data elements cover the business data involved in the target business scenario, that is, each data item.
[0114] Furthermore, based on the above data element construction method steps, refer to... Figure 5 As shown, this is a logical schematic diagram of a data element construction method provided in an embodiment of this application. It obtains various data items (e.g., Data.Item1, Data.Item2, and Data.Item3) involved in the target business scenario and their respective data item names (in order: Data.Item.Na1, Data.Item.Na2, and Data.Item.Na3). Then, it filters out the respective keyword unit sets (in order: Key.Unit.Set1, Key.Unit.Set2, and Key.Unit.Set3) of each data item name, selecting those that satisfy the preset keyword unit conditions Keywo. The keyword unit subsets of rd.Unit.Condition (namely Key.Cell.Subset1, Key.Cell.Subset2, and Key.Cell.Subset3, respectively) are then processed. Next, based on the keyword units contained in each of the obtained keyword unit subsets, data elements corresponding to each data item are generated (e.g., Data.Ele1, Data.Ele2, and Data.Ele3). Finally, when it is determined that each obtained data element satisfies the preset data element matching condition Da.Ele.Mat.Cri, each data element is saved to the preset target data element set Tar.Data.Ele.Col.
[0115] For example, see Figure 6 As shown, this is a schematic diagram of a specific application process for constructing data elements according to an embodiment of this application, wherein the word root is a keyword unit, and the specific method flow is as follows:
[0116] S601: Extract data items.
[0117] It should be noted that the premise of constructing data elements is that data management must be carried out for an industry that currently lacks industry data elements or is missing relevant data elements. Therefore, under this premise, it is certain that the data management client can access all the business system data that needs to be managed.
[0118] Furthermore, in order to extract the corresponding data items, the physical model data table structure information of these business system data can be extracted to form a two-dimensional data table. This data table includes, but is not limited to: data item auto-incrementing ID, data table Chinese name, data table English name, data item Chinese name, data item English name, data type, and the business system to which it belongs, as shown in Table 2 below:
[0119] Table 2
[0120]
[0121] S602: Extract word roots.
[0122] Specifically, during step S601, the server primarily uses natural language processing techniques to extract noun roots from the data item names one by one, facilitating the next step of pairing analysis. This includes the following steps:
[0123] S6021: Obtain the content segmentation corpus.
[0124] Among them, the content segmentation corpus is the set of keyword units.
[0125] For example, the server can extract the Chinese name column of the data item from the physical model data table structure information output in step S601, and perform word segmentation and part-of-speech tagging through natural language processing related technologies to construct a content segmentation corpus corresponding to each data item.
[0126] The Chinese word segmentation methods used in this application include, but are not limited to: vocabulary-based FMM, BMM, and N-shortest path methods; word segmentation methods based on N-gram language models using statistical analysis models; and sequence labeling-based HMM, CRF, and word perceptron.
[0127] It should be noted that the Python third-party toolkit jieba, by integrating the HMM algorithm, allows users to easily write code to perform word segmentation and part-of-speech tagging on Chinese words, which can greatly improve the efficiency of word segmentation.
[0128] For example, using the above techniques, a content segmentation corpus was constructed as shown in Table 3, and parts of speech were labeled.
[0129] Table 3
[0130]
[0131] S6022: Filter noun roots.
[0132] Specifically, when executing step S6022, after the server obtains the content segmentation corpus, it can further filter out word roots of specific part-of-speech rules from the content segmentation corpus. Since the constructed data elements consist of three parts: object class words, characteristic words, and representation words, and these three parts are mostly nouns, the word root filtering rule can be: only retain noun word roots and remove other part-of-speech word roots.
[0133] For example, if we still take the content segmentation corpus 3 after tagging parts of speech as an example, the corresponding noun roots can be obtained, as shown in Table 4:
[0134] Table 4
[0135]
[0136]
[0137] It should be noted that the above-mentioned noun root list is the set of nouns mentioned in the embodiments of this application and the subset of keyword units that meet the keyword unit conditions.
[0138] S6023: Word root combination.
[0139] Specifically, during step S6023, after the server filters out the noun roots, it can combine the noun root table for each data item one by one. Assuming that the noun root table for a certain data item has N roots (i.e., the number of keyword units), the root combination rules are as follows:
[0140] 1. If N≤3, this data item returns a word root pairing result, where the relative order of the noun word roots contained in the combination remains unchanged.
[0141] 2. If N>3, this data item returns a permutation of 3 word roots taken from N different word roots. Each permutation (keyword unit sequence) is a word root combination (i.e., keyword unit combination), in which the relative order of the noun word roots contained in the permutation remains unchanged.
[0142] It should be noted that in this embodiment, 3 is a preset threshold for the number of units. Therefore, based on the above rules, the server can obtain the root word combination results corresponding to Table 4, as shown in Table 5:
[0143] Table 5
[0144]
[0145] S603: Statistics on root word pairings.
[0146] Specifically, when executing step S603, after the server extracts the word roots, it can group the obtained word root combinations in the matching test set by "word root pairing". Each group needs to count the number of times each pairing occurs (i.e., the number of successful pairings). At the same time, it can also count the data items of each word segmentation and sort the statistical results of each group according to the number of times the pairings occur.
[0147] For example, taking the word root combinations in Table 5 as an example, the frequency and data items of each word root combination and its respective pairings are shown in Table 6:
[0148] Table 6
[0149] Root word combination 1 3485 Data item 1, Hit data item 2 Root word combination 2 2641 Hit data item 3 Root word combination 3 1780 Hit data item 4, hit data item 5 Root word combination 4 1258 Hit data item 6, hit data item 7 Root word combination 5 421 Hit data item 8 6 word root combinations 255 Hit data item 9, hit data item 10
[0150] S604: Construct data elements.
[0151] Specifically, during step S604, after the server statistically analyzes the root word pairings, it can select at least one root word combination from the obtained root word combinations that can be used to construct data elements. Then, based on the selected at least one root word combination, the corresponding data elements are constructed. The specific method is as follows:
[0152] S6041: Select the root combination to construct the data element.
[0153] For example, taking the word root pairing results recorded in Table 6 as an example, we calculate the cumulative distribution of each result, that is, the proportion of the cumulative pairing counts to the total number of pairing counts. Assuming the total number of pairing counts is 10,000, the cumulative distribution results are shown in Table 7:
[0154] Table 7
[0155]
[0156]
[0157] Furthermore, based on the table above, assuming the cumulative distribution threshold is 90%, that is, if the cumulative distribution of word root combinations is not less than the cumulative distribution threshold of 90%, then the corresponding word root combinations (i.e., word root combination 1, word root combination 2, word root combination 3 and word root combination 4) can be selected to construct data elements.
[0158] S6042: Construct data elements using word root combinations.
[0159] For example, when performing step S6042, after obtaining at least one word root combination, the server can construct the corresponding data element according to the composition of the data element (any combination of object class, characteristic word and representation word, such as "object class + characteristic word + representation word" or "characteristic word + representation word").
[0160] S605: Verify whether the data element meets the requirements. If yes, proceed to S606; otherwise, proceed to S604 until the constructed data element meets the requirements, i.e., the data element has passed the verification.
[0161] Specifically, during step S605, after the data elements are constructed, the server can apply the constructed data elements to the business system data of the data items obtained in step S601. Therefore, after performing several data item pairing operations, if the sum of the success probabilities of each data element pairing, i.e., the total success probability of pairing, is greater than the preset success probability threshold, then it can be determined that each data element meets the data element matching conditions, and thus it can be determined that the constructed data elements meet the requirements, i.e., pass the verification; otherwise, the requirements are not met, i.e., the verification fails.
[0162] S606: Stop the construction of data elements.
[0163] In summary, the data element construction method provided in this application embodiment obtains each data item involved in the target business scenario and its respective data item name, and filters out the keyword unit subsets of each data item name that meet the preset keyword unit conditions from the keyword unit set of each data item name. Then, based on the keyword units contained in each obtained keyword unit subset, data elements corresponding to each data item are generated. Each data element represents at least one keyword unit of the corresponding data item. When it is determined that each obtained data element meets the preset data element matching conditions, each data element is saved to the preset target data element set.
[0164] This approach generates data elements corresponding to each data item based on the keyword units contained in each subset of obtained keyword units, thus achieving automatic construction of data elements. Furthermore, when it is determined that each obtained data element meets the preset data element matching conditions, each data element is saved to the preset target data element set. This ensures the probability of successful pairing of the constructed data elements with data items in the target business scenario, improving the accuracy of data element construction. As such, it avoids the technical drawbacks of existing technologies, such as the need to pre-construct data elements or construct new data elements based on existing data elements, and the low efficiency of traditional data element construction methods. Therefore, it improves the efficiency of data element construction.
[0165] Furthermore, based on the same technical concept, embodiments of this application provide a data element construction apparatus for implementing the above-described method flow of embodiments of this application. See also... Figure 7As shown, the data element construction device includes: an acquisition module 701, a generation module 702, and a filtering module 703, wherein:
[0166] The acquisition module 701 is used to acquire each data item involved in the target business scenario and its respective data item name, and to filter out the keyword unit subset of each data item name that meets the preset keyword unit conditions from the keyword unit set of each data item name.
[0167] The generation module 702 is used to generate data elements corresponding to each data item based on the keyword units contained in each of the obtained subsets of keyword units; wherein each data element represents: at least one keyword unit of the corresponding data item;
[0168] The filtering module 703 is used to save each data element to a preset target data element set when it is determined that each obtained data element meets the preset data element matching conditions.
[0169] In one possible embodiment, when selecting a subset of keyword units for each data item name that satisfy preset keyword unit conditions from the keyword unit sets for each data item name, the acquisition module 701 is specifically used for:
[0170] Perform the following operations for each data item name:
[0171] Divide a data item name into a set of keyword units for that data item name, and determine the part-of-speech of each keyword unit contained in the set of keyword units.
[0172] From each keyword unit, at least one keyword unit with the part of speech of noun is selected, and the set of nouns consisting of at least one keyword unit is taken as the subset of keyword units that meet the keyword unit conditions.
[0173] In one possible embodiment, when generating data elements corresponding to each data item based on the keyword units contained in each of the obtained subsets of keyword units, the generation module 702 is specifically used for:
[0174] Based on the keyword units contained in each subset of keyword units and the number of keyword units, the keyword unit combination corresponding to each data item is obtained;
[0175] When matching data items in the matching test set with a set number of test pairings, obtain the matching success probability of each keyword unit combination, and select at least one keyword unit combination from each keyword unit combination according to the preset matching success probability selection conditions.
[0176] Based on at least one keyword unit combination, generate the corresponding data elements for each data item.
[0177] In one possible embodiment, when obtaining the keyword unit combination corresponding to each data item based on the keyword units contained in each subset of keyword units and the number of keyword units therein, the generation module 702 is specifically used for:
[0178] For each subset of keyword units, perform the following operations:
[0179] Determine the number of keyword units contained in a subset of keyword units;
[0180] If the number of keyword units is not greater than the preset threshold for the number of units, then the sequence of keyword units contained in a subset of keyword units will be used as the keyword unit combination of the corresponding data item.
[0181] If the number of keyword units exceeds the threshold, then select the keyword unit sequence that meets the threshold from the keyword units contained in a subset of keyword units, and use the keyword unit sequence as the keyword unit combination of the corresponding data item.
[0182] In one possible embodiment, when generating data elements corresponding to each data item based on at least one keyword unit combination, the generation module 702 is specifically used for:
[0183] If at least one keyword unit combination satisfies the preset data element composition conditions, then the keyword unit combination can be used as the data element of the corresponding data item.
[0184] In one possible embodiment, when it is determined that each of the obtained data elements satisfies a preset data element matching condition, the filtering module 703 is specifically used for:
[0185] When matching data items with a set target number of matching times in the target business scenario, obtain the matching success probability of each data element, and obtain the corresponding total matching success probability based on the obtained matching success probabilities.
[0186] When the total probability of successful pairing is greater than the preset probability threshold for successful pairing, each data element is determined to meet the data element matching condition.
[0187] Based on the same technical concept, embodiments of this application also provide an electronic device that can implement the data element construction method flow provided in the above embodiments of this application. In one embodiment, the electronic device can be a server, a terminal device, or other electronic devices. Figure 8As shown, the electronic device may include:
[0188] At least one processor 801 and a memory 802 connected to at least one processor 801. In this embodiment, the specific connection medium between the processor 801 and the memory 802 is not limited. Figure 8 The example shown is the connection between processor 801 and memory 802 via bus 800. Bus 800 is... Figure 8 The connections between other components are indicated by thick lines and are for illustrative purposes only, not as limiting information. The 800 bus can be divided into address bus, data bus, control bus, etc., for ease of representation. Figure 8 The term is represented by a single thick line, but this does not imply that there is only one bus or one type of bus. Alternatively, the processor 801 can also be called a controller; there is no restriction on the name.
[0189] In this embodiment, memory 802 stores instructions executable by at least one processor 801. By executing the instructions stored in memory 802, at least one processor 801 can execute a data element construction method described above. Processor 801 can implement... Figure 7 The functions of each module in the device shown.
[0190] The processor 801 is the control center of the device. It can connect to various parts of the control device through various interfaces and lines. By running or executing instructions stored in memory 802 and calling data stored in memory 802, the processor can perform various functions and process data, thereby monitoring the device as a whole.
[0191] In one possible design, processor 801 may include one or more processing units. Processor 801 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may also not be integrated into processor 801. In some embodiments, processor 801 and memory 802 may be implemented on the same chip; in some embodiments, they may also be implemented on separate chips.
[0192] The processor 801 can be a general-purpose processor, such as a CPU, digital signal processor, application-specific integrated circuit, field-programmable gate array or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, capable of implementing or executing the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of a data element construction method disclosed in the embodiments of this application can be directly manifested as execution by a hardware processor, or execution by a combination of hardware and software modules within the processor.
[0193] Memory 802, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules. Memory 802 may include at least one type of storage medium, such as flash memory, hard disk, multimedia card, card-type memory, random access memory (RAM), static random access memory (SRAM), programmable read-only memory (PROM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), magnetic storage, magnetic disk, optical disk, etc. Memory 802 can be any other medium capable of carrying or storing desired program code in the form of instructions or data structures that can be accessed by a computer, but is not limited thereto. In the embodiments of this application, memory 802 can also be a circuit or any other device capable of implementing storage functions for storing program instructions and / or data.
[0194] By designing and programming the processor 801, the code corresponding to the data element construction method described in the foregoing embodiments can be embedded into the chip, thereby enabling the chip to execute the code during runtime. Figure 2 The illustrated embodiment presents the steps of a data element construction method. How to design and program the processor 801 is a technique well-known to those skilled in the art and will not be described further here.
[0195] Based on the same inventive concept, embodiments of this application also provide a storage medium storing computer instructions that, when executed on a computer, cause the computer to perform a data element construction method described above.
[0196] In some possible implementations, this application also provides a method for constructing data elements that can also be implemented as a program product including program code that, when the program product is run on a device, causes the control device to perform the steps of a method for constructing data elements according to various exemplary embodiments of this application as described above.
[0197] It should be noted that although several units or sub-units of the device have been mentioned in the detailed description above, this division is merely exemplary and not mandatory. In fact, according to embodiments of this application, the features and functions of two or more units described above can be embodied in one unit. Conversely, the features and functions of one unit described above can be further divided and embodied by multiple units.
[0198] Furthermore, although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0199] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0200] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a server, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0201] Program code for performing the operations of this application can be written using any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0202] In cases involving remote computing devices, the remote computing device can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computing device (e.g., via the Internet using an Internet service provider).
[0203] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0204] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0205] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Therefore, if such modifications and variations fall within the scope of the claims of this application and their equivalents, this application also intends to include such modifications and variations.
Claims
1. A method of data element construction, characterized by, include: Obtain all data items involved in the target business scenario and their respective data item names, and filter out the keyword unit subsets of each data item name that meet the preset keyword unit conditions from the keyword unit set of each data item name; wherein, the keyword unit conditions represent: whether the part of speech of the keyword units classified into the keyword unit subset meets the preset part of speech requirements; the part of speech requirements are any one of various parts of speech; the various parts of speech include at least verbs, nouns, and adjectives; Based on the keyword units contained in each of the obtained subsets of keyword units, data elements corresponding to each data item are generated; wherein, each data element represents: at least one keyword unit of the corresponding data item; When it is determined that each obtained data element meets the preset data element matching conditions, the data element is saved to the preset target data element set.
2. The method of claim 1, wherein, The step of selecting a subset of keyword units from the keyword unit sets of each data item name that meets the preset keyword unit conditions includes: For each of the data item names, perform the following operations: A data item name is divided to obtain a set of keyword units for the data item name, and the part-of-speech of each keyword unit contained in the set of keyword units is determined. From the various keyword units, at least one keyword unit with the part of speech of noun is selected, and the set of nouns formed by the at least one keyword unit is taken as the subset of keyword units that satisfy the keyword unit condition.
3. The method of claim 1, wherein, The process of generating data elements corresponding to each data item based on the keyword units contained in each obtained subset of keyword units includes: Based on the keyword units contained in each of the keyword unit subsets and the number of keyword units therein, the keyword unit combination corresponding to each data item is obtained. When matching data items in a matching test set with a set number of matching tests, obtain the matching success probability of each keyword unit combination, and select at least one keyword unit combination from the keyword unit combinations according to the preset matching success probability selection conditions. Based on the combination of at least one keyword unit, data elements corresponding to each data item are generated.
4. The method of claim 3, wherein, The step of obtaining the keyword unit combination corresponding to each data item based on the keyword units contained in each subset of keyword units and the number of keyword units in each subset includes: For each of the aforementioned keyword unit subsets, perform the following operations respectively: Determine the number of keyword units contained in a subset of keyword units; If the number of keyword units is not greater than a preset threshold for the number of units, then the sequence of keyword units contained in the subset of keyword units will be used as the keyword unit combination of the corresponding data item. If the number of keyword units is greater than the threshold number of units, then from the keyword units contained in the subset of keyword units, a sequence of keyword units that meets the threshold number of units is selected, and the sequence of keyword units is used as the keyword unit combination of the corresponding data item.
5. The method of claim 3, wherein, The step of generating data elements corresponding to each data item based on the combination of at least one keyword unit includes: If, among the at least one keyword unit combination, there exists a keyword unit combination that satisfies the preset data element composition conditions, then the keyword unit combination is used as a data element of the corresponding data item.
6. The method according to any one of claims 1-5, characterized in that, The determined data elements satisfy preset data element matching conditions, including: When matching data items in the target business scenario with a set number of matching attempts, obtain the matching success probability of each data element, and obtain the corresponding total matching success probability based on the obtained matching success probabilities. When the total probability of successful pairing is greater than the preset probability threshold for successful pairing, it is determined that each data element meets the data element matching condition.
7. A data element construction apparatus, characterized in that, include: The acquisition module is used to acquire each data item involved in the target business scenario and its respective data item name, and to filter out the keyword unit subsets of each data item name that meet the preset keyword unit conditions from the keyword unit set of each data item name; wherein, the keyword unit conditions represent: whether the part of speech of the keyword unit classified into the keyword unit subset meets the preset part of speech requirements; the part of speech requirements are any one of various parts of speech; the various parts of speech include at least verbs, nouns, and adjectives; The generation module is used to generate data elements corresponding to each data item based on the keyword units contained in each of the obtained subsets of keyword units; wherein each data element represents: at least one keyword unit of the corresponding data item; The filtering module is used to save each data element to a preset target data element set when it is determined that each obtained data element meets the preset data element matching conditions.
8. The apparatus as claimed in claim 7, characterized in that, When selecting a subset of keyword units that satisfy preset keyword unit conditions from the keyword unit sets for each data item name, the acquisition module is specifically used for: For each of the data item names, perform the following operations: A data item name is divided to obtain a set of keyword units for the data item name, and the part-of-speech of each keyword unit contained in the set of keyword units is determined. From the various keyword units, at least one keyword unit with the part of speech of noun is selected, and the set of nouns formed by the at least one keyword unit is taken as the subset of keyword units that satisfy the keyword unit condition.
9. The apparatus as claimed in claim 7, characterized in that, When generating the data elements corresponding to each data item based on the keyword units contained in each of the obtained subsets of keyword units, the generation module is specifically used for: Based on the keyword units contained in each of the keyword unit subsets and the number of keyword units therein, the keyword unit combination corresponding to each data item is obtained. When matching data items in a matching test set with a set number of matching tests, obtain the matching success probability of each keyword unit combination, and select at least one keyword unit combination from the keyword unit combinations according to the preset matching success probability selection conditions. Based on the combination of at least one keyword unit, data elements corresponding to each data item are generated.
10. The apparatus as claimed in claim 9, characterized in that, When obtaining the keyword unit combination corresponding to each data item based on the keyword units contained in each subset of keyword units and the number of keyword units in each subset, the generation module is specifically used for: For each of the aforementioned keyword unit subsets, perform the following operations respectively: Determine the number of keyword units contained in a subset of keyword units; If the number of keyword units is not greater than a preset threshold for the number of units, then the sequence of keyword units contained in the subset of keyword units will be used as the keyword unit combination of the corresponding data item. If the number of keyword units is greater than the threshold number of units, then from the keyword units contained in the subset of keyword units, a sequence of keyword units that meets the threshold number of units is selected, and the sequence of keyword units is used as the keyword unit combination of the corresponding data item.
11. The apparatus as claimed in claim 9, characterized in that, When generating the data elements corresponding to each data item based on the combination of at least one keyword unit, the generation module is specifically used for: If, among the at least one keyword unit combination, there exists a keyword unit combination that satisfies the preset data element composition conditions, then the keyword unit combination is used as a data element of the corresponding data item.
12. The apparatus according to any one of claims 7-11, characterized in that, When each of the obtained data elements satisfies the preset data element matching conditions, the filtering module is specifically used for: When matching data items in the target business scenario with a set number of matching attempts, obtain the matching success probability of each data element, and obtain the corresponding total matching success probability based on the obtained matching success probabilities. When the total probability of successful pairing is greater than the preset probability threshold for successful pairing, it is determined that each data element meets the data element matching condition.
13. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1-6.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1-6.
15. A computer program product, characterized in that, When the computer program product is invoked by a computer, it causes the computer to perform the method as described in any one of claims 1-6.
Citation Information
Patent Citations
Benchmarking method and system for data items, files and databases
CN110196834A
Search method and device based on user intention recognition
CN111400436A