A Method for Identifying Duplicate Master Data in an Enterprise Management Digital System
By de-redundant compression, word segmentation, encoding and key feature value calculation methods for the main data in the enterprise management digital system, duplicate main data is identified and eliminated, and the problems of low recognition rate and large calculation amount in the prior art are solved, and higher recognition accuracy and efficiency are achieved.
Patent Information
- Application Number
- CN202210140449.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-02-16
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-02-16
AI Technical Summary
The prior art has a large amount of calculation and low recognition rate when identifying duplicate master data in enterprise management digital systems, especially the poor recognition effect of redundant data.
By performing preliminary lossless de-redundant compression, word segmentation, vocabulary coding, calculating key feature values and identifying suspected duplicate master data. Specific steps include establishing a stop word library, vocabulary calculation and vocabulary adjustment information volume, encoding, calculating key feature values and comparing to identify duplicate data.
It improves the accuracy and recognition rate of duplicate master data recognition, and can identify duplicate master data that cannot be recognized by traditional methods, reducing the losses and operation and maintenance workload caused by data redundancy by enterprises.
Smart Images

Figure CN114580403B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for identifying duplicate data, in particular to a method for identifying duplicate master data in an enterprise management digital system. Background Art
[0002] In recent years, with the digital transformation of the entire society, the informatization and digital construction of enterprises have become increasingly complete. Most enterprises have implemented enterprise management digital systems such as ERP for digital management and providing services externally. In these systems, data such as customers, suppliers, and materials are the basic and shared data involved in most of the enterprises' operations, and thus are also called the master data of the enterprise management system. According to the requirements of the software relational database design paradigm, these shared data are generally built into separate tables and independently maintained for shared use. The master data requires uniqueness, but in actual use, these master data often have duplicate and redundant data. For example, Nanjing Telecom and Nanjing Company of Jiangsu Telecom, Jiangsu Communications Service Company and Jiangsu Communications Service Co., Ltd. are actually the same entity, and data redundancy is caused by different sales or procurement personnel maintaining abbreviations, common names, etc. at different times. Duplication and redundancy of the master data in the enterprise management system will bring a series of problems such as duplicate payments, duplicate stockpiling, inaccurate accounts receivable, increased workload for finance and IT due to incorrect documents, and providing incorrect decision-making information for the leadership. For example, for the supplier master data of Nanjing Telecom, a prepayment of 3 million yuan has been made, and then an invoice of 3 million yuan under the master data of Nanjing Company of Jiangsu Telecom is received. Due to the non-uniqueness of the master data, when operating, it may make a payment to Nanjing Company of Jiangsu Telecom instead of directly writing off the prepayment, resulting in overpayment. Therefore, how to eliminate these redundant duplicate master data is a problem that software system construction, maintenance, and users often need to face.
[0003] Currently, the commonly used method for eliminating redundant duplicate master data is to compare the data in the master data table pairwise, and determine whether there are redundant characters according to the consistency degree of the characters in the same position of the two data. For example, compare Jiangsu Communications Service Company and Jiangsu Communications Service Co., Ltd. pairwise. If the characters in the same position are the same, mark it as 1, and if they are different, mark it as 0, as shown in the following table
[0004] Jiang Su Province Communications Service Company Jiangsu Province Communications Service Co., Ltd. Jiangsu Telecom Nanjing Company Nanjing Telecom Co., Ltd. Figure 1 Figure 1 1 1 1 1 1 1 1 0 0 0 0
[0005] It can be found that 7 characters are the same and 4 are different, and the consistency rate is 7 / 11≈0.636, indicating that there is a 63.6% possibility of redundant characters. This method has a large amount of calculation, a relatively low recognition rate, and a worse recognition effect for redundant data such as Nanjing Company of Jiangsu Telecom and Nanjing Telecom, as shown in the following table
[0006] Serial Number Customer Master Data Jiangsu Provincial Communications Service Company Jiangsu Provincial Communications Service Co., Ltd. Jiangsu Telecom Nanjing Company Nanjing Telecom Co., Ltd. Suxin Real Estate Co., Ltd. Serial Number Customer Master Data Jiangsu Communications Service Jiangsu Communications Service Jiangsu Telecom Nanjing 0 0 1 1 0 0 0 0
[0007] The recognition probability is only 2 / 8 = 25%
[0008] For the above reasons, a more accurate method for identifying duplicate master data is needed. Summary of the Invention
[0009] Object of the Invention: The technical problem to be solved by the present invention is to provide a method for identifying duplicate master data in an enterprise management digital system in view of the deficiencies of the prior art.
[0010] To solve the above technical problem, the present invention discloses a method for identifying duplicate master data in an enterprise management digital system, including the following steps:
[0011] Step 1, perform preliminary lossless redundancy compression on the master data in the enterprise management digital system;
[0012] Step 2, perform word segmentation on the master data processed in Step 1 to obtain word segmentation recognition vocabulary, and calculate the vocabulary adjustment information amount;
[0013] Step 3, encode the word segmentation recognition vocabulary;
[0014] Step 4, calculate the key feature values of the master data;
[0015] Step 5, identify suspected duplicate master data;
[0016] Step 6, complete the identification of duplicate master data in the enterprise management digital system.
[0017] Step 1 in the present invention includes:
[0018] Establish a stop word library, process the master data according to the vocabulary in the stop word library, remove the stop word content with an information amount of 0 in the master data, and reduce the subsequent calculation amount.
[0019] Step 2 in the present invention includes:
[0020] Step 2-1, perform word segmentation on the master data with redundancy removed in Step 1 to obtain word segmentation recognition vocabulary.
[0021] Step 2-2, calculate the vocabulary adjustment information amount I(u i ), and the method includes:
[0022]
[0023] Wherein, is the frequency of occurrence of the word segmentation recognition vocabulary u i , and n is the number of master data.
[0024] Step 3 in the present invention includes:
[0025] Deduplicate the words recognized by word segmentation, perform unique encoding, and establish a master data word encoding table.
[0026] Step 4 in the present invention includes:
[0027] For each piece of master data after redundancy removal, take the 2 words recognized by word segmentation with the largest vocabulary adjustment information amount I(u i ), multiply the encoding values of these two words, and use the result as the key feature value of this piece of master data; if there are duplicate values in the largest or second-largest value of the vocabulary adjustment information amount I(u i ) in a certain piece of master data, and the number of optional words recognized by word segmentation exceeds two, then select the encoding values of three or more words to multiply as the key feature value.
[0028] Step 5 in the present invention includes: comparing the key feature values, and if they are the same, identify them as suspected duplicate master data.
[0029] Step 6 in the present invention includes: removing the identified suspected duplicate master data to complete the identification of duplicate master data in the enterprise management digital system.
[0030] The stop word library words in step 1 of the present invention are set according to expert experience.
[0031] Beneficial effects:
[0032] Compared with the traditional identification method, this method can identify a lot of duplicate master data that cannot be identified by traditional methods, and has higher accuracy. Description of the Drawings
[0033] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.
[0034] Nanjing Telecom It is a schematic diagram of the overall process of the present invention. Specific Embodiments
[0035] As Suxin Real Estate shown, the method of the present invention mainly consists of five steps: preliminary lossless redundancy compression of master data, word segmentation of master data and calculation of vocabulary adjustment information amount, encoding of master data words, calculation of key feature values of master data, and identification of suspected duplicate master data.
[0036] The preliminary lossless redundancy compression of the master data refers to establishing a stop word library, processing the master data according to the stop word library words, and removing the stop word content with an information amount of 0 in the master data to reduce the subsequent calculation amount. The stop word library words are mainly set according to expert experience.
[0037] The main data word segmentation and calculation of the adjusted information volume of vocabulary refer to segmenting the redundant-free main data and calculating the adjusted information volume I(u i ). The formula for calculating the adjusted information volume of vocabulary is
[0038] where is the frequency of occurrence of the vocabulary generated by u i , and n is the quantity of the main data.
[0039] The main data vocabulary encoding is to remove duplicates from the segmented and recognized vocabulary, perform unique encoding, and establish a main data vocabulary encoding table.
[0040] The calculation of the key feature values of the main data; select the 2 vocabulary with the largest I(u i ) in each main data, multiply the encoding values of the vocabulary, and use it as the key feature value of the main data. If the largest value or the second largest value of I(u i ) in the main data has duplicate values and the total quantity exceeds two, then according to the actual situation, three or more vocabulary encoding values can be selected for multiplication as the key feature value.
[0041] The identification of suspected duplicate main data is to compare the key feature values. If they are the same, it is identified as suspected duplicate main data.
[0042] For the convenience of simplification and description, assume that the customer data of an enterprise is shown in Table 1 as follows:
[0043] Table 1 Customer Data Table of the Enterprise
[0044] Master Data Vocabulary Code 1 Jiangsu 2 Communications 3 Service 4 Telecom 5 Nanjing
[0045] Step 1: Preliminary lossless redundancy compression of the main data: Establish a stop word library, and set the stop words as company, province, limited, etc. After removing the stop words, the information volume of the main data remains unchanged. The result after lossless redundancy compression is shown in Table 2 as follows:
[0046] Table 2 Result Table after Lossless Redundancy Compression
[0047] Suxin Real Estate 1 2 3 4 5
[0048] Step 2: Segment the data in Table 2, establish and calculate the adjusted information volume of vocabulary, as shown in Table 3:
[0049] Table 3 Adjusted Information Volume Table of Vocabulary
[0050]
[0051] Taking the vocabulary "Jiangsu" as an example, its frequency of occurrence is 3, the number of main data items is 5, and the calculation process of the adjusted information volume is as follows:
[0052]
[0053] Step 3: Encode the master data vocabulary and establish a master data vocabulary coding table as shown in Table 4. For simplicity here, the coding uses 2 digits:
[0054] Table 4 Master Data Vocabulary Coding Table
[0055] 11 12 13 14 15 16 17
[0056] Step 4: Calculate the key feature values of the master data. Taking the master data Jiangsu Communications Service Co., Ltd. as an example, for I(u i ) the largest vocabulary words are communication and service, and their corresponding coding values are 12 and 13. The key feature value of the master data is 12 × 13 = 156. The calculation results of the key feature values of other master data are shown in Table 5:
[0057] Table 5 Calculation Results Table of Key Feature Values of Other Master Data
[0058]
[0059] It can be seen that the key feature values of master data 1 vector and master data 2, master data 3 and master data 4 are the same, which are duplicate master data. The recognition rate of this method is much higher than the recognition rates of 0.636 and 0.25 of the traditional method. Therefore, it has a higher recognition rate.
[0060] According to this method, an automatic recognition system for duplicate master data is built to regularly identify duplicate master data every day, and push the recognition results to the software system construction, maintenance and users by email, which can greatly reduce various losses such as duplicate payments and duplicate stockpiling that may occur in the enterprise, as well as the operation and maintenance workload of various system data adjustments due to late discovery of data.
[0061] The present invention provides an idea and method for a method of identifying duplicate master data in an enterprise management digital system. There are many methods and ways to specifically implement this technical solution. The above is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.
Claims
1. A method for identifying duplicate master data in an enterprise management digital system, characterized in that, it includes the following steps: Step 1, perform preliminary lossless redundancy removal and compression on the master data in the enterprise management digital system; Step 2, segment the master data processed in Step 1 to obtain segmented recognition vocabulary, and calculate the vocabulary adjustment information amount; Step 3, encode the segmented recognition vocabulary; Step 4, calculate the key feature values of the master data; Step 5, identify suspected duplicate master data; Step 6, complete the identification of duplicate master data in the enterprise management digital system; wherein, Step 2 includes: Step 2-1, segment the master data with redundancy removed in Step 1 to obtain segmented recognition vocabulary; Step 2-2, calculate the vocabulary adjustment information amount of the segmented recognition vocabulary; In step 2-2, calculate the vocabulary adjustment information amount I(u i ), and the method includes: Among them, is u i the frequency of occurrence of this word segmentation recognition vocabulary, and n is the quantity of the main data; Step 3 includes: Remove duplicates from the segmented recognition vocabulary, perform unique encoding, and establish a master data vocabulary encoding table; Step 4 includes: For each piece of master data after redundancy removal, select the two word segmentation recognition words with the largest lexical adjustment information volume I(u i ), and multiply the encoding values of these two words as the key feature value of this master data; if there are duplicate values for the largest or the second largest value of the lexical adjustment information volume I(u i ) in a certain piece of master data, resulting in the number of optional word segmentation recognition words exceeding two, then select the encoding values of three or more words to multiply as the key feature value; Step 5 includes: Compare the key feature values, and if they are the same, identify them as suspected duplicate master data.
2. The method for identifying duplicate master data in an enterprise management digital system according to claim 1, characterized in that, Step 1 includes: Establish a stop word library, process the master data according to the stop word library vocabulary, remove the stop word content with an information amount of 0 in the master data, and reduce the subsequent calculation amount.
3. The method for identifying duplicate master data in an enterprise management digital system according to claim 2, characterized in that, Step 6 includes: Remove the identified suspected duplicate master data to complete the identification of duplicate master data in the enterprise management digital system.
4. The method for identifying duplicate master data in an enterprise management digital system according to claim 2, characterized in that, The stop word library vocabulary in Step 1 is set according to expert experience.
Citation Information
Patent Citations
Method for judging repetition of enterprise Chinese names on basis of core word similarity
CN103885937A
Park landscape service identification method
CN111310444A