A company name matching method
By adopting the company name matching method in the B2B field, and using cleaning, buffer truncation, pinyin conversion and double-array Trie tree technology, the problem of low efficiency of company name matching in the existing technology is solved, and efficient matching is achieved in the environment of large data volume.
Patent Information
- Application Number
- CN202210193820.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-01
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-03-01
AI Technical Summary
When the prior art processes large data volumes in the B2B field, the company name matching efficiency is low, making it difficult to effectively process unstandard, erroneous, and fuzzy company name data.
A company name matching method is adopted, which includes reading and cleaning the company name, creating a three-dimensional array through buffer truncation and pinyin conversion, building a double-array Trie tree, and using weight value matching and filtering results.
It improves the efficiency and accuracy of company name matching, and can quickly match results in a large amount of data environment, avoiding the problem of Hash trees occupying a lot of space and low query efficiency.
Smart Images

Figure CN114547151B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer data analysis, and in particular to a company name matching method. Background Art
[0002] For companies with B2B business models, the corresponding customers are also companies, and the premise of customer analysis is the basic attributes of customers. The basic data of a company often uses the company name as the primary key for association matching, so being able to effectively deal with the problems caused by company name matching is an important problem that must be overcome for companies with B2B business models. There are many sources of inaccurate company name data information, data errors, and fuzzy data. Some are caused by homophones in manual entry by sales staff, special characters in data stored in the business database, full-width and half-width input method problems, and simplified company name entry. These have a significant impact on the matching of company names, so how to deal with them is particularly important for subsequent analysis and mining.
[0003] Matching is essentially a data processing technology that can provide customers with extended data in multiple dimensions for business management. How to quickly obtain matching results in the face of large amounts of data in the big data era and effectively use this data to improve business capabilities directly determines the purpose and value of matching.
[0004] There are some cases in the prior art that match by formalized company name (RXIO), but in practical applications, especially in the B2B field, especially when faced with huge amounts of data, it still runs very hard and the efficiency does not meet expectations. If it is not improved, it will be difficult to deploy and use in practice.
[0005] Therefore, a company name matching method with better performance and more applicability is needed. Summary of the invention
[0006] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a company name matching method, which mainly provides a corresponding processing method for the situations where the company name data in the big data system is non-standard, the data is wrong, and the data is vague, thereby improving the matching degree and data utilization.
[0007] In order to solve the above technical problems, the present invention provides a company name matching method, which is characterized by comprising the following steps:
[0008] Step 1: Read the company information in the business database, obtain and clean the company name;
[0009] Step 2: Buffer and truncate the cleaned company name to obtain a formalized company name and store it in a created temporary table containing business database data information;
[0010] Step 3: Convert the formalized company name into a combination of English letters, and use the initial consonants, final consonants and homophones of Pinyin to create a three-dimensional array to express the formalized company name;
[0011] Step 4: construct a double array Trie tree of the formalized company name, store the attributes in the temporary table into the corresponding leaf nodes, and store the leaf nodes into memory;
[0012] Step 5: Read the company information of the third party and process it according to steps 1 to 4 to obtain the corresponding formalized company name;
[0013] Step 6: Match the formalized company name of the third-party company with the double-array Trie tree, set the weight of each category in the double-array Trie tree, and when the double-array Trie tree of one category is matched, take the result as a data row and jump out, and enter the double-array Trie tree of the next category after accumulating the weight value internally. When the double-array Trie tree of all categories is matched, start matching the next company information;
[0014] Step 7: Use the accumulated results of the weight values after matching in step 6 to build a temporary table; filter the results according to the matching threshold.
[0015] In the step 1, the reading of company information includes reading a mapping between the full name and the abbreviation of the company constructed in the company information to convert the abbreviation into the full name of the company; the cleaning includes removing special symbols, full-width characters, and half-width characters through regular expressions.
[0016] In the step 2, the buffer truncation specifically involves splitting the company name into preset keywords, which include region, keyword, industry and company suffix, and combining the preset keywords except the company suffix into a formalized company name.
[0017] In the step three, converting the formalized company name into a combination of English letters specifically includes: converting the preset keywords separated out in step two into a combination of English letters to express the formalized company name composed of the preset keywords.
[0018] In the step 4, constructing a dual-array Trie tree of formalized company names includes: constructing a dual-array Trie tree with preset keywords as categories.
[0019] In step six, matching the formalized company name of the third-party company with the dual-array Trie tree includes: matching the preset keywords in the formalized company name of the third-party company with the dual-array Trie tree one by one, and setting the weight of the dual-array Trie tree with each preset keyword as a category.
[0020] In the step 7, using the accumulated results of the weight values to construct a temporary table includes: according to the accumulated results of the weight values, taking the company name matching results that reach a preset threshold, and inserting them into the specified temporary table through the Scala language.
[0021] The beneficial effects achieved by the present invention are as follows: the present invention proposes a company name matching method, which uses full pinyin and establishes a three-dimensional array as a mapping to divide Chinese characters into three parts: vowels, consonants, and homophones, and converts the letter matching brought by pinyin into the matching of three numbers to improve efficiency. The construction of a double array prevents the construction of a Hash tree from occupying a large amount of space. It is difficult to ensure O(1) when the amount of data is large, and the Hash contains a large number of pointers, which is not friendly to languages containing gc, so that the results can be quickly matched even in a large amount of data environment. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 A schematic diagram of a method flow of an exemplary embodiment of the present invention;
[0023] Figure 2 is a schematic diagram of a module structure in an exemplary embodiment of the present invention;
[0024] Figure 3 It is a schematic diagram of data flow in an exemplary embodiment of the present invention. DETAILED DESCRIPTION
[0025] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.
[0026] In the present invention, the company name is formalized from the region, keyword, industry and company suffix by custom rules after reading the configuration file data, and the company name is mapped by converting Chinese characters into numbers and a dictionary tree containing the company information of the business library is created to achieve matching while matching other data information. Finally, a temporary table is constructed with the matched data and inserted into the formulated table, such as Figure 1 The data processing flow of an embodiment of the present invention shown includes the following steps:
[0027] Step 1: Read the company information in the business database, obtain and clean the company name;
[0028] Step 2: Buffer and truncate the cleaned company name to obtain a formalized company name, and store it in a created temporary table containing business database data information;
[0029] Step 3, convert the formalized company name into a combination of English letters, and use the pinyin initials, finals and homophones to create a three-dimensional array to express the formalized company name;
[0030] Step 4: Construct a double-array Trie tree for the formal company names, store the attributes in the temporary table into the corresponding leaf nodes, and store the leaf nodes in memory.
[0031] Step 5: Read the company information of the third party, process it according to Steps 1 to 4 to obtain the corresponding formal company names.
[0032] Step 6: Match the formal company names of the third-party companies with the double-array Trie tree, set the weights for each category in the double-array Trie tree. When the matching of one category of the double-array Trie tree is completed, take the result as a data row and jump out, and accumulate the weight values internally and then enter the double-array Trie tree of the next category. When the matching of all categories of the double-array Trie tree is completed, start the matching of the next company information.
[0033] Step 7: Use the accumulated result of the weight values after the matching in Step 6 to construct a temporary table.
[0034] Step 8: Filter out the results according to the matching threshold.
[0035] In an exemplary embodiment of the present invention, the specific steps are as follows:
[0036] Step 11: Read the company information in the business library, obtain and clean the company names. The cleaning includes removing special symbols, full-width and half-width characters through regular expressions, and converting possible abbreviations into the full company names by reading the mapping between the full company names and abbreviations constructed in the company information. For example, "CNOOC" corresponds to "China National Offshore Oil Corporation", effectively preventing the loss of the form or key information of the company name.
[0037] Step 12: Perform buffer truncation on the cleaned data. After the company names in the business library are cleaned and mapped (abbreviation initialization), they become standard company names. Therefore, split the standard company names into formal company names in the form of a combination of region (Region), keyword (X), industry (Industry), and company suffix (Org_Suffix), and create a temporary table containing the data information of the business library. In the embodiment of the present invention, a row of data about the company name information in the temporary table is as follows:
[0038] Region(Nanjing) Keyword(XXX) Industry (Technology) Company Suffix (Limited Liability Company)
[0039] Step 13: Convert the keywords in the split company names into English pinyin and, according to the characteristics of Chinese pinyin, use initials, finals, and the homophone library to form a three-dimensional array corresponding to numbers, expressing a Chinese character skillfully using numbers, effectively shortening the number of bytes occupied by the company name, improving the execution efficiency, and alleviating the program pressure. The example is as follows, taking "Ke" corresponding to "868" as an example:
[0040] Initial consonant array:
[0041] ... k ... ... 8 ...
[0042] Finals array:
[0043] ... e ... ... 6 ...
[0044] Homophone table:
[0045] ... division ... ... 8 ...
[0046] Step 14: Construct a double-array Trie tree with preset keywords as categories, store the corresponding attributes of the formalized company name in the temporary table into the corresponding leaf nodes, and store the leaf nodes in memory.
[0047] Step 15: Read the company information that the third party needs to match with the business database, and clean and convert it according to steps 11 to 14.
[0048] Step 16: Match the processed third-party company information with the double-array Trie tree one by one, and set the weight of each category double-array Trie tree. If the match is completed, the relevant information is directly taken as a data row and jumped out, and the weight value is accumulated internally to enter the double-array Trie tree of the next category. The matching of all category double-array Trie trees is completed, and the next data is started.
[0049] Step 17: Build a temporary table with the above processed results and take the matching results that reach the preset threshold according to the accumulated weight value, and insert them into the specified table through the Scala language.
[0050] Step 18: Filter the results based on the matching threshold.
[0051] like Figure 2 As shown, an exemplary embodiment of the present invention discloses a company name matching system based on the above method, including: a data source module, a data preprocessing module, a dictionary tree construction and data matching module, and a data filtering module connected in sequence.
[0052] The data source module is used to read business library related data, third-party library data and read company name truncation initialization condition configuration, which may come from business systems, text logs and other data structure sources.
[0053] The data preprocessing module cleans and standardizes the company names in the business database and the third-party company names, and preprocesses the data for the field tree construction module and the data matching module.
[0054] The dictionary tree construction module and data matching module express the company name through the digital expression module, and construct a double-array Trie tree for each pre-processed business library company name through the double-array Trie tree matching module. In order to avoid the need to associate the third-party company information table and the business library company information table again after processing to obtain the corresponding information and cause re-matching, the company information corresponding to the business library is stored in the leaf node of the tree through the leaf node module when the double-array Trie tree is reconstructed, that is, the leaf node is used to store the basic company information. According to the preset matching logic, if the match is successful, the weight value is calculated by the calculation module, and it is used as a row of records for the subsequent construction of a temporary table. If the match is successful, it will jump out and start the next match.
[0055] The data filtering module constructs a temporary table with rows of data generated by the matching module through Scala sample classes, and finally inserts the data into the corresponding library table after threshold filtering.
[0056] like Figure 3 As shown in the figure, it is a specific example of the data matching module. Nanjing Lao Wang Aquatic Products is split into "Nanjing", "Lao Wang", and "Aquatic Products" according to the keywords, and the pinyin initials, finals, and homophones are converted into numbers, such as "456789", "545668", and "851999". The double-array Trie tree constructed according to the region, keyword, and industry is matched respectively to determine whether it is a hit. If it is a hit, the corresponding weight is accumulated, otherwise it is recorded as 0, and the sum of the weights is calculated.
[0057] In the present invention, since the essence of the Trie tree is a deterministic finite state automaton (DFA), the core idea is to trade space for time, and the common prefix of the string is used to reduce the query time overhead to achieve the purpose of improving efficiency. However, due to the serious sparseness of the Trie tree, the space utilization rate is low. In order to make the Trie tree occupy less space and ensure the query efficiency at the same time, it is finally proposed to use two linear arrays to represent the Trie tree, that is, the double array Trie (DoubleArray Trie).
[0058] In addition to constructing based on a double array, the present invention also compresses the leaf nodes, shortens the time of finding unoccupied space for child nodes, reduces the construction time, reduces the space occupied by the array, and improves the query speed.
[0059] The present invention proposes a company name matching method based on full spelling and Aho-Corasick algorithm. Full spelling is used to establish a three-dimensional array as a mapping to divide Chinese characters into three parts: initial consonants, finals, and homophones. Then English letters are converted into more concise digital combinations, and the letter matching brought by pinyin is converted into the matching of three numbers to improve efficiency. The construction of a double array improves the defects that a large amount of space is occupied when constructing a Hash tree, it is difficult to ensure O(1) when the amount of data is large, and the Hash contains a large number of pointers, which is not friendly to the language containing gc, so that the results can be quickly matched even in a large amount of data environment.
[0060] The above embodiments do not limit the present invention in any way. Any other improvements and applications made to the above embodiments in an equivalent transformation manner belong to the protection scope of the present invention.
Claims
1. A company name matching method, characterized in that: The steps include: Step 1: Read the company information in the business database, obtain and clean the company name; Step 2: Buffer and truncate the cleaned company name to obtain a formalized company name and store it in a created temporary table containing business database data information; Step 3: Convert the formalized company name into a combination of English letters, and use the initial consonants, final consonants and homophones of Pinyin to create a three-dimensional array to express the formalized company name; Step 4: construct a double array Trie tree of the formalized company name, store the attributes in the temporary table into the corresponding leaf nodes, and store the leaf nodes into memory; Step 5: Read the company information of the third party and process it according to steps 1 to 4 to obtain the corresponding formalized company name; Step 6: Match the formalized company name of the third-party company with the double-array Trie tree, set the weight of each category in the double-array Trie tree, and when the double-array Trie tree of one category is matched, take the result as a data row and jump out, and enter the double-array Trie tree of the next category after accumulating the weight value internally. When the double-array Trie tree of all categories is matched, start matching the next company information; Step 7: Use the accumulated results of the weight values after matching in step 6 to build a temporary table, and filter the results according to the matching threshold.
2. A company name matching method as claimed in claim 1, characterized in that: In the step 1, reading the company information includes reading the mapping between the full name and the abbreviation constructed in the company information to convert the abbreviation into the full name of the company; and the cleaning includes removing special symbols, full-width characters, and half-width characters through regular expressions.
3. A company name matching method as claimed in claim 2, characterized in that: In the step 2, the buffer truncation specifically involves splitting the company name into preset keywords, which include region, keyword, industry and company suffix, and combining the preset keywords except the company suffix into a formalized company name.
4. A company name matching method as claimed in claim 3, characterized in that: In the step three, converting the formalized company name into a combination of English letters specifically includes: converting the preset keywords separated out in step two into a combination of English letters to express the formalized company name composed of the preset keywords.
5. A company name matching method as claimed in claim 4, characterized in that: In the step 4, constructing a dual-array Trie tree of formalized company names includes: constructing a dual-array Trie tree with preset keywords as categories.
6. A company name matching method as claimed in claim 5, characterized in that: In step six, matching the formalized company name of the third-party company with the dual-array Trie tree includes: matching the preset keywords in the formalized company name of the third-party company with the dual-array Trie tree one by one, and setting the weight of the dual-array Trie tree with each preset keyword as a category.
7. A company name matching method as claimed in claim 6, characterized in that: In the step 7, using the accumulated results of the weight values to construct a temporary table includes: according to the accumulated results of the weight values, taking the company name matching results that reach a preset threshold, and inserting them into the specified temporary table through the Scala language.
Citation Information
Patent Citations
Syllable segmentation method and device
CN109377980A
Method and device of matching speech input to text
WO2014201834A1