Industry tag matching methods, devices, equipment and storage media
By differentiating between listed and unlisted companies and using different matching methods, this technology utilizes multi-dimensional data scoring and correlation calculation to solve the problem of inaccurate industry labels in existing technologies, achieving more accurate matching of enterprise industry labels.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHINA MOBILE GROUP ZHEJIANG
- Filing Date
- 2022-12-07
- Publication Date
- 2026-07-31
AI Technical Summary
Existing technologies use the same logic when matching industry labels to listed and unlisted companies, resulting in inaccurate industry labels.
Different matching methods are used for industry tag matching depending on whether the company is publicly listed. For publicly listed companies, tags are determined by acquiring and scoring the amount of matching data across multiple primary business data dimensions; for privately listed companies, tags are determined by acquiring and calculating the relevance across multiple secondary business data dimensions.
It improves the accuracy and reliability of industry tag matching, avoiding tag omissions and inaccuracies caused by a single data dimension.
Smart Images

Figure CN116956898B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to an industry tag matching method, apparatus, device and storage medium. Background Technology
[0002] Industry tags, also known as industry classification tags, are used to indicate the industry to which a company belongs.
[0003] Companies are categorized as listed or unlisted, and the difference in information transparency between the two leads to different types of information that can be collected. Currently, the same matching logic is used to assign industry tags to both listed and unlisted companies, resulting in inaccurate industry tags. Summary of the Invention
[0004] In view of this, embodiments of this application provide an industry label matching method, apparatus, device, and storage medium, aiming to improve the accuracy of matching industry labels for enterprises.
[0005] To achieve the above objectives, this application provides an industry tag matching method, the method comprising:
[0006] Determine whether the company to be matched is a listed company;
[0007] If so, then industry tags will be matched to the companies to be matched based on the matching method for listed companies;
[0008] If not, then industry tags will be matched to the companies to be matched based on the non-listed company matching method.
[0009] For example, matching industry tags to the companies to be matched based on the listed company matching method includes:
[0010] Obtain the first business data of the enterprise to be matched, each first business data includes a first data dimension, and each first data dimension includes at least the product name;
[0011] Each industry's first preset business dataset is matched with each first business data to obtain multiple first matched data volumes; each first preset business dataset includes the first data dimension included in each first business data.
[0012] The companies to be matched are scored based on the amount of first matching data, resulting in multiple scores.
[0013] Industry tags are matched to listed companies based on the scores.
[0014] For example, the step of matching each first preset business dataset of each industry with each first business data to obtain multiple first matched data volumes includes:
[0015] Construct a TrieTree for the first pre-defined business dataset of each industry;
[0016] Each of the first business data is segmented into words to obtain multiple first word segmentation results;
[0017] Each TrieTree is matched with its corresponding first word segmentation result to obtain multiple first matching data sets.
[0018] For example, the scoring of the companies to be matched is based on each first matching data volume to obtain multiple scores, including:
[0019] The following formula is used to calculate the score of the enterprise to be matched in any first data dimension. :
[0020]
[0021] in, This is the default value; This represents the first set of matching data.
[0022] For example, the method of matching industry tags to the company to be matched based on non-listed companies includes:
[0023] Obtain the second business data of the enterprise to be matched. Each second business data includes a second data dimension, and the product name is not included in each second data dimension.
[0024] The second preset business datasets for each industry are matched with each second business data to obtain multiple second matched data volumes; the second preset business datasets include the second data dimensions included in each second business data.
[0025] Obtain the weighted ratio of each second data dimension in each industry;
[0026] The first correlation between the company to be matched and any industry is calculated using the following formula. :
[0027]
[0028] in, This serves as the identifier for the second data dimension. For the first The second matching data volume of the second data dimension For the first The weighting ratio of the second data dimension;
[0029] Determine the correlation thresholds for each industry;
[0030] If the first relevance of the enterprise to be matched If the relevance threshold of any industry is greater than or equal to that of any industry, then the industry label of the company to be matched is determined to be the industry label of that industry.
[0031] For example, obtaining the weighted ratio of each second data dimension in each industry includes:
[0032] The weighting ratio for any industry is determined as follows:
[0033] Match the second preset business dataset of any industry with the third business data of each listed company in any industry to obtain multiple third matching data volumes;
[0034] The sum of the third matching data volume for each second data dimension is calculated separately to obtain the total matching data volume for each second data dimension.
[0035] Calculate the weighting ratio of each second data dimension using the following formula. :
[0036]
[0037] in, This serves as the identifier for the second data dimension. For the first The total amount of matching data in each of the second data dimensions.
[0038] For example, determining the correlation threshold for each industry includes:
[0039] The correlation threshold for any industry is determined in the following way:
[0040] Determine the second degree of correlation for each non-listed company in any of the aforementioned industries;
[0041] Sort the second degree of correlation by size;
[0042] Determine the third degree of correlation for each listed company in any of the aforementioned industries;
[0043] Sort the third degree of correlation by size;
[0044] Calculate the standard deviation of each second correlation degree using the following formula. :
[0045]
[0046] in, This represents the total number of non-listed companies. For non-listed companies, For the first The second degree of connection of an unlisted company The average of all second-degree associations;
[0047] exist as well as When, determine the The correlation threshold is denoted as , where To be the minimum second degree of association, It represents the minimum third degree of association.
[0048] Furthermore, to achieve the above objectives, this application also provides an industry label matching device, the device comprising:
[0049] The determination module is used to determine whether the company to be matched is a listed company;
[0050] The first matching module is used to match industry tags to the company to be matched based on the matching method for listed companies if the case is true.
[0051] The second matching module is used to match industry tags to the company to be matched based on the non-listed company matching method if no.
[0052] In addition, to achieve the above objectives, this application also provides an industry label matching device, which includes a memory, a processor, and an industry label matching program stored in the memory and executable on the processor. When the industry label matching program is executed by the processor, it implements the steps of the industry label matching method as described above.
[0053] In addition, to achieve the above objectives, this application also provides a computer-readable storage medium storing an industry label matching program, which, when executed by a processor, implements the steps of the industry label matching method as described above.
[0054] Compared to existing technologies that use the same matching logic to match industry tags for both listed and unlisted companies, leading to inaccurate industry tags, this application determines whether the company to be matched is a listed company. If so, it matches industry tags to the company based on the matching method for listed companies; otherwise, it matches industry tags to the company based on the matching method for unlisted companies. This application first classifies the companies to be matched into listed and unlisted companies. For listed companies, it matches industry tags to the company based on the matching method for listed companies; for unlisted companies, it matches industry tags to the company based on the matching method for unlisted companies. This differentiates the matching of industry tags for listed and unlisted companies, avoiding the use of the same matching logic for matching industry tags for both listed and unlisted companies, thereby improving the accuracy of matching industry tags to the companies to be matched. Attached Figure Description
[0055] Figure 1 This is a flowchart illustrating the first embodiment of the industry label matching method of this application;
[0056] Figure 2 This is a schematic diagram of the functional modules of a preferred embodiment of the industry label matching device of this application;
[0057] Figure 3 This is a schematic diagram of the hardware operating environment involved in the embodiments of this application.
[0058] The realization of the purpose, functional features and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation
[0059] It should be understood that the specific embodiments described herein are merely illustrative of this application and are not intended to limit this application.
[0060] This application provides an industry label matching method, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the industry label matching method of this application.
[0061] This application provides embodiments of an industry label matching method. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order. The industry label matching method can be applied to servers or personal computers. For ease of description, the following description of the various steps of the industry label matching method is omitted. The industry label matching method includes:
[0062] Step S110: Determine whether the company to be matched is a listed company;
[0063] The difference in information transparency between listed and unlisted companies mainly lies in the structured nature of their product data and the concentration of data. Listed companies are required to publish financial reports regularly, such as annual and quarterly reports, which include a detailed statement of their main business revenue and expenditure. This statement contains structured product data. Unlisted companies are not required to publish financial reports, so it may be impossible to obtain their structured product data. Listed companies have high information transparency and a wider range of data collection channels, while unlisted companies have low information transparency and a narrower range of data collection channels.
[0064] Whether a company to be matched is a listed company can be determined manually, by searching on the Internet, or by looking up a table, etc. This embodiment does not specifically limit the method of determination. Among them, determining by searching a search engine can be based on web crawling technology, for example, crawling relevant information on whether the company to be matched is listed from web pages; the table used for determining by looking up a table is predetermined, that is, the information of each company is collected in advance, and a mapping relationship between each company and whether it is listed is established.
[0065] Step S120: If yes, then match industry tags to the company to be matched based on the listed company matching method.
[0066] The matching method for listed companies is to automatically retrieve the first business data on the Internet, or to match industry tags based on the first business data retrieved when determining whether the company to be matched is a listed company.
[0067] Step S130: If not, then match industry tags to the company to be matched based on the non-listed company matching method.
[0068] For non-listed companies, the matching method involves adding supplementary information obtained from offline channels based on the retrieved information to obtain second business data, which is then used to match industry tags.
[0069] Compared to existing technologies that use the same matching logic to match industry tags for both listed and unlisted companies, leading to inaccurate industry tags, this application determines whether the company to be matched is a listed company. If so, it matches industry tags to the company based on the matching method for listed companies; otherwise, it matches industry tags to the company based on the matching method for unlisted companies. This application first classifies the companies to be matched into listed and unlisted companies. For listed companies, it matches industry tags to the company based on the matching method for listed companies; for unlisted companies, it matches industry tags to the company based on the matching method for unlisted companies. This differentiates the matching of industry tags for listed and unlisted companies, avoiding the use of the same matching logic for matching industry tags for both listed and unlisted companies, thereby improving the accuracy of matching industry tags to the companies to be matched.
[0070] For example, matching industry tags to the companies to be matched based on the listed company matching method includes:
[0071] Step a: Obtain the first business data of the enterprise to be matched. Each first business data includes a first data dimension, and each first data dimension includes at least the product name.
[0072] The first data dimension includes product name, patent name, and company name, etc.
[0073] Specifically, for product names, the main business revenue and expenditure details in the financial statements of the companies to be matched are identified, extracted, and manually reviewed to obtain the company's main products; and the corresponding main product terms are identified from unstructured text to obtain the first business data corresponding to the product name, such as XX router.
[0074] For patent names, we identify, extract, and manually review the patent names of the companies to be matched, and obtain the first business data corresponding to the patent name, such as XXX method and system.
[0075] For example, the first business data is XXX Co., Ltd.
[0076] Step b involves matching each first preset business dataset of each industry with each first business data to obtain multiple first matched data volumes; each first preset business dataset includes the first data dimension included in each first business data.
[0077] The difference between the first preset business dataset and the first business data lies in the fact that after obtaining the business data of benchmark enterprises, it is necessary to perform word segmentation processing on it. The process of obtaining the business data of benchmark enterprises is basically the same as that of obtaining the first business data, and will not be repeated here.
[0078] Among them, benchmark enterprises are determined according to industry classification standards and based on the annual revenue and number of employees of listed companies. Specifically, a preset number of the most representative enterprises in each industry are selected, and the preset number can be 10, 20, etc. This embodiment does not make a specific limitation.
[0079] The first matching data volume is the number of times that words in the first preset business dataset are successfully matched with words in the first business data.
[0080] To improve the matching degree between the data in the first preset business dataset and the first business data, semantic analysis of the benchmark enterprise's business data can be performed through NLP (Natural Language Processing) semantic understanding, thereby configuring synonyms or near-synonyms for the benchmark enterprise's business data, so that the first preset business dataset contains not only the benchmark enterprise's business data, but also its synonyms or near-synonyms.
[0081] Furthermore, for any industry, each of the first preset business datasets can also be undifferentiated by the second data dimension, that is, the business data of each benchmark enterprise can be placed in one dataset without distinction, in order to reduce data storage costs.
[0082] Step c: Score the companies to be matched based on each first matching data volume to obtain multiple scores.
[0083] Step d: Match industry tags to listed companies based on the multiple scores.
[0084] Calculate the total score using the following formula. :
[0085]
[0086] in, The first matching data volume corresponding to the product name. This represents the first matching data volume corresponding to the patent name. The score corresponding to the product name. The score corresponding to the patent name. The score corresponds to the company name.
[0087] The above formula is an empirical formula. When the business data also includes other second data dimensions, the total score will be... You can also add scores related to other second data dimensions, or scores and the amount of data matching the first data dimension. This embodiment does not specifically limit the specific method.
[0088] After obtaining the total score Then, compare the score with the total score threshold for the corresponding industry. If... If the score is greater than or equal to the total score threshold, a match is successful, meaning the industry label for that industry is determined to be the industry label for the company to be matched. If the score is less than the total score threshold, the match fails. The total score threshold is an empirical value.
[0089] It should be noted that for any industry, as long as If the total score is greater than or equal to the threshold, then the industry label of any industry can be determined as the industry label of the company to be matched. Existing technologies determine the target industry from a large number of industries based on the company name and business scope, and then determine the industry label of the target industry as the industry label of the company to be matched. Therefore, the existing technology determines only one industry label, while this embodiment can determine multiple industry labels. Thus, it can be seen that the existing technology may miss industry labels belonging to the company to be matched when matching industry labels, while this embodiment can remedy this deficiency, thereby improving the accuracy of matching industry labels.
[0090] Furthermore, compared to the first business data with multiple second data dimensions in this embodiment, the existing technology uses relatively simple data when matching industry tags for companies to be matched, only including the company name and business scope. This results in less information for matching industry tags, leading to low accuracy and reliability of the matched industry tags. Correspondingly, this embodiment uses first business data with multiple second data dimensions to match industry tags for companies to be matched, improving the accuracy and reliability of the matched industry tags.
[0091] For example, the step of matching each first preset business dataset of each industry with each first business data to obtain multiple first matched data volumes includes:
[0092] Step b1: Construct the TrieTree for the first preset business dataset of each industry;
[0093] Step b2: Match each TrieTree with the corresponding first business data to obtain multiple first matching data volumes.
[0094] A TrieTree is a variant of a hash tree, and it has a tree-like structure. In the first preset business dataset, a word is a string, and in the TrieTree, a word is a node.
[0095] Matching is achieved using algorithms such as the Aho-Corasick automaton (AC) and the Aho-Corasick Boyer-Moore (ACBM).
[0096] Taking the Aho-Corasick automaton as an example, the Aho-Corasick automaton mainly relies on constructing a finite state machine. It performs noun matching based on a trie and counts the number of matches. After all nodes in the trie have completed matching, the first set of matched data is obtained.
[0097] For example, the scoring of the companies to be matched is based on each first matching data volume to obtain multiple scores, including:
[0098] The following formula is used to calculate the score of the enterprise to be matched in any first data dimension. :
[0099]
[0100] in, This is the default value; This represents the first set of matching data.
[0101] This preset value is an empirical value, such as the value corresponding to the product name. It scores 1.7 points, corresponding to the patent name. It scores 2.1 points, corresponding to the company name. It scored 3.6 points.
[0102] For example, the method of matching industry tags to the company to be matched based on non-listed companies includes:
[0103] Step e: Obtain the second business data of the enterprise to be matched. Each second business data includes a second data dimension, and each second data dimension does not include the product name.
[0104] The second data dimension in the second business data includes company name, business scope, company introduction, patent name, etc.
[0105] Step f involves matching the second preset business datasets of each industry with each second business data to obtain multiple second matched data volumes; the second preset business datasets include the second data dimensions included in each second business data.
[0106] The second pre-defined business dataset includes the segmentation results obtained after segmenting industry standard vocabulary, as well as synonyms or near-synonyms for each segmentation result. Similarly, the synonyms or near-synonyms for the segmentation results can be obtained by performing semantic analysis on the industry standard vocabulary using NLP semantic understanding and then configuring them.
[0107] The industry standard terminology is extracted from relevant information of listed companies in a particular industry. This relevant information includes the second data dimension, which is part of the second business data.
[0108] The second matching data volume refers to the number of times words in the second preset business dataset successfully match words in the second business data. For example, the second matching data volumes for company name, business scope, company introduction, patent name, etc., are A1, A2, A3, ..., An, respectively. It should be noted that the specific implementation method of the matching process in this embodiment is basically the same as the specific implementation method of matching the first preset business dataset with the first business data described above, and will not be repeated here.
[0109] Step g: Obtain the weighting ratio of each second data dimension in each industry;
[0110] Step h: Calculate the first correlation between the company to be matched and any industry using the following formula. :
[0111]
[0112] in, This serves as the identifier for the second data dimension. For the first The second matching data volume of the second data dimension For the first The weighting ratio of the second data dimension;
[0113] Step i, determine the correlation threshold for each industry;
[0114] Step j, if the first relevance of the enterprise to be matched... If the relevance threshold of any industry is greater than or equal to that of any industry, then the industry label of the company to be matched is determined to be the industry label of that industry.
[0115] First relevance of the companies to be matched If the relevance score is less than the relevance threshold for any industry, the match fails. The relevance threshold is an empirical value.
[0116] It should be noted that for any industry, as long as If the relevance threshold is greater than or equal to the relevance threshold, then the industry label of any industry can be determined as the industry label of the company to be matched. Existing technologies determine the target industry from a large number of industries based on the company name and business scope, and then determine the industry label of the target industry as the industry label of the company to be matched. Therefore, the existing technology determines only one industry label, while this embodiment can determine multiple industry labels. Thus, it can be seen that the existing technology may miss industry labels belonging to the company to be matched when matching industry labels, while this embodiment can remedy this deficiency, thereby improving the accuracy of matching industry labels.
[0117] Furthermore, compared to the second business data with multiple second data dimensions in this embodiment, the existing technology uses relatively simple data when matching industry tags for companies to be matched, only including the company name and business scope. This results in less information for matching industry tags, leading to low accuracy and reliability of the matched industry tags. Correspondingly, this embodiment uses second business data with multiple second data dimensions to match industry tags for companies to be matched, improving the accuracy and reliability of the matched industry tags.
[0118] For example, obtaining the weighted ratio of each second data dimension in each industry includes:
[0119] The weighting ratio for any industry is determined as follows:
[0120] Step g1: Match the second preset business dataset of any industry with the third business data of each listed company in any industry to obtain multiple third matching data volumes;
[0121] Step g2: Calculate the sum of the third matching data volume for each second data dimension to obtain the total matching data volume for each second data dimension.
[0122] The impact of word segmentation matching results on matching industry tags varies across different second data dimensions. Therefore, a word segmentation correlation model is established to quantify the contribution (impact) of the third matching data volume in different second data dimensions to industry tags. The process of establishing the word segmentation correlation model is as follows: For any industry tag, a list of listed companies belonging to that industry is generated based on the industry tag; the AC automaton algorithm is used to obtain the third matching data volume of each listed company in each second data dimension, and the total matching data volume of listed companies in each second data dimension is calculated, denoted as B1, B2, B3, ..., Bn.
[0123] Step g3: Calculate the weighting ratio of each second data dimension using the following formula. :
[0124]
[0125] in, This serves as the identifier for the second data dimension. For the first The total amount of matching data in each of the second data dimensions.
[0126] For example, determining the correlation threshold for each industry includes:
[0127] The correlation threshold for any industry is determined in the following way:
[0128] Step i1: Determine the second correlation degree of each non-listed company in any of the industries;
[0129] Step i2: Sort the second degree of correlation by size;
[0130] Step i3: Determine the third degree of correlation for each listed company in any of the industries.
[0131] Step i4: Sort the third degree of correlation by size;
[0132] Step i5: Calculate the standard deviation of each second correlation degree using the following formula. :
[0133]
[0134] in, This represents the total number of non-listed companies. For non-listed companies, For the first The second degree of connection of an unlisted company The average of all second-degree associations;
[0135] Step i6, in as well as When, determine the The correlation threshold is denoted as , where To be the minimum second degree of association, It represents the minimum third degree of association.
[0136] The degree of relevance between a company and an industry label is called the relevance score. The magnitude of the relevance score determines whether a company can be assigned that industry label.
[0137] Sorting can be done from largest to smallest or from smallest to largest.
[0138] It should be noted that, since there is no need to perform step-by-step analysis of the companies to be matched in order to determine the target industry to which the companies belong and to determine the industry label of the target industry as the industry label of the companies to be matched, in the context of big data, the industry label matching method of this application completes the matching process automatically and quickly, thereby making the matching of industry labels simple and fast, simplifying the matching process and reducing the consumption of human and material resources.
[0139] With the acceleration of marketization in my country, various ownership structures are developing rapidly, leading to a large number of cross-industry enterprises. Enterprises often need to strategize and adjust their development plans during their growth; therefore, they need to analyze their industrial chain and clearly define their industry. Thus, the appropriate matching of industry tags is particularly important.
[0140] In addition, refer to Figure 2 , Figure 2 This application also provides an industry label matching device, the device comprising:
[0141] Module 10 is used to determine whether the company to be matched is a listed company.
[0142] The first matching module 20 is used to match industry tags to the company to be matched based on the matching method of listed companies if the case is true.
[0143] The second matching module 30 is used to match industry tags to the company to be matched based on the non-listed company matching method if no.
[0144] For example, the first matching module 20 is specifically used for:
[0145] Obtain the first business data of the enterprise to be matched, each first business data includes a first data dimension, and each first data dimension includes at least the product name;
[0146] Each industry's first preset business dataset is matched with each first business data to obtain multiple first matched data volumes; each first preset business dataset includes the first data dimension included in each first business data.
[0147] The companies to be matched are scored based on the amount of first matching data, resulting in multiple scores.
[0148] Industry tags are matched to listed companies based on the scores.
[0149] For example, the first matching module 20 is further configured to:
[0150] Construct a TrieTree for the first pre-defined business dataset of each industry;
[0151] Each of the first business data is segmented into words to obtain multiple first word segmentation results;
[0152] Each TrieTree is matched with its corresponding first word segmentation result to obtain multiple first matching data sets.
[0153] For example, the first matching module 20 is further configured to:
[0154] The following formula is used to calculate the score of the enterprise to be matched in any first data dimension. :
[0155]
[0156] in, This is the default value; This represents the first set of matching data.
[0157] For example, the second matching module 30 is specifically used for:
[0158] Obtain the second business data of the enterprise to be matched. Each second business data includes a second data dimension, and the product name is not included in each second data dimension.
[0159] The second preset business datasets for each industry are matched with each second business data to obtain multiple second matched data volumes; the second preset business datasets include the second data dimensions included in each second business data.
[0160] Obtain the weighted ratio of each second data dimension in each industry;
[0161] The first correlation between the company to be matched and any industry is calculated using the following formula. :
[0162]
[0163] in, This serves as the identifier for the second data dimension. For the first The second matching data volume of the second data dimension For the first The weighting ratio of the second data dimension;
[0164] Determine the correlation thresholds for each industry;
[0165] If the first relevance of the enterprise to be matched If the relevance threshold of any industry is greater than or equal to that of any industry, then the industry label of the company to be matched is determined to be the industry label of that industry.
[0166] For example, the second matching module 30 is further configured to:
[0167] The weighting ratio for any industry is determined as follows:
[0168] Match the second preset business dataset of any industry with the third business data of each listed company in any industry to obtain multiple third matching data volumes;
[0169] The sum of the third matching data volume for each second data dimension is calculated separately to obtain the total matching data volume for each second data dimension.
[0170] Calculate the weighting ratio of each second data dimension using the following formula. :
[0171]
[0172] in, This serves as the identifier for the second data dimension. For the first The total amount of matching data in each of the second data dimensions.
[0173] For example, the second matching module 30 is further configured to:
[0174] The correlation threshold for any industry is determined in the following way:
[0175] Determine the second degree of correlation for each non-listed company in any of the aforementioned industries;
[0176] Sort the second degree of correlation by size;
[0177] Determine the third degree of correlation for each listed company in any of the aforementioned industries;
[0178] Sort the third degree of correlation by size;
[0179] Calculate the standard deviation of each second correlation degree using the following formula. :
[0180]
[0181] in, This represents the total number of non-listed companies. For non-listed companies, For the first The second degree of connection of an unlisted company The average of all second-degree associations;
[0182] exist as well as When, determine the The correlation threshold is denoted as , where To be the minimum second degree of association, It represents the minimum third degree of association.
[0183] The specific implementation of the industry label matching device in this application is basically the same as the embodiments of the above-described industry label matching method, and will not be repeated here.
[0184] In addition, this application also provides an industry label matching device. For example... Figure 3 As shown, Figure 3 This is a schematic diagram of the hardware operating environment involved in the embodiments of this application.
[0185] In one possible implementation, Figure 3 This is a schematic diagram of the hardware operating environment of the device that can be matched with industry labels.
[0186] like Figure 3 As shown, the industry label matching device may include a processor 301, a communication interface 302, a memory 303, and a communication bus 304. The processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304. The memory 303 is used to store computer programs. When the processor 301 executes the program stored in the memory 303, it implements the steps of the industry label matching method.
[0187] The communication bus 304 mentioned in the aforementioned industry label matching device can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus 304 can be divided into address bus, data bus, and control bus, etc. For ease of illustration, it is represented by only one thick line in the figure, but this does not indicate that there is only one bus or one type of bus.
[0188] Communication interface 302 is used for communication between the aforementioned industry label matching device and other devices.
[0189] The memory 303 may include random access memory (RMD) or non-volatile memory (NM), such as at least one disk storage device. Optionally, the memory 303 may also be at least one storage device located remotely from the aforementioned processor 301.
[0190] The processor 301 mentioned above can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0191] The specific implementation of the industry label matching device in this application is basically the same as the embodiments of the above-described industry label matching method, and will not be repeated here.
[0192] Furthermore, embodiments of this application also propose a computer-readable storage medium storing an industry label matching program, which, when executed by a processor, implements the steps of the industry label matching method described above.
[0193] The specific implementation of the computer-readable storage medium in this application is basically the same as the embodiments of the above-described industry label matching method, and will not be described again here.
[0194] In addition, to achieve the above objectives, this application also provides a computer program product, comprising: a computer program, which, when executed by a processor, implements the steps of the industry tag matching method as described above.
[0195] The specific implementation method of the computer program product in this application is basically the same as the various embodiments of the above-mentioned industry label matching method, and will not be described again here.
[0196] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0197] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0198] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes: several instructions to cause a terminal device (which may be a mobile phone, computer, server, device, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0199] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. An industry tag matching method, characterized in that, The method includes: Determine whether the company to be matched is a listed company; If so, then industry tags will be matched to the companies to be matched based on the matching method for listed companies; If not, then industry tags will be matched to the companies to be matched based on the non-listed company matching method; The step of matching industry tags to the companies to be matched based on the non-listed company matching method includes: Obtain the second business data of the enterprise to be matched. Each second business data includes a second data dimension, and the product name is not included in each second data dimension. Obtain the weighting ratio of each second data dimension in each industry, and determine the correlation threshold for each industry; The second preset business datasets of each industry are matched with each second business data to obtain multiple second matching data volumes; Based on the second data dimension, the second matching data volume, and the weighted ratio of the second data dimension, the first correlation degree between the enterprise to be matched and any industry is calculated; If the first relevance of the company to be matched is greater than or equal to the relevance threshold of any industry, then the industry label of the company to be matched is determined to be the industry label of any industry.
2. The method as described in claim 1, characterized in that, The method of matching industry tags to the companies to be matched based on listed companies includes: Obtain the first business data of the enterprise to be matched, each first business data includes a first data dimension, and each first data dimension includes at least the product name; Each industry's first preset business dataset is matched with each first business data to obtain multiple first matched data volumes; each first preset business dataset includes the first data dimension included in each first business data. The companies to be matched are scored based on the amount of first matching data, resulting in multiple scores. Industry tags are matched to listed companies based on the scores.
3. The method as described in claim 2, characterized in that, The step involves matching each first preset business dataset from each industry with each first business data to obtain multiple first matched data volumes, including: Construct a TrieTree for the first pre-defined business dataset of each industry; Each TrieTree is matched with its corresponding first business data to obtain multiple first matching data volumes.
4. The method as described in claim 3, characterized in that, The process of scoring the companies to be matched based on each first matching data volume yields multiple scores, including: The following formula is used to calculate the score of the enterprise to be matched in any first data dimension. : in, This is the default value; This represents the first set of matching data.
5. The method as described in claim 1, characterized in that, The second preset business dataset includes the second data dimension of each second business data. The step of calculating the first correlation between the enterprise to be matched and any industry based on the weighted ratio of the second data dimension, the second matching data volume, and the second data dimension includes: The first correlation between the company to be matched and any industry is calculated using the following formula. : in, This serves as the identifier for the second data dimension. For the first The second matching data volume of the second data dimension For the first The weighted ratio of the second data dimension.
6. The method as described in claim 5, characterized in that, The process of obtaining the weighted ratio of each second data dimension in each industry includes: The weighting ratio for any industry is determined as follows: Match the second preset business dataset of any industry with the third business data of each listed company in any industry to obtain multiple third matching data volumes; The sum of the third matching data volume for each second data dimension is calculated separately to obtain the total matching data volume for each second data dimension. Calculate the weighting ratio of each second data dimension using the following formula. : in, This serves as the identifier for the second data dimension. For the first The total amount of matching data in each of the second data dimensions.
7. The method as described in claim 6, characterized in that, The determination of the correlation threshold for each industry includes: The correlation threshold for any industry is determined in the following way: Determine the second degree of correlation for each non-listed company in any of the aforementioned industries; Sort the second degree of correlation by size; Determine the third degree of correlation for each listed company in any of the aforementioned industries; Sort the third degree of correlation by size; Calculate the standard deviation of each second correlation degree using the following formula. : in, This represents the total number of non-listed companies. For non-listed companies, For the first The second degree of connection of an unlisted company The average of all second-degree associations; exist as well as When, determine the The correlation threshold is denoted as , where To be the minimum second degree of association, It represents the minimum third degree of association.
8. An industry label matching device, characterized in that, The device includes: The determination module is used to determine whether the company to be matched is a listed company; The first matching module is used to match industry tags to the company to be matched based on the matching method for listed companies if the case is true. The second matching module is used to match industry tags to the company to be matched based on the non-listed company matching method if no; The second matching module is further configured to acquire each second business data of the enterprise to be matched, each second business data including a second data dimension, each second data dimension excluding product names, acquire the weighting ratio of each second data dimension in each industry, determine the correlation threshold of each industry, match the second preset business dataset of each industry with each second business data to obtain multiple second matching data volumes, and calculate the first correlation between the enterprise to be matched and any industry based on the second data dimension, the second matching data volume and the weighting ratio of the second data dimension. If the first correlation of the enterprise to be matched is greater than or equal to the correlation threshold of any industry, then the industry label of the enterprise to be matched is determined to be the industry label of any industry.
9. An industry label matching device, characterized in that, The industry label matching device includes a memory, a processor, and an industry label matching program stored in the memory and executable on the processor. When executed by the processor, the industry label matching program implements the steps of the industry label matching method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores an industry label matching program, which, when executed by a processor, implements the steps of the industry label matching method as described in any one of claims 1 to 7.