Enterprise industry classification method and apparatus

By extracting unique keywords and calculating matching scores from enterprise industry standard classification documents, the problem of incomplete enterprise industry classification caused by data resource deficiencies in existing technologies is solved, enabling rapid and accurate classification of enterprise industries, especially suitable for small sample enterprises.

CN119004177BActive Publication Date: 2026-07-28BEIJING CHIBO INFORMATION ENG CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CHIBO INFORMATION ENG CO LTD
Filing Date
2024-08-07
Publication Date
2026-07-28

AI Technical Summary

Technical Problem

Existing enterprise industry classification technologies cannot effectively match micro-companies or individual businesses when data resources are lacking or external databases are incomplete. Furthermore, artificial intelligence methods are not applicable to small samples and their algorithms are not transparent, making it difficult to broadly cover all enterprises.

Method used

By extracting unique keywords and industry classification parameters from enterprise industry standard classification files, and calculating matching scores based on whether the enterprise name contains these keywords, the most matching industry is determined, supporting rapid classification of enterprise industries.

Benefits of technology

It improves the universality of enterprise industry classification, especially suitable for small sample enterprises, and can quickly and accurately identify the industry to which an enterprise belongs, especially suitable for enterprises with obvious differences in nature.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119004177B_ABST
    Figure CN119004177B_ABST
Patent Text Reader

Abstract

The application provides an enterprise industry classification method and device, and relates to the technical field of data processing. The method comprises the following steps: extracting independent keywords and enterprise industry classification parameters corresponding to each industry from an enterprise industry standard classification file; determining the number of target independent keywords of each target industry matched with the enterprise to be classified according to whether the enterprise name of the enterprise to be classified contains the independent keywords; and calculating the matching scores of the enterprise to be classified corresponding to each target industry according to the number of target independent keywords, the number of occurrences, the total number of keywords, the number of keywords and the number of times, and determining the target industry with the highest matching score as the enterprise industry classification of the enterprise to be classified. The device executes the above method. The enterprise industry classification method provided in the application can improve the universality of enterprise industry classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and specifically to a method and apparatus for classifying enterprise industries. Background Technology

[0002] There is a widespread demand for enterprise industry classification. Existing enterprise industry classification technology mainly achieves this by matching enterprise names (or enterprise unified identification codes) with data collected in external databases. This type of existing technology can achieve a perfect match when the data resources are complete, but it has the following problems when the data resources are incomplete: (1) The matching is completely limited by the selection of external databases and the data collected. If the external database is defective, the data is incorrect or missing, or the identification codes are not uniform, there is no solution; (2) Traditional research focuses on large enterprises, so external databases do not often collect micro-companies or individual businesses that are small in size but numerous in number. Existing enterprise industry classification technology cannot cover the latter extensively.

[0003] In addition to the aforementioned traditional technologies, artificial intelligence methods (such as deep learning and clustering) have been used in recent years to classify enterprises and industries. The general problem with these methods is that they often require a large training set, are not suitable for small samples, and the algorithms themselves are opaque, unpredictable, prone to overfitting and non-standardization. Summary of the Invention

[0004] To address the problems in the prior art, embodiments of the present invention provide a method and apparatus for classifying enterprise industries, which can at least partially solve the problems existing in the prior art.

[0005] On the one hand, this invention proposes a method for classifying enterprise industries, including:

[0006] Based on the enterprise industry standard classification document, we extract the unique, deduplicated keywords and enterprise industry classification parameters corresponding to each industry.

[0007] The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file corresponding to each industry, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries.

[0008] The number of target independent keywords for each target industry that matches the enterprise to be classified is determined based on whether the enterprise name to be classified contains the independent keyword.

[0009] Based on the number of independent target keywords, the frequency of occurrence, the total number of keywords, the number of keywords, and the frequency, the matching score of the enterprise to be classified with each target industry is calculated, and the target industry with the highest matching score is determined as the enterprise industry classification of the enterprise to be classified.

[0010] The deduplicated independent keywords extracted from the enterprise industry standard classification document and corresponding to each industry include:

[0011] Extract the column containing the category name from the enterprise industry standard classification document, and perform preliminary extraction of the text content in the column containing the category name to obtain initial keywords;

[0012] The initial keywords are modified to obtain unique, deduplicated keywords corresponding to each industry.

[0013] The step of refining the initial keywords to obtain unique, deduplicated keywords corresponding to each industry includes:

[0014] Remove words and phrases irrelevant to the industry, and correct any mis-split words in the remaining words based on the category name corresponding to the column containing the category name;

[0015] Correcting incorrectly split words containing conjunctions, splitting the corrected words into keywords with a maximum of a preset character count, removing duplicates from these keywords, and obtaining unique keywords corresponding to each industry.

[0016] The step of determining the number of target independent keywords for each target industry that matches the enterprise to be classified based on whether the enterprise name contains the independent keyword includes:

[0017] The industry corresponding to at least one independent keyword contained in the company name shall be the target industry;

[0018] For each occurrence of an independent keyword contained in the company name within a target industry, the number of target independent keywords for that target industry is incremented by 1, thus obtaining the number of target independent keywords for each target industry that matches the company to be classified.

[0019] The enterprise industry classification method also includes:

[0020] The industries corresponding to independent keywords that are not included in the company name are directly identified as industry categories unrelated to the company to be classified.

[0021] The step of calculating the matching score between the enterprise to be classified and each target industry based on the number of target independent keywords, the frequency of occurrence, the total number of keywords, the number of keywords, and the frequency of occurrence includes:

[0022] The matching score corresponding to each target industry is calculated based on the following expression:

[0023]

[0024] Where w is the number of the target independent keywords, and n is the number of the target independent keywords. i Where N is the number of occurrences, M is the total number of keywords, and k is the number of keywords. i Let i represent the number of times, i represents the i-th target independent keyword, and I represents the target independent keyword set composed of all target independent keywords.

[0025] The enterprise industry classification method further includes, after the step of determining the target industry with the highest matching score as the enterprise industry classification category, the step of determining the enterprise industry classification category as follows:

[0026] In response to the review operation on the enterprise industry classification, if the review operation result is determined to be the first review operation result of incorrect classification, a prompt message is generated to optimize the enterprise industry classification parameters and / or independent keywords.

[0027] On one hand, the present invention proposes an enterprise industry classification device, comprising:

[0028] The acquisition unit is used to extract unique, deduplicated keywords and enterprise industry classification parameters corresponding to each industry based on the enterprise industry standard classification file.

[0029] The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file corresponding to each industry, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries.

[0030] The determining unit is used to determine the number of target independent keywords for each target industry that matches the enterprise to be classified, based on whether the enterprise name of the enterprise to be classified contains the independent keyword.

[0031] The classification unit is used to calculate the matching score between the enterprise to be classified and each target industry based on the number of target independent keywords, the number of occurrences, the total number of keywords, the number of keywords, and the number of occurrences, and to determine the target industry with the highest matching score as the enterprise industry classification of the enterprise to be classified.

[0032] In another aspect, embodiments of the present invention provide an electronic device, including: a processor, a memory, and a bus, wherein,

[0033] The processor and the memory communicate with each other via the bus;

[0034] The memory stores program instructions that can be executed by the processor, and the processor can execute the following methods by calling the program instructions:

[0035] Based on the enterprise industry standard classification document, we extract the unique, deduplicated keywords and enterprise industry classification parameters corresponding to each industry.

[0036] The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file corresponding to each industry, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries.

[0037] The number of target independent keywords for each target industry that matches the enterprise to be classified is determined based on whether the enterprise name to be classified contains the independent keyword.

[0038] Based on the number of independent target keywords, the frequency of occurrence, the total number of keywords, the number of keywords, and the frequency, the matching score of the enterprise to be classified with each target industry is calculated, and the target industry with the highest matching score is determined as the enterprise industry classification of the enterprise to be classified.

[0039] This invention provides a non-transitory computer-readable storage medium, comprising:

[0040] The non-transitory computer-readable storage medium stores computer instructions that cause the computer to perform the following methods:

[0041] Based on the enterprise industry standard classification document, we extract the unique, deduplicated keywords and enterprise industry classification parameters corresponding to each industry.

[0042] The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file corresponding to each industry, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries.

[0043] The number of target independent keywords for each target industry that matches the enterprise to be classified is determined based on whether the enterprise name to be classified contains the independent keyword.

[0044] Based on the number of independent target keywords, the frequency of occurrence, the total number of keywords, the number of keywords, and the frequency, the matching score of the enterprise to be classified with each target industry is calculated, and the target industry with the highest matching score is determined as the enterprise industry classification of the enterprise to be classified.

[0045] The enterprise industry classification method and apparatus provided in this invention extracts deduplicated independent keywords and enterprise industry classification parameters corresponding to each industry based on an enterprise industry standard classification file. The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears across different industries. The number of target independent keywords matching the enterprise to be classified is determined based on whether the enterprise name contains the independent keywords. A matching score is calculated between the enterprise to be classified and each target industry based on the number of target independent keywords, the number of occurrences, the total number of keywords, the number of keywords, and the number of occurrences. The target industry with the highest matching score is determined as the enterprise industry classification of the enterprise to be classified, thereby improving the universality of enterprise industry classification. Attached Figure Description

[0046] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. In the drawings:

[0047] Figure 1 This is a flowchart illustrating an enterprise industry classification method provided in an embodiment of the present invention.

[0048] Figure 2 This is a schematic diagram illustrating a portion of the content extracted from an enterprise industry standard classification file, as provided in an embodiment of the present invention.

[0049] Figure 3 This is a flowchart illustrating an enterprise industry classification method provided in another embodiment of the present invention.

[0050] Figure 4This is a flowchart illustrating an enterprise industry classification method provided in another embodiment of the present invention.

[0051] Figure 5 This is a schematic diagram of the structure of an enterprise industry classification device provided in an embodiment of the present invention.

[0052] Figure 6 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0053] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings. Here, the illustrative embodiments and descriptions of the present invention are used to explain the present invention, but are not intended to limit the present invention. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of this application can be arbitrarily combined with each other.

[0054] Explanation of relevant terms:

[0055] Sliding Window: A sliding window is a problem-solving technique related to data structures and algorithms, suitable for array or list problems. It was initially used for flow control. The sliding window protocol is applied in TCP / IP and is one of the core strategies for TCP flow control, used to manage network data transmission flow and prevent congestion. Later, it evolved into the sliding window algorithm, which primarily solves different problems by maintaining a window and moving its two boundaries.

[0056] Sliding window resource pool: Borrowing the idea of ​​thread pool, multiple sliding windows are maintained in the sliding window resource pool, allowing concurrent execution of tasks.

[0057] Figure 1 This is a flowchart illustrating an enterprise industry classification method according to an embodiment of the present invention, as shown below. Figure 1 As shown, the enterprise industry classification method provided in this embodiment of the invention includes:

[0058] Step S1: Extract the unique, deduplicated keywords and industry classification parameters corresponding to each industry based on the enterprise industry standard classification document;

[0059] The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file corresponding to each industry, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries.

[0060] Step S2: Determine the number of target independent keywords for each target industry that matches the enterprise to be classified, based on whether the enterprise name contains the independent keyword.

[0061] Step S3: Calculate the matching score between the enterprise to be classified and each target industry based on the number of target independent keywords, the number of occurrences, the total number of keywords, the number of keywords, and the number of occurrences. Determine the target industry with the highest matching score as the enterprise industry classification of the enterprise to be classified.

[0062] In step S1 above, the device extracts deduplicated independent keywords and enterprise industry classification parameters corresponding to each industry based on the enterprise industry standard classification file.

[0063] The enterprise industry classification parameters include the total number of keywords that repeatedly appear in the enterprise industry standard classification file corresponding to each industry, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries. The apparatus can be a computer device that performs the method, such as a server. It should be noted that the acquisition and analysis of data involved in this embodiment of the invention are authorized by the user. The enterprise industry standard classification file can specifically be a file corresponding to GB / T 4754-2017, such as... Figure 2 As shown, a certain keyword can appear multiple times in the same row of the same industry in the "Description" column. You can remove duplicates of this keyword, and the resulting unique keyword is the keyword that is retained after deduplication.

[0064] The unique, deduplicated keywords extracted from the enterprise industry standard classification document and corresponding to each industry include:

[0065] Extract the column containing the category name from the enterprise industry standard classification document, and perform preliminary extraction of the text content in the column containing the category name to obtain initial keywords;

[0066] The initial keywords are modified to obtain unique, deduplicated keywords corresponding to each industry.

[0067] The column containing the category name indicates the source of the industry keywords, such as... Figure 2 As shown, the industry classification codes, category names, and descriptions are listed at each level, for example... Figure 2 The industry classification code for "rice planting" is "A0111", where "A" is the category code, the next two "01" are the major category code, and the last two "1" are the sequence codes for the intermediate and minor categories, respectively.

[0068] Performing a correction process on the initial keywords to obtain deduplicated independent keywords corresponding to each industry respectively, including:

[0069] Removing words and phrases irrelevant to the industry, and correcting mis-split words among the remaining words according to the category names corresponding to the column where the category name is located;

[0070] Correcting mis-split words with conjunctions, splitting the completed corrected words to obtain keywords not exceeding the preset number of characters, and removing duplicates from the keywords not exceeding the preset number of characters to obtain deduplicated independent keywords corresponding to each industry respectively.

[0071] Such as Figure 3 As shown, the extraction of independent keywords mainly includes a splitting process and three correction processes, which are described as follows:

[0072] First, the open-source jieba library in Python can be used to perform a rough first-step splitting of the industry category names, and then common conjunctions (such as "and", "and others", "other", "of", etc.) and the punctuation marks that appear (such as the comma and Chinese and English brackets) are removed to obtain the initial keywords.

[0073] Since the category names are mostly short sentences, there are also many mistakes and omissions in the jieba library splitting, and the above initial keywords are corrected:

[0074] Screen out relatively obvious useless words and adjectives from all the single characters split out, remove words that are obviously irrelevant to the industry (such as "type", "contain", "in", "west", "similar"), and then combine the category names to correct the adjectives and mis-split words into common words (such as "light" combined with Figure 2 corrected to "light type", "year" combined with Figure 2 corrected to "annuity", etc.).

[0075] Correct mis-split words with conjunctions. For example, conjunctions such as "and others", "and" occasionally fail to be separated from other nouns (such as "salt and others"), and in this step, the mis-split words are corrected to normal words (such as "salt").

[0076] Continue to split keywords with four characters and five characters (the longest word after jieba splitting is five characters). Among the keywords with four characters, except for special cases (such as "carbon dioxide"), they are all split into two words according to 2 + 2, and among the keywords with five characters, except for special cases, they are split into two words according to the common language rule of 2 + 3 or 3 + 2.

[0077] Table 1

[0078]

[0079] Table 1 shows 10 examples illustrating the keyword breakdown results of industry category names. Approximately one longer category and keyword breakdown result was randomly selected from every 200 data points. Keywords could be extracted from all categories, including major, minor, and intermediate categories. Overall, two-character words were the most common, covering representative terms from various industries.

[0080] Before matching industries, it is necessary to analyze and statistically analyze the keyword list.

[0081] first Figure 2 The last column extracts all keywords related to each industry (category), with N representing the total number of keywords extracted for each industry (including duplicates) (i.e., the total number of keywords corresponding to each industry, including duplicates, in the enterprise industry standard classification file). i Let M represent the number of times each independent keyword appears (i.e., the number of times each independent keyword appears in the enterprise's industry standard classification file), and let M represent the total number of unique keywords in each industry (i.e., the number of unique keywords in each industry).

[0082] After calculating N and n i After M, we then analyze the keywords between industries, using k i This represents the number of times a keyword appears across different industries (i.e., the number of times each individual keyword appears in different industries). For example, if a keyword appears in two industries simultaneously, then k... i =2, if a keyword appears in three industries simultaneously, then k = 2. i =3, and so on. After all statistics are completed, a table containing industry (category code), keywords (unique keywords after deduplication), and enterprise industry classification parameters (N, n) will be generated. i M, k i The industry matching table is used for subsequent processing.

[0083] In step S2 above, the device determines the number of target independent keywords for each target industry that matches the enterprise to be classified based on whether the enterprise name contains the independent keyword. For example... Figure 4 As shown, the company names of companies to be categorized can be preprocessed before determining whether they contain independent keywords. Figure 4The left side represents the processing rules for company names, and the right side represents the processing rules for the industry matching table. Here, we will use a company, "XX Aviation Industry Group Co., Ltd.", as an example. First, company names generally contain many identical but industry-irrelevant words, such as "limited liability" and "group." The method of this invention includes preprocessing of the company name to remove these irrelevant words. The main purposes are twofold: first, to improve the speed of the subsequent matching algorithm; and second, to reduce matching interference that irrelevant words may cause.

[0084] For example, the remaining part of "XX Aviation Industry Group Co., Ltd." after removing irrelevant words is "aviation industry".

[0085] The step of determining the number of target independent keywords for each target industry that matches the enterprise to be classified based on whether the enterprise name contains the independent keyword includes:

[0086] The industry corresponding to at least one independent keyword contained in the company name shall be the target industry;

[0087] For each occurrence of an independent keyword contained in the company name within a target industry, the number of target independent keywords for that target industry is incremented by 1, thus obtaining the number of target independent keywords for each target industry that matches the company to be classified.

[0088] The enterprise industry classification method also includes:

[0089] The industries corresponding to independent keywords that are not included in the company name are directly identified as industry categories unrelated to the company to be classified. For industry categories unrelated to the company to be classified, the corresponding industry matching score calculation process can be skipped.

[0090] In step S3 above, the device calculates the matching score between the enterprise to be classified and each target industry based on the number of target independent keywords, the number of occurrences, the total number of keywords, the number of keywords, and the number of occurrences, and determines the target industry with the highest matching score as the enterprise industry classification of the enterprise to be classified.

[0091] The process of calculating the matching score between the enterprise to be classified and each target industry based on the number of target independent keywords, the frequency of occurrence, the total number of keywords, the number of keywords, and the frequency of occurrence includes:

[0092] The matching score corresponding to each target industry is calculated based on the following expression:

[0093]

[0094] Where w is the number of the target independent keywords, and n is the number of the target independent keywords. i Where N is the number of occurrences, M is the total number of keywords, and k is the number of keywords. i Let be the number of times, where i represents the i-th target independent keyword, and I represents the set of all target independent keywords. Referring to the above explanation, for industry categories unrelated to the company to be classified, w = 0, therefore the corresponding score = 0.

[0095] The underlying logic of this formula considers the importance of each keyword within the industry, as well as the number of keywords contained in a company name. For example, if a keyword appears frequently in an industry, then the keyword frequency... The score will be relatively high, and companies containing this keyword are more likely to belong to this industry, so it should account for a larger proportion in the score calculation formula.

[0096] For example, if a keyword appears in multiple industries, then the industry representativeness of this keyword is relatively poor, and it should be given a relatively small weight in the score statistics.

[0097] Since the number of keywords varies across different industries, generally speaking, the more unique keywords an industry has, the larger it is and the more companies it includes. Therefore, M is used as an adjustment item for industry size in the score.

[0098] The more industry-related keywords a company name contains, the more likely the company belongs to that industry. The formula uses 'w' to adjust the final matching score.

[0099] After the matching score for each industry is calculated, an industry matching score table containing all industries can be generated. If the matching score for all industries is 0, it means that the company name does not contain any industry keywords. In this case, we consider the industry classification to be unsuccessful.

[0100] If the industry matching score is not all zeros, then we consider the industry with the highest score to be the most likely correct industry category, and the algorithm will return the category code of this industry, indicating a successful industry match. For example, after the industry matching algorithm calculates the matching score for "XX Aviation Industry Group Co., Ltd.", it will generate the matching score table shown in Table 2:

[0101] Table 2

[0102] A 0 K 0 B 0 L 0 C 2.806 M 0.3636 D 0 N 0 E 0 O 0 F 0 P 0 G 1.700 Q 0 H 0 R 0 I 0 S 0 J 0 T 0

[0103] Among them, category C (manufacturing) has the highest score, and the algorithm of this invention will classify the enterprise into industry C.

[0104] This invention can quickly classify companies in a sample by industry. The results show that most matching failures are due to the company name itself not containing industry information, such as "XXX (Group) Co., Ltd." or "XXX XXX Group Holding Co., Ltd."

[0105] In addition, the examples of successful matches in the sample were manually checked, and the accuracy rate of matching was around 70%. The occurrence of matching errors was mainly affected by two factors:

[0106] (1) Some industry keywords are not commonly found in company names. For example, the word "retail" is often used in industry names, but actual retail companies often use words such as "supermarket" and "department store".

[0107] (2) Industry keywords are mainly two-character words, and there is overlap between keywords in related industries. Therefore, matching errors are likely to occur when single-character combinations appear in company names. For example, in "XXX Power Grid," both "electricity" and "grid" contain industry information, but in the keyword table, "electricity" mainly appears as two-character words such as "power generation." The matching accuracy of the algorithm can be improved in some cases by further refining the company industry classification parameters and / or the keyword table.

[0108] The enterprise industry classification method provided in this invention can be used for rapid classification of samples of any size, and is particularly suitable for large-sample classification studies that are difficult to process manually and for situations where sample enterprises lack effective external database matching information. The more industry-related information contained in the enterprise name, the higher the classification accuracy. For industries with significant differences in nature, such as manufacturing and finance, the algorithm classification accuracy is very high. For industries with common characteristics, such as manufacturing and transportation, the algorithm classification accuracy is relatively low. Therefore, the method of this invention is more suitable for distinguishing enterprises with significant differences in industry nature.

[0109] The enterprise industry classification method provided in this invention extracts deduplicated independent keywords and enterprise industry classification parameters corresponding to each industry from an enterprise industry standard classification file. The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears across different industries. The number of target independent keywords matching the enterprise to be classified is determined based on whether the enterprise name contains the independent keywords. A matching score is calculated between the enterprise to be classified and each target industry based on the number of target independent keywords, the number of occurrences, the total number of keywords, the number of keywords, and the number of occurrences. The target industry with the highest matching score is determined as the enterprise industry classification of the enterprise to be classified, thus improving the universality of enterprise industry classification.

[0110] Furthermore, the unique, deduplicated keywords extracted from the enterprise industry standard classification document, corresponding to each industry, include:

[0111] Extract the column containing the category name from the enterprise industry standard classification document, and perform preliminary extraction on the text content in the column containing the category name to obtain initial keywords; the above embodiments can be referred to for explanation, and will not be repeated here.

[0112] The initial keywords are then modified to obtain unique, deduplicated keywords corresponding to each industry. This can be referred to the above embodiments for further explanation, and will not be repeated here.

[0113] Furthermore, the step of refining the initial keywords to obtain unique, deduplicated keywords corresponding to each industry includes:

[0114] Remove words and phrases irrelevant to the industry, and correct any incorrectly split words in the remaining words based on the category name corresponding to the column containing the category name; refer to the above embodiments for explanation, and will not be repeated here.

[0115] Correcting incorrectly split words containing conjunctions, splitting the corrected words into keywords of no more than a preset character limit, and deduplicating these keywords to obtain unique, deduplicated keywords corresponding to each industry. Refer to the above implementation example for further details.

[0116] Further, the step of determining the number of target independent keywords for each target industry that matches the enterprise to be classified based on whether the enterprise name contains the independent keyword includes:

[0117] The industry corresponding to at least one independent keyword contained in the company name is taken as the target industry; the above embodiments can be referred to for explanation, and will not be repeated here.

[0118] For each occurrence of a unique keyword contained in the company name within a target industry, the count of target unique keywords for that target industry is incremented by 1, thus obtaining the count of target unique keywords for each target industry matching the company to be classified. This can be referred to the above embodiment for further explanation and will not be repeated here.

[0119] Furthermore, the enterprise industry classification method also includes:

[0120] The industries corresponding to independent keywords that are not included in the company name are directly identified as industry categories unrelated to the company to be classified. This can be referred to the above embodiments for explanation, and will not be repeated here.

[0121] Further, the step of calculating the matching score between the enterprise to be classified and each target industry based on the number of target independent keywords, the frequency of occurrence, the total number of keywords, the number of keywords, and the frequency of occurrence includes:

[0122] The matching score corresponding to each target industry is calculated based on the following expression:

[0123]

[0124] Where w is the number of the target independent keywords, and n is the number of the target independent keywords. i Where N is the number of occurrences, M is the total number of keywords, and k is the number of keywords. i Let be the number of times, where i represents the i-th target independent keyword, and I represents the set of target independent keywords composed of all target independent keywords. Refer to the above embodiments for further explanation; details will not be repeated here.

[0125] Furthermore, after the step of determining the target industry with the highest matching score as the industry classification of the enterprise to be classified, the industry classification method further includes:

[0126] In response to the review operation on the enterprise's industry classification, if the review operation result is determined to be an incorrect classification (first review operation result), a prompt message is generated to optimize the enterprise's industry classification parameters and / or independent keywords. This can be referred to the above embodiment for further explanation and will not be repeated here.

[0127] It should be noted that the enterprise industry classification method provided in this embodiment of the invention can be used in the financial field, or in any technical field other than the financial field. This embodiment of the invention does not limit the application field of the enterprise industry classification method.

[0128] Figure 5 This is a schematic diagram of the structure of an enterprise industry classification device provided in an embodiment of the present invention, as shown below. Figure 5 As shown, the enterprise industry classification device provided in this embodiment of the invention includes an acquisition unit 501, a determination unit 502, and a classification unit 503, wherein:

[0129] The acquisition unit 501 is used to extract, according to the enterprise industry standard classification file, the deduplicated independent keywords and enterprise industry classification parameters corresponding to each industry respectively; wherein, the enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries respectively; the determination unit 502 is used to determine the number of target independent keywords matching the enterprise to be classified for each target industry based on whether the enterprise name of the enterprise to be classified contains the independent keywords; the classification unit 503 is used to calculate the matching score between the enterprise to be classified and each target industry based on the number of target independent keywords, the number of occurrences, the total number of keywords, the number of keywords, and the number of occurrences, and determine the target industry with the highest matching score as the enterprise industry classification of the enterprise to be classified.

[0130] Specifically, the acquisition unit 501 in the device is used to extract deduplicated independent keywords and enterprise industry classification parameters corresponding to each industry according to the enterprise industry standard classification file; wherein, the enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries; the determination unit 502 is used to determine the number of target independent keywords matching the enterprise to be classified for each target industry based on whether the enterprise name of the enterprise to be classified contains the independent keywords; the classification unit 503 is used to calculate the matching score between the enterprise to be classified and each target industry based on the number of target independent keywords, the number of occurrences, the total number of keywords, the number of keywords, and the number of occurrences, and determine the target industry with the highest matching score as the enterprise industry classification of the enterprise to be classified.

[0131] The enterprise industry classification device provided in this embodiment of the invention extracts deduplicated independent keywords and enterprise industry classification parameters corresponding to each industry based on an enterprise industry standard classification file. The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears across different industries. The device determines the number of target independent keywords for each target industry that matches the enterprise to be classified based on whether the enterprise name contains the independent keywords. It calculates the matching score between the enterprise to be classified and each target industry based on the number of target independent keywords, the number of occurrences, the total number of keywords, the number of keywords, and the number of occurrences. The target industry with the highest matching score is determined as the enterprise industry classification of the enterprise to be classified, thereby improving the universality of enterprise industry classification.

[0132] The embodiments of the present invention provide an enterprise industry classification device that can be used to execute the processing flow of the above-described method embodiments. Its functions will not be repeated here, but can be referred to the detailed description of the above-described method embodiments.

[0133] Figure 6 This is a schematic diagram of the physical structure of an electronic device provided in an embodiment of the present invention, such as... Figure 6 As shown, the electronic device includes: a processor 601, a memory 602, and a bus 603;

[0134] The processor 601 and the memory 602 communicate with each other via the bus 603.

[0135] The processor 601 is used to call program instructions in the memory 602 to execute the methods provided in the above-described method embodiments, including, for example:

[0136] Based on the enterprise industry standard classification document, we extract the unique, deduplicated keywords and enterprise industry classification parameters corresponding to each industry.

[0137] The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file corresponding to each industry, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries.

[0138] The number of target independent keywords for each target industry that matches the enterprise to be classified is determined based on whether the enterprise name to be classified contains the independent keyword.

[0139] Based on the number of independent target keywords, the frequency of occurrence, the total number of keywords, the number of keywords, and the frequency, the matching score of the enterprise to be classified with each target industry is calculated, and the target industry with the highest matching score is determined as the enterprise industry classification of the enterprise to be classified.

[0140] This embodiment discloses a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions, and when the program instructions are executed by a computer, the computer can perform the methods provided in the above-described method embodiments, such as:

[0141] Based on the enterprise industry standard classification document, we extract the unique, deduplicated keywords and enterprise industry classification parameters corresponding to each industry.

[0142] The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file corresponding to each industry, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries.

[0143] The number of target independent keywords for each target industry that matches the enterprise to be classified is determined based on whether the enterprise name to be classified contains the independent keyword.

[0144] Based on the number of independent target keywords, the frequency of occurrence, the total number of keywords, the number of keywords, and the frequency, the matching score of the enterprise to be classified with each target industry is calculated, and the target industry with the highest matching score is determined as the enterprise industry classification of the enterprise to be classified.

[0145] This embodiment provides a computer-readable storage medium storing a computer program that causes the computer to execute the methods provided in the above-described method embodiments, including, for example:

[0146] Based on the enterprise industry standard classification document, we extract the unique, deduplicated keywords and enterprise industry classification parameters corresponding to each industry.

[0147] The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file corresponding to each industry, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries.

[0148] The number of target independent keywords for each target industry that matches the enterprise to be classified is determined based on whether the enterprise name to be classified contains the independent keyword.

[0149] Based on the number of independent target keywords, the frequency of occurrence, the total number of keywords, the number of keywords, and the frequency, the matching score of the enterprise to be classified with each target industry is calculated, and the target industry with the highest matching score is determined as the enterprise industry classification of the enterprise to be classified.

[0150] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0151] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0152] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0153] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0154] In the description of this specification, the references to terms such as "an embodiment," "a specific embodiment," "some embodiments," "for example," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0155] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A method for classifying enterprise industries, characterized in that, include: Based on the enterprise industry standard classification document, we extract the unique, deduplicated keywords and enterprise industry classification parameters corresponding to each industry. The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file corresponding to each industry, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries. The number of target independent keywords for each target industry that matches the enterprise to be classified is determined based on whether the enterprise name to be classified contains the independent keyword. Based on the number of independent target keywords, the frequency of occurrence, the total number of keywords, the number of keywords, and the frequency, the matching score of the enterprise to be classified with each target industry is calculated, and the target industry with the highest matching score is determined as the enterprise industry classification of the enterprise to be classified.

2. The enterprise industry classification method according to claim 1, characterized in that, The unique, deduplicated keywords extracted from the enterprise industry standard classification document and corresponding to each industry include: Extract the column containing the category name from the enterprise industry standard classification document, and perform preliminary extraction of the text content in the column containing the category name to obtain initial keywords; The initial keywords are modified to obtain unique, deduplicated keywords corresponding to each industry.

3. The enterprise industry classification method according to claim 2, characterized in that, The process of refining the initial keywords to obtain unique, deduplicated keywords corresponding to each industry includes: Remove words and phrases irrelevant to the industry, and correct any mis-split words in the remaining words based on the category name corresponding to the column containing the category name; Correcting incorrectly split words containing conjunctions, splitting the corrected words into keywords with a maximum of a preset character count, removing duplicates from these keywords, and obtaining unique keywords corresponding to each industry.

4. The enterprise industry classification method according to claim 1, characterized in that, The step of determining the number of target independent keywords for each target industry that matches the enterprise to be classified based on whether the enterprise name contains the independent keyword includes: The industry corresponding to at least one independent keyword contained in the company name shall be the target industry; For each occurrence of an independent keyword contained in the company name within a target industry, the number of target independent keywords for that target industry is incremented by 1, thus obtaining the number of target independent keywords for each target industry that matches the company to be classified.

5. The enterprise industry classification method according to claim 4, characterized in that, The enterprise industry classification method also includes: The industries corresponding to independent keywords that are not included in the company name are directly identified as industry categories unrelated to the company to be classified.

6. The enterprise industry classification method according to claim 1, characterized in that, The process of calculating the matching score between the enterprise to be classified and each target industry based on the number of target independent keywords, the frequency of occurrence, the total number of keywords, the number of keywords, and the frequency of occurrence includes: The matching score corresponding to each target industry is calculated based on the following expression: Where w is the number of the target independent keywords, and n is the number of the target independent keywords. i Where N is the number of occurrences, M is the total number of keywords, and k is the number of keywords. i Let i represent the number of times, i represents the i-th target independent keyword, and I represents the target independent keyword set composed of all target independent keywords.

7. The enterprise industry classification method according to any one of claims 1 to 6, characterized in that, After the step of determining the target industry with the highest matching score as the industry classification of the enterprise to be classified, the industry classification method further includes: In response to the review operation on the enterprise industry classification, if the review operation result is determined to be the first review operation result of incorrect classification, a prompt message is generated to optimize the enterprise industry classification parameters and / or independent keywords.

8. A business industry classification device, characterized in that, include: The acquisition unit is used to extract unique, deduplicated keywords and enterprise industry classification parameters corresponding to each industry based on the enterprise industry standard classification file. The enterprise industry classification parameters include the total number of keywords that appear repeatedly in the enterprise industry standard classification file corresponding to each industry, the number of times each independent keyword appears in the enterprise industry standard classification file, the number of unique keywords in each industry, and the number of times each independent keyword appears in different industries. The determining unit is used to determine the number of target independent keywords for each target industry that matches the enterprise to be classified, based on whether the enterprise name of the enterprise to be classified contains the independent keyword. The classification unit is used to calculate the matching score between the enterprise to be classified and each target industry based on the number of target independent keywords, the number of occurrences, the total number of keywords, the number of keywords, and the number of occurrences, and to determine the target industry with the highest matching score as the enterprise industry classification of the enterprise to be classified.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.