Intelligent enterprise classification method based on keyword optimization
Through Textr and K-means, the industrial keyword thesaurus is generated, and the BERT network model and the inverse document frequency algorithm are combined to optimize enterprise classification. The problems of manual labeling are solved in the existing methods, and the problems of time-consuming and labor-intensive and noisy data sensitivity are achieved, achieving efficient and accurate enterprise classification.
Patent Information
- Application Number
- CN202510461492.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-14
- Publication Date
- 2025-07-18
AI Technical Summary
The existing enterprise classification methods rely on manual labeling of keywords to be time-consuming and labor-intensive and error-prone. Traditional text feature extraction methods ignore word relationships and semantic information. Machine learning models are sensitive to noisy data and are vulnerable to malicious attacks.
The industrial keyword lexicon is automatically generated using the Texttrank algorithm and the K-means algorithm, the correlation between the enterprise and keywords is calculated through word vectors and cosine similarity, multiple data sets are generated using the Bagging algorithm and input into the BERT network model for classification, and the keyword database is optimized with the inverse document frequency algorithm.
It realizes automatic, efficient and accurate enterprise classification, improves keyword extraction efficiency and accuracy, enhances the robustness and long-term effectiveness of the model, and reduces labor costs.
Smart Images

Figure CN120336965A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the cross - field of natural language processing and machine learning, and relates to an enterprise intelligent classification method based on keyword optimization. Background Art
[0002] Traditional enterprise classification methods mainly rely on manual operation, which is inefficient and error - prone. In recent years, with the rapid development of artificial intelligence technology, intelligent classification methods based on machine learning have gradually attracted attention. These methods can automatically extract features from enterprise data and use machine learning algorithms for classification, with higher efficiency and accuracy. However, there are still some problems in existing intelligent classification methods:
[0003] The classification effect of machine learning models depends on the selection of keywords, and existing methods often require manual keyword annotation, which is time - consuming, labor - intensive and error - prone.
[0004] Traditional text feature extraction methods, such as TF - IDF, often ignore the relationship and semantic information between words, resulting in poor classification effects.
[0005] Some machine learning models are sensitive to noise data and are easily affected by malicious attacks or data tampering.
[0006] To solve the above problems, the present invention proposes an enterprise intelligent classification method based on keyword optimization. This method uses the Textrank algorithm and the K - means algorithm to automatically generate an industrial keyword library, and calculates the relevance between enterprises and keywords through word vectors and cosine similarity, so as to achieve automatic, efficient and accurate enterprise classification. Summary of the Invention
[0007] In view of this, the purpose of the present invention is to provide an enterprise intelligent classification method based on keyword optimization.
[0008] To achieve the above - mentioned purpose, the present invention provides the following technical solutions:
[0009] An enterprise intelligent classification method based on keyword optimization, comprising the following steps:
[0010] Step 1: Crawl words related to the enterprise to be classified from the Internet. After word segmentation and stop - word removal, use the TextRank algorithm combined with the K - means algorithm to generate an industrial keyword library; obtain the introduction information of the enterprise to be classified;
[0011] Step 2: Perform word segmentation on the enterprise introduction information and remove stop - words to obtain a word - segmentation result;
[0012] Step 3: Convert the word - segmentation result into word vectors to form candidate enterprise introduction information;
[0013] Step 4: Calculate the preliminary relevance evaluation score based on the candidate enterprise introduction information and the industrial keyword thesaurus.
[0014] Step 5: Generate enterprise keyword introduction information and the comprehensive relevance score according to the preliminary relevance evaluation score.
[0015] Step 6: Generate multiple data sets using the Bagging algorithm and input them into multiple BERT network models for training respectively.
[0016] Step 7: Generate the final enterprise classification result by voting on the output results of the multiple BERT network models.
[0017] Step 8: Update the industrial keyword thesaurus according to the enterprise classification result using the inverse document frequency algorithm.
[0018] Furthermore, in the above Step 4, the method for calculating the preliminary relevance evaluation score includes:
[0019] Calculate the cosine similarity between each word in the candidate enterprise introduction information and each word in the industrial keyword thesaurus.
[0020] Calculate the preliminary relevance evaluation score through maximum similarity matching and a weight function.
[0021] Furthermore, in the above Step 5, the method for generating the enterprise keyword introduction information is: Select the K words with the highest similarity to the industrial keyword thesaurus in the candidate enterprise introduction information and generate the enterprise keyword introduction information in combination with the part-of-speech weight function.
[0022] Furthermore, in the above Step 6, the method for generating multiple data sets includes: Randomly sample the candidate enterprise introduction information to generate multiple independent training sets and input them into the BERT network model for training respectively.
[0023] Furthermore, in the above Step 7, the method for generating the final enterprise classification result is: Vote and count the classification results of the multiple BERT network models, and determine the final classification based on the majority voting result.
[0024] Furthermore, in the above Step 8, the method for updating the industrial keyword thesaurus includes: Calculate the inverse document frequency TF-IDF of each keyword according to the enterprise classification result and replace the keyword group with the lowest weight in the thesaurus.
[0025] Furthermore, the above Step 8 further includes: Compare the updated keyword weights with the high-frequency keywords in the classification result and optimize the keyword replacement strategy through a random forest model.
[0026] Furthermore, the weight function in the above Step 4 includes:
[0027] A part-of-speech weight function that assigns higher weights to nouns, and its calculation formula is:
[0028]
[0029] where f speech (·) represents the part-of-speech weight function, C2 is a weight constant between 0 and 1, and query e [r] represents the r-th word of the e-th word segmentation of the enterprise profile to be classified.
[0030] Furthermore, in the fifth step, the calculation formula for the comprehensive correlation score is:
[0031]
[0032] where correlation(·,·) represents the function for calculating the preliminary correlation evaluation score, preliminary(·,·) represents the function for calculating the preliminary correlation evaluation score, f weight (·) represents the word weight function, f speech (·) represents the part-of-speech weight function, query e [r] represents the r-th word of the e-th word segmentation of the enterprise profile to be classified, and A p,q represents the q-th auxiliary keyword array of the p-th industry; m represents the number of keyword arrays of the p-th industry; L represents the number of industries.
[0033] Furthermore, the method for generating the industry keyword library in the first step includes: performing semantic vector clustering on candidate keywords through a hierarchical clustering algorithm to generate multi-level keyword arrays, and each array corresponds to different sub-categories of the industry.
[0034] The beneficial effects of the present invention are as follows:
[0035] (1) Utilizing the Textrank algorithm and the K-means algorithm to automatically generate the industry keyword library, which avoids the time-consuming and laborious manual annotation of keywords and improves the efficiency and accuracy of keyword extraction.
[0036] (2) Calculating the correlation between the enterprise and the keywords through word vectors and cosine similarity can better capture the semantic information between words and improve the accuracy of enterprise classification.
[0037] (3) Using the Bagging algorithm to generate multiple data sets and inputting them into multiple BERT network models for processing can effectively improve the robustness of the model and reduce the impact of noise data or malicious attacks on the classification results.
[0038] (4) By using the inverse document frequency algorithm to update and iterate keywords, the keyword library can be continuously optimized, improving the long-term effectiveness of enterprise classification.
[0039] (5) This method can automatically and efficiently classify enterprises, improving the efficiency of enterprise classification and reducing labor costs.
[0040] Other advantages, objectives, and features of the present invention will, to some extent, be described in the subsequent specification, and to some extent, will be obvious to those skilled in the art based on an examination of the following text, or can be learned from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following specification. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail and preferably in conjunction with the accompanying drawings, where:
[0042] Figure 1 is a flowchart of the present invention;
[0043] Figure 2 is a schematic diagram of the data structure of the thesaurus;
[0044] Figure 3 is a diagram of the optimal result of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The following illustrates the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention schematically. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0046] Among them, the accompanying drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and should not be construed as a limitation of the present invention; in order to better illustrate the embodiments of the present invention, some components in the accompanying drawings will be omitted, enlarged, or reduced, and do not represent the dimensions of actual products; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the accompanying drawings may be omitted.
[0047] In the accompanying drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, they are based on the orientation or positional relationship shown in the accompanying drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the accompanying drawings are only for illustrative purposes and should not be construed as a limitation of the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0048] Please refer to Figure 1 and Figure 2 , the present invention includes the following steps:
[0049] Step 1
[0050] Scrape relevant words of the enterprise to be classified from the Internet. After word segmentation and removal of stop words, use the TextRank algorithm and the k-mean algorithm to generate an industrial keyword thesaurus.
[0051] Obtain the introduction information of the enterprise to be classified, and generate long text data according to the enterprise introduction information.
[0052] Step 2
[0053] Perform word segmentation on the long text data of the enterprise introduction information and remove stop words.
[0054] Step 3
[0055] Convert the enterprise introduction information after removing stop words into word vectors to form candidate enterprise introduction information. Obtain a preliminary correlation evaluation score according to the candidate enterprise introduction information and keyword information.
[0056] Step 4
[0057] According to the basic correlation score, obtain the enterprise keyword introduction information and the comprehensive correlation score.
[0058] Step 5
[0059] Use the Bagging algorithm to generate multiple data sets and input them into multiple BERT network models for processing.
[0060] Step 6
[0061] Use Bagging to generate the enterprise classification result
[0062] Step 7
[0063] Update and iterate the keywords using the inverse document frequency algorithm based on the enterprise classification result
[0064] For Step 1
[0065] The enterprise to be classified is recorded as: c e
[0066] where c e represents the e-th enterprise to be classified;
[0067] The profile of the enterprise to be classified is recorded as: introduction e ;
[0068] where introduction e represents the profile of the e-th enterprise;
[0069] Tokenize introduction e and remove stop words to obtain the tokenized profile of the enterprise, recorded as: query e = [y e,1 , y e,2 ,..., y e,r ;
[0070] where, introduction e represents the tokenized profile of the e-th enterprise to be classified, query e [r] = y e,r represents the r-th word of the tokenized profile of the e-th enterprise to be classified;
[0071] (1) Obtain the cosine similarity based on the tokenized profile of the enterprise to be classified and the industrial keyword library, recorded as:
[0072]
[0073] where, cossim(·,·) represents the function for calculating cosine similarity, w2v(·,·) represents the function for converting words into word vectors, query e [r] represents the r-th word of the tokenized profile of the e-th enterprise to be classified, A p,q [t] represents the t-th word in the q-th auxiliary keyword array of the p-th industry;
[0074] (2) Calculate the word similarity based on the cosine similarity, recorded as:
[0075] sim(query e [r], A p,q [t]) = cossim(w2v(query e [r]), w2v(A p,q [t]))
[0076] Among them, sim(·, ·) represents the function for calculating the word similarity, cossim(·, ·) represents the function for calculating the cosine similarity, and query e [r] represents the r-th word of the e-th word segmentation of the enterprise profile to be classified, and A p,q [t] represents the t-th word in the q-th auxiliary keyword array of the p-th industry;
[0077] (3) According to the word similarity, calculate the preliminary relevance evaluation score, denoted as:
[0078] preliminary(query e [r], A p,q ) = max(sim(query e [r], A p,q [t])), t ∈ [1, O]
[0079] Among them, preliminary(·, ·) represents the function for calculating the preliminary relevance evaluation score, sim(·, ·) represents the function for calculating the word similarity, and query e [r] represents the r-th word of the e-th word segmentation of the enterprise profile to be classified, and A p,q represents the q-th auxiliary keyword array of the p-th industry; O represents the number of words in the q-th auxiliary keyword array of the p-th industry;
[0080] (4) According to the preliminary relevance evaluation score, calculate the comprehensive relevance evaluation score, denoted as:
[0081]
[0082] Among them, correlation(·, ·) represents the function for calculating the preliminary relevance evaluation score, preliminary(·, ·) represents the function for calculating the preliminary relevance evaluation score, f weight (·) represents the word weight function, f speech (·) represents the part-of-speech weight function, word_n represents noun, word_v represents verb, C2 is a weight constant between 0 and 1, and query e [r] represents the r-th word of the e-th word segmentation of the enterprise profile to be classified, and A p,q represents the q-th auxiliary keyword array of the p-th industry; m represents the number of keyword arrays in the p-th industry; L represents the number of industries.
[0083] (5) Generate enterprise keyword introduction information according to the relevance evaluation score
[0084] Use the e-th enterprise profile word segmentation query eGenerate enterprise keyword introduction information for the K words with the highest similarity to industrial keyword A, denoted as:
[0085] querykey e =[querykey e [1], querykey e [2],..., querykey e [k]]
[0086] Among them, querykey e [1] represents the e-th enterprise profile word segmentation query e The first word with the highest similarity to industrial keyword A, and k represents the number of words taken in order of the highest similarity ranking.
[0087] (6) Feed the enterprise keyword introduction information into BERT for training.
[0088] (7) Optimize the training results using a random forest model.
[0089] (8) Update the industrial keyword library A according to the model classification results.
[0090] Embodiment
[0091] For step zero, the industry is denoted as: Artif p , p ∈ [0, L]
[0092] Among them, among them, Artif p represents the name of the p-th emerging industry, and L represents the number of industries to be classified;
[0093] Scrape relevant description information on the Internet according to the name of the emerging industry, denoted as:
[0094]
[0095] Segment the relevant description information scraped on the Internet and remove stop words to form candidate keywords, denoted as:
[0096]
[0097] Among them, represents the candidate keywords of the p-th emerging industry, and KA p,d represents the d-th candidate keyword of the p-th emerging industry, d ∈ [1, D],
[0098] D represents the number of candidate keywords;
[0099] Use word2vec technology to Map all words in it to a multi-dimensional word vector space, and use the Textrank algorithm and hierarchical clustering algorithm to generate a keyword library, denoted as:
[0100] A p = [A p,1 , A p,2 ,..., A p,m
[0101] A p,q = [A p,q [1], A p,q [2],..., A p,q [t],..., A p,q [O]], q ∈ [1, m], t ∈ [1, O]
[0102] A p,q represents the q-th auxiliary keyword array of the p-th strategic industry; A p,q [t] represents the t-th keyword of the q-th auxiliary keyword array of the p-th strategic industry; m represents the number of keyword arrays of the p-th strategic industry; O represents the number of words in the q-th auxiliary keyword array of the p-th strategic industry; L represents the number of strategic industries.
[0103] For step seven, update the strategic industry keyword library A according to the model classification results
[0104] Use the multi-Bert model to train the enterprise classification results optimized by the Bagging model, and the classification result is Sort, denoted as:
[0105]
[0106] Sort represents the enterprises in the classified industry, represents the j-th enterprise classified into the p-th industry, represents the k-th keyword of the j-th enterprise classified into the p-th industry, L p represents the number of strategic industries, L j represents the number of enterprises classified into the p-th industry, L j,k represents the number of keywords of the j-th enterprise classified into the p-th industry.
[0107] Use the inverse document frequency to calculate the inverse document frequency of each keyword. Regard all keywords of each type of strategic enterprise after classification as a document, and regard the keywords of all types of strategic enterprises as a document set, denoted as:
[0108]
[0109] TF_idf(·,·) represents the function for calculating the inverse document frequency, Total(Soft p represents the total number of keywords of all enterprises classified into the p-th industry, represents the keyword appears in the industry represents the keyword containing the number of industries with strategic industries, L p represents the number of strategic industries;
[0110] Update the weight of the keyword of industry P, denoted as:
[0111]
[0112] A weight p,q =[TF_idf weight (A p,q [1]), TF_idf weight (A p,q [2]),..., TF_idf weight (A p,q [O])]
[0113] A weight p =[A weight p,1 , A weight p,2 ,..., A weight p,m
[0114] where, A p,q [t] represents the t-th keyword in the q-th auxiliary keyword array of the p-th strategic industry, represents the k-th keyword of the j-th enterprise classified into the p-th industry, TF_idf(·,·) represents the function to calculate the inverse document frequency, TF_idf weight (A p,q [t])(·) represents the function to update the weight of the industry keyword;
[0115] Update the keyword of industry P, denoted as:
[0116]
[0117] A p,min =[Sort p [1], Sort p [2],..., Sort p [O]]
[0118] represents the auxiliary phrase with the lowest average updated weight in the p-th industry. Sort p [1] represents the keyword with the highest inverse document frequency among all enterprise keywords classified into the p-th industry, Sort p [2] represents the keyword with the second highest inverse document frequency among all enterprise keywords classified into the p-th industry. Replace the newly obtained auxiliary phrase A p,min with the auxiliary phrase with the lowest average update weight in the p-th industry to obtain the updated keyword library for the p-th industry. Update all industries in sequence to obtain the updated keyword library for the industries.
[0119] Dataset description: 35,292 pieces of data on enterprise names and enterprise profile information were crawled from the Internet. After cleaning and deduplication, 10,871 pieces of valid data were obtained. Manual annotation was carried out, and the annotation content was of the types "high-end motorcycles", "light alloy materials", "light textiles", "biomedicine", "new energy and new energy storage", and "new displays". Traditional machine learning methods and deep neural network models were respectively selected for comparison. The accuracy rate of the decision tree method was 81.48%, the accuracy rate of the CNN model was 90.69%, and the accuracy rate of the RNN model was 90.97%. In this paper, the Bagging method was divided into 3 batches, and the test set and training set were randomly selected each time. The optimal result was 92.60%, as Figure 3 shown.
[0120] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. An enterprise intelligent classification method based on keyword optimization, characterized in that: It includes the following steps: Step 1: Crawl relevant words of the enterprise to be classified from the Internet. After word segmentation and stop word removal, use the TextRank algorithm combined with the K-means algorithm to generate an industrial keyword thesaurus; obtain the introduction information of the enterprise to be classified; Step 2: Segment the enterprise introduction information and remove stop words to obtain a word segmentation result; Step 3: Convert the word segmentation result into word vectors to form candidate enterprise introduction information; Step 4: Based on the candidate enterprise introduction information and the industrial keyword thesaurus, calculate a preliminary relevance evaluation score; Step 5: According to the preliminary relevance evaluation score, generate enterprise keyword introduction information and a comprehensive relevance score; Step 6: Use the Bagging algorithm to generate multiple data sets and input them into multiple BERT network models for training respectively; Step 7: Generate the final enterprise classification result by voting on the output results of the multiple BERT network models; Step 8: According to the enterprise classification result, use the inverse document frequency algorithm to update the industrial keyword thesaurus.
2. The enterprise intelligent classification method based on keyword optimization according to claim 1, wherein: In the above Step 4, the method for calculating the preliminary relevance evaluation score includes: Calculate the cosine similarity between each word in the candidate enterprise introduction information and each word in the industrial keyword thesaurus; Calculate the preliminary relevance evaluation score through maximum similarity matching and a weight function.
3. The enterprise intelligent classification method based on keyword optimization according to claim 1, characterized in that: In the above Step 5, the method for generating enterprise keyword introduction information is: select the K words with the highest similarity to the industrial keyword thesaurus in the candidate enterprise introduction information and generate enterprise keyword introduction information in combination with a part-of-speech weight function.
4. A method for intelligent classification of enterprises based on keyword optimization according to claim 1, characterized in that: In the above Step 6, the method for generating multiple data sets includes: randomly sample the candidate enterprise introduction information to generate multiple independent training sets and input them into the BERT network model for training respectively.
5. A method for enterprise intelligent classification based on keyword optimization according to claim 1, characterized in that: In the above Step 7, the method for generating the final enterprise classification result is: conduct a vote count on the classification results of the multiple BERT network models and determine the final classification based on the majority vote result.
6. The enterprise intelligent classification method based on keyword optimization according to claim 1, characterized in that: In the above Step 8, the method for updating the industrial keyword thesaurus includes: count the inverse document frequency TF-IDF of each keyword according to the enterprise classification result and replace the keyword group with the lowest weight in the thesaurus.
7. An enterprise intelligent classification method based on keyword optimization according to claim 6, characterized in that: The above Step 8 further includes: compare the updated keyword weights with the high-frequency keywords in the classification result and optimize the keyword replacement strategy through a random forest model.
8. A method for intelligent classification of enterprises based on keyword optimization according to claim 1, characterized in that: The weight function in the above Step 4 includes: A part-of-speech weight function that assigns a higher weight to nouns, and the calculation formula is: Among them, f speech (·) represents a part-of-speech weight function, and C2 is a weight constant between 0 and 1. query e [r] represents the r-th word of the e-th word segmentation of the enterprise profile to be classified.
9. The enterprise intelligent classification method based on keyword optimization according to claim 1, characterized in that: In the above Step 5, the calculation formula for the comprehensive relevance score is: Among them, correlation(·,·) represents the function for calculating the preliminary correlation evaluation score, preliminary(·,·) represents the function for calculating the preliminary correlation evaluation score, f weight (·) represents the word weight function, f speech (·) represents the part-of-speech weight function, query e [r] represents the r-th word of the e-th enterprise profile word segmentation to be classified, A p,q represents the q-th auxiliary keyword array of the p-th industry; m represents the number of keyword arrays of the p-th industry; L represents the number of industries.
10. A method for intelligent classification of enterprises based on keyword optimization according to claim 1, characterized in that: The method for generating the industrial keyword thesaurus in the above Step 1 includes: perform semantic vector clustering on candidate keywords through a hierarchical clustering algorithm to generate a multi-level keyword array, and each array corresponds to different sub-categories of the industry.