Bag-of-words similarity analysis-based overseas enterprise attribution industry determination method and system
Through the method based on bag-of-words similarity analysis, overseas enterprises are identified in industry, solving the consistency, accuracy and sustainability of industry identification in the existing technology, and achieving efficient and accurate industry identification.
Patent Information
- Application Number
- CN202411868235.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-16
AI Technical Summary
The existing technology has problems in the identification of language structure differences, industry structure differences, identification consistency problems, identification errors caused by translation, and identification difficulties caused by huge number of companies and industry bases.
Using bag-of-words similarity analysis method, multi-dimensional identification of overseas enterprises and industries is achieved through steps such as obtaining enterprise text information, information preprocessing, feature extraction, text and dictionary similarity calculation, dictionary industry recognition, text and text similarity calculation and enterprise industry recognition.
It improves the accuracy and efficiency of overseas enterprise industry identification, reduces translation risks and costs, solves the consistency and sustainability of industry identification, and enhances the interpretability of identification results.
Smart Images

Figure CN120011561A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to but is not limited to the field of data analysis technology, and in particular relates to a method and system for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis. Background Art
[0002] In the current global economic environment, through in-depth research on overseas companies and industries and the establishment of a global full industrial chain system, which is also an important part of the global economy, it can promote economic development and innovation on a global scale. For enterprises, in order to better expand overseas business, they need to understand the global market and industry trends, so that they can formulate more accurate market strategies and product strategies, improve their own competitiveness, and at the same time, they can also discover new suppliers, customers and partners through understanding the business of overseas companies, and expand business opportunities. The existing methods for identifying overseas companies and industries are mainly to obtain relevant information through existing database queries, news information descriptions, or by translating relevant texts and performing text analysis.
[0003] For general industry identification methods, one technology involving foreign languages (English) is a text-based industry category identification method. By performing Chinese and English information extraction on the target text of the industry to be identified, a set of related Chinese and English text word vectors is obtained, and the length of the text word vector set is determined. Then, based on the length of the text word vector set, the industry category is determined.
[0004] Another type is a text-based industry recognition model, which mainly determines the target vector of each word in the word set of the determined sample text based on the feature vector of each word, which contains the semantic information of the words adjacent to the word, and obtains the word set (in Chinese and English) of the sample text; performs a connection operation on the target vector of each word in the word set of the sample text to obtain the target sample text, and trains the determined basic industry recognition model based on the target sample text to obtain the trained industry recognition model, which is used to analyze the text of the industry to be identified and obtain the industry category that matches the text of the industry to be identified.
[0005] Since English texts, or some texts in other languages, lack a clear language structure, such as using different word forms for different tenses and relatively flexible word order, text analysis is somewhat difficult.
[0006] In view of the above analysis, the technical problems that urgently need to be solved in the existing technology are: the consistency problem of industry identification due to differences in language structure and industry structure of different countries, industry identification errors caused by translation, and the problem that industry identification cannot be performed through a large model due to the large number of enterprises and industry base. Summary of the invention
[0007] In view of the problems existing in the prior art, the present invention provides a method and system for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis.
[0008] The present invention is implemented as follows: a method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis comprises the following steps:
[0009] Step 1: Obtain enterprise text information, obtain enterprise-related introduction texts from enterprise official websites, enterprise information, and sorted enterprise information websites (platforms), etc., and record the total number of texts as N;
[0010] Step 2: Information preprocessing: preprocessing the acquired text information and dictionary;
[0011] Step 3: feature extraction: extracting text features and dictionary features from the preprocessed text information and dictionary;
[0012] Step 4: Calculate the similarity between the text and the dictionary, using the Jaccard algorithm to calculate the overlap between the text word bag E and the industry group word bag D;
[0013] Step 5: dictionary industry identification: multiply the word feature matrix WS with the industry feature vector to obtain the industry similarity DEJacE, and normalize DEJacE to obtain the dictionary industry recognition rate;
[0014] Step 6: Calculate the text similarity between the text to be identified and the enterprise text of the identified industry;
[0015] Step 7: Text industry recognition: output the industries of the identified text word bags EO from high to low according to EEJacE, and the text industry recognition rate is EEJacE;
[0016] Step 8: Enterprise industry identification: Through the above-mentioned dictionary industry identification and text industry identification, conduct multi-dimensional industry identification of overseas enterprises, and finally output the identified industry and industry identification rate.
[0017] Furthermore, the information preprocessing in step 2 specifically includes:
[0018] Text preprocessing: First, the text is divided into sentences using “.” as the position coordinate, and all sentences are numbered, such as sentence 1, sentence 2, sentence 3, etc. All sentences are preliminarily processed by removing punctuation, converting uppercase and lowercase, and removing stop words. Then, the processed sentences are segmented, and each word is stemmed and restored to its word form. The processed word form and the position of the original text where the word is located (sentence number S) are stored.
[0019] Dictionary preprocessing: All industries in the industry dictionary are represented by their positions according to the industry level. For example, Crude oil exploration and exploitation, its industry branch is "Energy-Conventional energy-Oil and gas-Oil and gas exploration and exploitation-Crude oil exploration and exploitation", its industry number is "A01020102", and its level is L5, while Oil and gas exploration and exploitation, its industry number is "A010201", and its level is L4. For each industry name in the industry dictionary, stop words are removed, text is segmented, stems are extracted, and word forms are restored. The processed word form ind, the industry level L of the word, and the number of word components num in the industry are stored.
[0020] Furthermore, the feature extraction in step 3 specifically includes:
[0021] Extract text features and build a text bag of words E. Use bag-of-words to represent the preprocessed text words. The total number of words in the bag of words is M. The repetition value of word w in the document is R (R>0), and n is the number of documents containing word w. Use the TF-IDF algorithm IDF(w)=|ln(R / M)|*(n / N) to calculate the feature value IDF of each word. Build a sentence feature matrix WS for all words in the text in the order of the original sentence. The feature matrix value WS of two adjacent words is 1, and the feature matrix value of non-adjacent words is 0.
[0022] Extract dictionary features, build dictionary word bags, and divide the dictionary into several industry groups according to the inclusion relationship of the industry dictionary level L. For example, the industry group of the first-level industry includes all its sub-level industries, and the industry group of the second-level industry includes all its sub-level industries, until the industry L-1 level. After removing word segmentation, stem extraction, word form restoration and removing duplicate values for each industry group, use bag-of-words to represent it and build the industry group word bag D.
[0023] Furthermore, the calculation of the similarity between the text and the dictionary in step 4 specifically includes:
[0024] First, the intersection inter and union uni of the two word bags are counted, and their ratio inter / uni is calculated, which is recorded as DEJacO. The industry group word bags D with DEJacO=0 are eliminated, and the industry group word bags D are sorted from high to low according to DEJacO to form the industry group to be identified.
[0025] Further, the establishment of the industry feature vector in step five includes:
[0026] For each industry word bag D in the industry group to be recognized, list the industries IND that contain the words in the intersection inter in step four in the industry group in sequence. When an industry IND appears once, its Rp(IND) value is incremented by one, where the initial Rp(IND)=0. For example, if the intersection contains (oil, gas, exploration}, when traversing the industry group according to oil, the Rp(IND) of all industries containing oil is incremented by 1. When traversing the industry group according to exploration, the Rp(IND) of all industries containing exploration is incremented by 1. Finally, calculate the Rp value of each industry IND. For example, Rp(oil exploration)=2.
[0027] Calculate the industry coincidence degree OLR through the formula OLR(IND)=Rp(IND) / num, and calculate the industry feature value according to the four elements of the word feature value IDF, the industry coincidence degree OLR, the word position, and the industry level. Different weights W(S) are given according to the word position. The earlier the statement, the greater the weight, and the weights decrease in sequence according to the statement order; different weights W(L) are given according to different industry levels. The lower the level (the larger L), the greater the weight, and the weights increase in sequence according to the degree of hierarchical subdivision. The industry feature value is IDF(w)*OLR(IND)*W(S)*W(L), and the industry feature vector is established according to the industry feature value.
[0028] Further, step six uses the Jaccard algorithm to calculate the coincidence degree between the text word bag E and the recognized text word bag EO, count the intersection inter and union uni of the two word bags, calculate their ratio inter / uni, denoted as EEJacO. Set a similarity threshold TH for EEJacO, and eliminate the text word bag E0 where EEJacO<TH. Then, the text word bag and the text word bag of the recognized industry are further split into text sub-word bags according to the statement number S, and the statement similarity is calculated in sequence according to the number S, denoted as EESJacO.
[0029] Calculate the industry similarity EEJacE=EESJacO*W(S)*EEJacO according to the three elements of the statement similarity EESJacO, the statement position S, and the text similarity EEJacO.
[0030] Another object of the present invention is to provide an overseas enterprise attributed industry determination system based on bag-of-words similarity analysis for the overseas enterprise attributed industry determination method based on bag-of-words similarity analysis, including:
[0031] The enterprise text information acquisition module obtains enterprise-related introduction texts from the enterprise official website, enterprise information, and sorted enterprise information websites (platforms), etc. The total number of recorded texts is N;
[0032] An information preprocessing module preprocesses the acquired text information and dictionary;
[0033] A feature extraction module performs text feature extraction and dictionary feature extraction on the preprocessed text information and dictionary;
[0034] The module for calculating the similarity between text and dictionary uses the Jaccard algorithm to calculate the overlap between the text word bag E and the industry word bag D.
[0035] The dictionary industry recognition module multiplies the word feature matrix WS with the industry feature vector to obtain the industry similarity DEJacE, and normalizes DEJacE to obtain the dictionary industry recognition rate;
[0036] A text-to-text similarity calculation module calculates the text similarity between the text to be identified and the enterprise text of the identified industry;
[0037] The text industry recognition module outputs the industry of the identified text word bag EO from high to low according to EEJacE, and its text industry recognition rate is EEJacE;
[0038] The enterprise industry identification module conducts multi-dimensional industry identification of overseas enterprises through the above-mentioned dictionary industry identification and text industry identification, and finally outputs the identified industry and industry identification rate.
[0039] Another object of the present invention is to provide a computer device, the computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis.
[0040] Another object of the present invention is to provide a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis.
[0041] Another object of the present invention is to provide an information data processing terminal, which includes the overseas enterprise industry determination system based on bag-of-words similarity analysis.
[0042] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0043] First, the present invention uses the bag-of-words method, through the similarity analysis between multiple word bags, and finally obtains the matching degree of the industry based on the overlap of the word bags, and identifies the industry to which overseas companies belong. The present invention can convert the text into a word vector or a word vector matrix by using a word bag for similarity analysis, reduce the consideration of the text grammar and syntactic structure, eliminate unnecessary word meaning ambiguity, only consider the co-occurrence relationship of words, and can capture the global characteristics of the text. For large-scale text data, the calculation efficiency is very high. At the same time, the word bag model is simple and intuitive, easy to understand and implement. The result of the calculation output based on the overlap of the word bag has strong interpretability, and can avoid the risk of translation bias, and improve the feasibility, accuracy and sustainable update of identifying the industry to which overseas companies belong.
[0044] The present invention establishes word bags of multilingual industry dictionaries and company profiles, and calculates multi-dimensional similarity of word bags based on word bag overlap to obtain industry matching. It focuses on solving the problem of industry identification consistency caused by differences in language structure and industry structure in different countries, avoids industry identification errors caused by translation, avoids the problem of not being able to identify industries through large models due to the large number of companies and industry bases, improves the feasibility and accuracy of overseas enterprise industry identification, reduces the information gap between domestic and foreign companies, and also solves the problem of sustainable identification of companies and industries.
[0045] Second, as auxiliary evidence of the inventiveness of the claims of the present invention, it is also reflected in the following important aspects:
[0046] (1) The expected benefits and commercial value of the technical solution of the present invention after transformation are:
[0047] By accurately and efficiently identifying the industries to which overseas companies belong, this technology can help companies, research institutions and government departments more accurately grasp market dynamics, evaluate corporate cooperation potential and make strategic decisions. In the business field, this technology can quickly reduce international information asymmetry, promote cross-border trade and investment, and promote the process of global economic integration. At the same time, the efficiency and explainability of this technology can be easily integrated into existing enterprise information systems and data analysis platforms, providing customized industry intelligence services for companies in all walks of life and helping to enhance corporate competitiveness. In addition, with the continuous iteration and updating of technology, this technology can continue to adapt to market changes and provide solid guarantees for long-term and stable business value creation.
[0048] (2) The technical solution of the present invention fills the technical gap in the industry at home and abroad:
[0049] The technical solution of the present invention is groundbreaking in the industry and has successfully filled the technical gap in the field of using the bag-of-words method to identify the industry affiliation of overseas companies at home and abroad. Through the unique bag-of-words model construction and overlap analysis, this technology not only overcomes the challenges brought by language structure differences and translation errors, but also significantly improves the processing efficiency of large-scale text data, ensuring the accuracy and sustainability of industry identification. This innovative achievement not only provides strong technical support for related industries, but also opens up new directions for future technological development and application, and is of great milestone significance.
[0050] (3) The technical solution of the present invention solves the technical problems that people have been eager to solve but have never been able to solve successfully:
[0051] The technical solution of the present invention successfully overcomes a technical problem that has long troubled the industry and people have been eager to solve but have never made a breakthrough - how to accurately and efficiently automatically identify the industry to which overseas companies belong, especially in the face of language barriers, cultural differences and the challenges of massive data processing. By introducing the innovative bag-of-words method and bag-of-words overlap analysis, it not only significantly improves the accuracy and efficiency of industry identification, but also reduces the translation risk and cost in the identification process, providing strong technical support for multiple fields such as cross-border cooperation, market research, and policy formulation, and achieving a major leap from theory to practice.
[0052] (4) The technical solution of the present invention overcomes technical prejudice:
[0053] The technical solution of the present invention successfully overcomes technical prejudice and challenges the inherent cognition of industry identification technology in traditional concepts. In the past, people believed that industry identification must rely on complex natural language processing technology or a lot of translation work to solve the obstacles caused by language and cultural differences. However, the present invention uses the bag-of-words method and bag-of-words overlap analysis to prove that efficient and accurate industry identification can be achieved even without relying on complex syntactic analysis and translation. This technical solution not only simplifies the identification process and lowers the technical threshold, but also broadens the application scenarios of industry identification, opening up new paths for technological innovation and development in related fields.
[0054] Third, the existing technical problems solved by the present invention in industrial applications are:
[0055] 1. Difficulty in industry identification due to the complexity of multilingual texts
[0056] In the existing technology, the information of overseas enterprises mostly comes from official websites, news and other text data, which are multilingual and multicultural, and the text features are complex and diverse, making it difficult for traditional methods to handle them uniformly. Traditional industry attribution methods rely on manual annotation or static industry keyword matching, and lack the ability to process multilingual and unstructured texts, resulting in low efficiency and poor accuracy in industry attribution.
[0057] 2. Insufficient and dynamic industry classification dictionaries
[0058] Existing industry classification dictionaries are often unable to adapt to the ever-changing industry needs and technological development due to limited coverage and untimely updates. Static matching of industry keywords cannot dynamically capture changes in corporate business and lacks accurate judgment on the affiliation of emerging industries and mixed industries.
[0059] 3. Inaccuracy in calculating the correlation between text and industry
[0060] In the existing technology, most methods use simple keyword matching, ignoring the implicit contextual relationship and ambiguity in corporate texts. Existing similarity calculation methods are difficult to accurately evaluate the correlation strength between corporate texts and industry characteristics, resulting in a lack of reliability in classification results.
[0061] 4. Unable to conduct multi-dimensional industry attribution verification
[0062] A single industry identification method (such as based on industry dictionaries or simple classification algorithms) cannot verify industry attribution. Current technology is difficult to comprehensively consider the similarity between texts and the relevance between dictionaries and texts, resulting in low accuracy and credibility of industry classification.
[0063] Significant technical advancements of the present invention:
[0064] 1. Multi-dimensional industry attribution mechanism improves identification accuracy
[0065] The present invention combines dictionary industry identification with text industry identification, and adopts the bag-of-words model and the Jaccard similarity algorithm. It can not only dynamically capture the correlation between text features and industry dictionaries, but also analyze the similarity between the text to be identified and the text of the attributed industry, verify the industry affiliation from multiple dimensions, and greatly improve the accuracy and reliability of recognition.
[0066] 2. Dynamically update industry characteristics and improve adaptability
[0067] By extracting features and calculating similarities between enterprise texts and industry dictionaries, the present invention can dynamically identify industry keywords and feature vectors, adapting industry classification methods to the complex attribution issues of emerging industries and cross-industry enterprises, and effectively solving the limitations of industry classification dictionaries.
[0068] 3. Improve unstructured text processing capabilities
[0069] The present invention adopts preprocessing and feature extraction technology (such as removing stop words, unifying the format, etc.), and effectively simplifies the noise and interference in unstructured text through the bag-of-words model, thereby improving the ability to process complex corporate texts and providing an efficient solution for large-scale corporate industry attribution.
[0070] 4. Efficient similarity calculation model
[0071] Through the Jaccard similarity algorithm, the present invention accurately evaluates the overlap between text and industry dictionaries, and between texts, overcoming the neglect of contextual relationships and ambiguity issues in traditional methods, and significantly improving the accuracy of industry attribution results.
[0072] 5. Improve industrial application efficiency and intelligence level
[0073] The multi-dimensional analysis mechanism and automated process of the present invention can automatically and massively attribute enterprises to different industries, thereby reducing labor costs and time consumption, and improving data processing efficiency and intelligence levels in cross-border e-commerce, international trade, market research, and other fields.
[0074] 6. Solve the demand for accurate industry classification in the industry
[0075] The present invention is widely applicable to industry classification needs in the context of international enterprises, providing high-precision data support for market analysis, competitor evaluation, supply chain optimization, etc. of cross-border enterprises, significantly improving the industry classification capabilities and efficiency in the industry.
[0076] The present invention significantly improves the accuracy and adaptability of overseas enterprise industry affiliation, overcomes the difficulties in the prior art in processing unstructured text and multilingual enterprise information, and introduces multi-dimensional similarity analysis to improve the intelligence and reliability of classification, providing efficient and intelligent solutions for international trade, market research and other fields, and has important industrial application value and technological innovation significance. BRIEF DESCRIPTION OF THE DRAWINGS
[0077] Figure 1 It is a flow chart of a method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis provided by an embodiment of the present invention.
[0078] Figure 2 It is a structural diagram of a system for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis provided by an embodiment of the present invention.
[0079] Figure 3 This is a diagram of industry attribution matching results for U.S. stock companies provided by an embodiment of the present invention.
[0080] Figure 4 It is a detailed flow chart of a method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis provided by an embodiment of the present invention.
[0081] Figure 5 Schematic diagram of a word feature matrix provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0082] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0083] like Figure 1 As shown, a method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis provided by an embodiment of the present invention includes the following steps:
[0084] S101, obtaining enterprise text information, obtaining enterprise-related introduction texts from the enterprise official website, enterprise information, and sorted enterprise information websites, and the total number of recorded texts is N;
[0085] S102, information preprocessing, preprocessing the acquired text information and dictionary;
[0086] S103, feature extraction, performing text feature extraction and dictionary feature extraction on the preprocessed text information and dictionary;
[0087] S104, calculating the similarity between the text and the dictionary, using the Jaccard algorithm to calculate the overlap between the text word bag E and the industry group word bag D;
[0088] S105, dictionary industry recognition, multiplying the word feature matrix WS with the industry feature vector to obtain the industry similarity DEJacE, normalizing DEJacE to obtain the dictionary industry recognition rate;
[0089] S106, text-to-text similarity calculation, performing text similarity calculation between the text to be recognized and the enterprise text of the recognized industry;
[0090] S107, text industry recognition, output the industry of the identified text word bag EO from high to low according to EEJacE, and its text industry recognition rate is EEJacE;
[0091] S108, enterprise industry identification, through the above dictionary industry identification and text industry identification, conduct multi-dimensional industry identification of overseas enterprises, and finally output the identification industry and industry identification rate.
[0092] Collect relevant text information of the enterprise from the official website, enterprise information or sorted enterprise information website, covering the business description, product information and operation model of the enterprise. By recording the total number of enterprise introduction texts N, the foundation for subsequent industry attribution analysis is laid. These text data will serve as the original input for subsequent text processing and analysis.
[0093] The collected text information is preprocessed, including removing stop words, punctuation marks, and non-business related content, unifying the text format, and ensuring the standardization of words. At the same time, the industry classification dictionary used for analysis is also preprocessed to clean up redundant content and build a standardized industry keyword set. Through cleaning and standardization, the accuracy of text feature extraction and similarity calculation is improved.
[0094] Feature extraction is performed on the preprocessed text information and industry dictionary respectively. The bag-of-words model is used to convert the text and dictionary into word frequency vectors. Text feature extraction generates text word bags, and industry dictionary extraction generates industry group word bags. This bag-of-words representation simplifies the text semantic relationship while retaining the keyword distribution characteristics of the text.
[0095] Based on the bag-of-words model, the Jaccard algorithm is used to calculate the overlap between corporate text and industry dictionaries. The Jaccard algorithm quantifies the similarity between two sets by measuring the ratio of the intersection to the union.
[0096] The calculation results represent the similarity distribution between the enterprise text and the dictionaries of various industries, providing a preliminary similarity index for subsequent industry identification.
[0097] The feature matrix established by the text bag of words and the industry bag of words is multiplied by the industry feature vector to further calculate the industry similarity. The calculation results are normalized and adjusted to the range of [0,1] to obtain the dictionary industry recognition rate. Through this step, the strength of the association between the enterprise and the industries in the dictionary can be preliminarily identified.
[0098] The text similarity between the enterprise text to be identified and the enterprise text of the identified industry is calculated using the bag-of-words model and Jaccard algorithm. By analyzing the overlap between the text to be identified and the existing industry annotated text, the enterprise industry attribution judgment is further verified and optimized. The similarity calculation results provide an important reference for enterprise text industry identification.
[0099] Combining the industry identification results of the dictionary and the industry identification results of the text, through multi-dimensional similarity analysis, the industry affiliation of the enterprise is comprehensively judged, and the final industry identification results and industry recognition rate are output. This method not only utilizes the structured characteristics of the industry dictionary, but also combines the actual text content of the annotated industry to achieve accurate identification of the industry affiliation of overseas enterprises.
[0100] like Figure 4 As shown, an embodiment of the present invention provides a method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis, comprising the following steps:
[0101] Step 1: Obtain enterprise text information, obtain enterprise-related introduction texts from enterprise official websites, enterprise information, and sorted enterprise information websites (platforms), etc., and record the total number of texts as N;
[0102] Step 2: Information preprocessing: preprocessing the acquired text information and dictionary; specifically including:
[0103] Text preprocessing: First, the text is divided into sentences using “.” as the position coordinate, and all sentences are numbered, such as sentence 1, sentence 2, sentence 3, etc. All sentences are preliminarily processed by removing punctuation, converting uppercase and lowercase, and removing stop words. Then, the processed sentences are segmented, and each word is stemmed and restored to its word form. The processed word form and the position of the original text where the word is located (sentence number S) are stored.
[0104] Dictionary preprocessing: All industries in the industry dictionary are represented by their positions according to the industry level. For example, Crude oil exploration and exploitation, its industry branch is "Energy-Conventional energy-Oil and gas-Oil and gas exploration and exploitation-Crude oil exploration and exploitation", its industry number is "A01020102", and its level is L5, while Oil and gas exploration and exploitation, its industry number is "A010201", and its level is L4. For each industry name in the industry dictionary, stop words are removed, text is segmented, stems are extracted, and word forms are restored. The processed word form ind, the industry level L of the word, and the number of word components num in the industry are stored.
[0105] Step 3, feature extraction, performs text feature extraction and dictionary feature extraction on the preprocessed text information and dictionary; specifically includes:
[0106] Extract text features and build a text bag of words E. Use bag-of-words to represent the preprocessed text words. The total number of words in the bag of words is M. The repetition value of word w in the document is R (R>0), and n is the number of documents containing word w. Use the TF-IDF algorithm IDF(w)=|ln(R / M)|*(n / N) to calculate the feature value IDF of each word. Build a sentence feature matrix WS for all words in the text in the order of the original sentence. The feature matrix value WS of two adjacent words is 1, and the feature matrix value of non-adjacent words is 0.
[0107] Extract dictionary features, build dictionary word bags, and divide the dictionary into several industry groups according to the inclusion relationship of the industry dictionary level L. For example, the industry group of the first-level industry includes all its sub-level industries, and the industry group of the second-level industry includes all its sub-level industries, until the industry L-1 level. After removing word segmentation, stem extraction, word form restoration and removing duplicate values for each industry group, use bag-of-words to represent it and build the industry group word bag D.
[0108] Step 4: Calculate the similarity between the text and the dictionary. Use the Jaccard algorithm to calculate the overlap between the text word bag E and the industry word bag D. Specifically, it includes:
[0109] First, the intersection inter and union uni of the two word bags are counted, and their ratio inter / uni is calculated, which is recorded as DEJacO. The industry group word bags D with DEJacO=0 are eliminated, and the industry group word bags D are sorted from high to low according to DEJacO to form the industry group to be identified.
[0110] Step 5: dictionary industry identification, including:
[0111] For each industry group word bag D in the industry group to be identified, list the industry IND containing the word in the industry group in turn according to the words in the intersection inter in step 4. When the industry IND appears once, its Rp(IND) value is increased by one, where the initial Rp(IND) = 0. For example, the intersection contains (oil, gas, exploration}. When traversing the industry group according to oil, the repetition value of all industries containing oil is Rp(IND)+1. When traversing the industry group according to exploration, the repetition value of all industries containing exploration is Rp(IND)+1. Finally, the Rp value of each industry IND is calculated, for example, Rp(oil exploration) = 2.
[0112] Calculate the industry overlap rate OLR through the formula OLR(IND)=Rp(IND) / num, and calculate the industry eigenvalue based on four factors: the word eigenvalue IDF, the industry overlap rate OLR, the position of the word, and the industry level. Different weights W(S) are given according to the position of the word. The earlier the statement, the greater the weight, and the weights decrease in sequence according to the statement order; different weights W(L) are given according to different industry levels. The lower the level (the larger L), the greater the weight, and the weights increase in sequence according to the increase in the level subdivision degree. The industry eigenvalue is IDF(w)*OLR(IND)*W(S)*W(L), and an industry eigenvector is established based on the industry eigenvalue.
[0113] Multiply the word and sentence feature matrix WS by the industry eigenvector to obtain the industry similarity DEJacE, and perform normalization processing on DEJacE to obtain the dictionary industry recognition rate.
[0114] Step 6, text and text similarity calculation, calculate the text similarity between the text to be recognized and the enterprise text of the recognized industry; specifically including:
[0115] Use the Jaccard algorithm to calculate the overlap degree of the text bag of words E and the text bag of words EO of the recognized text, count the intersection inter and union uni of the two text bags of words, calculate their ratio inter / uni, and denote it as EEJacO. Set a similarity threshold TH for EEJacO, and剔除 the text bag of words E0 with EEJacO < TH. Then, split the text bag of words and the text bag of words of the recognized industry into text sub-bags according to the sentence number S, and calculate the sentence similarity in sequence according to the number S, denoted as EESJacO.
[0116] Calculate the industry similarity EEJacE = EESJacO * W(S) * EEJacO according to the three factors of the sentence similarity EESJacO, the sentence position S, and the text similarity EEJacO.
[0117] Step 7, text industry recognition, output the industries where the recognized text bag of words EO is located from high to low according to EEJacE, and its text industry recognition rate is EEJacE.
[0118] Step 8, enterprise industry recognition, through the above dictionary industry recognition and text industry recognition, perform multi-dimensional identification of the industries of overseas enterprises, and finally output the identified industry and the industry recognition rate.
[0119] As Figure 2 shown, the embodiment of the present invention provides a system for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis, including:
[0120] The enterprise text information acquisition module obtains enterprise-related introduction texts from the enterprise official website, enterprise information, and sorted enterprise information websites (platforms), etc. The total number of recorded texts is N;
[0121] An information preprocessing module preprocesses the acquired text information and dictionary;
[0122] A feature extraction module performs text feature extraction and dictionary feature extraction on the preprocessed text information and dictionary;
[0123] The module for calculating the similarity between text and dictionary uses the Jaccard algorithm to calculate the overlap between the text word bag E and the industry word bag D.
[0124] The dictionary industry recognition module multiplies the word feature matrix WS with the industry feature vector to obtain the industry similarity DEJacE, and normalizes DEJacE to obtain the dictionary industry recognition rate;
[0125] A text-to-text similarity calculation module calculates the text similarity between the text to be identified and the enterprise text of the identified industry;
[0126] The text industry recognition module outputs the industry of the identified text word bag EO from high to low according to EEJacE, and its text industry recognition rate is EEJacE;
[0127] The enterprise industry identification module conducts multi-dimensional industry identification of overseas enterprises through the above-mentioned dictionary industry identification and text industry identification, and finally outputs the identified industry and industry identification rate.
[0128] Example 1
[0129] A method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis comprises the following steps:
[0130] Step 1: Get the text information of Enterprise A. The total number of recorded texts is 3, including 2 in English and 1 in Chinese.
[0131] Step 2: preprocess the acquired text information and dictionary.
[0132] The text (A, together with its subsidiaries, provides professional infrastructure consulting services worldwide…) is divided into sentences. Text 1 has 6 sentences in total. The text is preprocessed by removing punctuation marks, converting uppercase and lowercase letters, removing stop words, text segmentation, stemming, and word form restoration. The processed word forms and the positions of the original text where the words are located (sentence number S) are stored, for example (infrastructure, 1; consult, 1; plan, 3…)
[0133] Dictionary preprocessing: All industries in the industry dictionary are represented by their positions according to the industry level. For example, Crude oil exploration and exploitation, its industry branch is "Energy-Conventional energy-Oil and gas-Oil and gas exploration and exploitation-Crude oil exploration and exploitation", its industry number is "A01020102", and its level is L5, while Oil and gas exploration and exploitation, its industry number is "A010201", and its level is L4. For each industry name (Oil and gas exploration and exploitation) in the industry dictionary, stop words are removed, text is segmented, stems are extracted, and word forms are restored. The processed word form ind, the industry level L of the word, and the number of word components num in the industry are stored, for example (Oil, 4, 4).
[0134] Step three: feature extraction. Feature extraction includes text feature extraction and dictionary feature extraction.
[0135] Extract text features and build a word bag E for the text. Use the word bag to represent text 1. The total number of words M = 70, the repetition value of the word infrastructure in the document R = 2, and the number of documents containing the word infrastructure n is 3. Use the TF-IDF algorithm to calculate its feature value IDF (infrastructure) = |ln (R / M) |* (n / N) = 3.56. Build a word feature matrix WS for all the words in the text in the original sentence order. The feature matrix value WS of two adjacent words is 1, and the feature matrix value of non-adjacent words is 0, such as Figure 5 shown.
[0136] Extract dictionary features, build dictionary word bags, divide the dictionary into several industry groups according to the inclusion relationship of the industry dictionary level L, remove word segmentation, stem extraction, word form restoration and remove duplicate values for each industry group, and use word bags to represent them to build industry group word bags D. For example, the industry group where the first-level industry (Energy) is located includes all its sub-industries (energy, conventional, oil, gas, exploration...), and the industry group where the third-level industry (Oil and gas) is located includes all its sub-industries (conventional, oil, gas, exploration...).
[0137] Step 4: Calculate the similarity between the text and the dictionary. Use the Jaccard algorithm to calculate the similarity between the text word bag E and the industry group word bag D. First, count the intersection inter and union uni of the two word bags. For example, the intersection inter (infrastructure, consult, service...), the union uni (infrastructure, consult, service, build, design...), calculate the similarity DEJacO = 0.24. Eliminate the industry group word bag D with DEJacO = 0, and sort the industry group word bag D from high to low according to DEJacO to form the industry group to be identified.
[0138] Step 5, dictionary industry identification, for each industry group word bag D in the industry group to be identified, according to the word (infrastructure) in the intersection inter in step 4, list the industry IND (Infrastructure Services, Infrastructure Consult, Infrastructure Agricultural equipment, Infrastructure Consult Service...) containing the word in the industry group in turn, and calculate the Rp (IND) value of each industry IND. Rp (Infrastructure Services) = 2, Rp (Infrastructure Agricultural equipment) = 1, Rp (Infrastructure Consult Service) = 3.
[0139] The industry overlap OLR is calculated by the formula OLR(IND)=Rp(IND) / num, OLR(Infrastructure Consult Service)=1.OLR(Infrastructure Agricultural equipment)=0.33. The industry feature value is calculated based on the four factors of word feature value IDF, industry overlap OLR, word position, and industry level, and the industry feature vector is established. Then, it is multiplied with the word feature matrix WS to obtain the industry similarity DEJacE. DEJacE is normalized to obtain the dictionary industry recognition rate (DEJacE(Infrastructure Consult Service)=0.92, DEJacE(Infrastructure Agricultural equipment)=0.21...}.
[0140] Step 6: Calculate the text-to-text similarity. Calculate the text similarity between the text to be identified of enterprise B (B, provides Railway, highway, transportation consult services...) and the text of the enterprise in the identified industry (A, together with its subsidiaries, provides professional infrastructure consulting services...). The intersection of the two word bags is (provide, consult, service...}, and the union is (infrastructure, provide, manage, service, facility, consult...}. The similarity is calculated as EEJacO=0.74. Then, the two text word bags are further split into S text sub-word bags according to the sentence number S. The sentence similarity is calculated in sequence according to the number S, EEJacO(S=1)=0.64, EEJacO(S=1)=0.47. The industry similarity EEJacE=0.87 is calculated based on the three elements of sentence similarity EEJacO, sentence position S, and text similarity EEJacO.
[0141] Step 7: Text industry identification. According to the EEJacE in step 6, the industries of the identified text word bag EO are output from high to low (Infrastructure Consult Service, 0.78; financial consultancy service, 0.16...).
[0142] Step 8: Enterprise industry identification: Through the above-mentioned dictionary industry identification and text industry identification, the overseas enterprises are identified in multiple dimensions, and the identified industries and industry identification rates are finally output (A, Infrastructure Consult Service, 0.92; B, Infrastructure Consult Service, 0.78…).
[0143] This paper proposes a method for determining the industry to which overseas enterprises belong based on bag-of-words similarity analysis, aiming to accurately identify the industry to which an enterprise belongs through in-depth analysis of enterprise text information. This method comprehensively solves the problems of the inconvenience of traditional manual force measurement, the difficulty of internal force monitoring, and the limitation of sensor layout through systematic steps, including text acquisition, preprocessing, feature extraction, similarity calculation, and industry identification.
[0144] First, the first step of the method is to obtain the text information of enterprise A. In practical applications, the text information of enterprises comes from various sources such as official websites, annual reports, and news releases. Taking enterprise A as an example, three texts were obtained, two of which were in English and the other in Chinese. These text information constitute the basic data for subsequent analysis and provide rich corpus resources for determining industry affiliation.
[0145] Next, in step 2, the acquired text information and industry dictionary are preprocessed. Taking text 1 as an example, it is split into six independent sentences through sentence processing, followed by a series of preprocessing steps such as punctuation removal, capitalization unification, stop word removal, word segmentation, stem extraction, and word form restoration. The processed words and their positions in the original text are recorded. At the same time, similar preprocessing is performed on each industry name in the industry dictionary, and its hierarchical position and word composition are recorded to lay the foundation for subsequent feature matching.
[0146] In step three, feature extraction is performed, which is divided into two parts: text feature extraction and dictionary feature extraction. Text feature extraction establishes a text word bag E, uses the TF-IDF algorithm to calculate the feature value of each word, and constructs a word feature matrix WS to represent the distribution and association of words in the text. Dictionary feature extraction establishes a dictionary word bag D, divides the industry dictionary into several industry groups according to the hierarchical relationship, and performs word segmentation and deduplication processing on each industry group to form a standardized industry word bag, providing a unified comparison benchmark for similarity calculation.
[0147] In step 4, the Jaccard similarity algorithm is used to calculate the similarity between the text word bag E and each industry group word bag D. By counting the number of words in the intersection and union, the similarity score DEJacO is calculated. The industry group word bags D with zero similarity are eliminated, and the remaining industry group word bags D are sorted from high to low according to the similarity score, and the industry group that is most consistent with the enterprise text is screened out to form a list of industry groups to be identified.
[0148] Then, in step five, dictionary industry identification is performed. For each industry word bag D to be identified, the specific industry IND containing the word is listed based on the intersection words with the text word bag E, and the Rp value of each industry IND is calculated, that is, the number of occurrences of a specific word in the industry. The industry overlap OLR is calculated using the formula OLR(IND)=Rp(IND) / num to reflect the degree of matching of each industry on specific words. Finally, the industry feature vector is established by combining the four elements of word feature value IDF, industry overlap OLR, word position and industry level, and multiplied by the word feature matrix WS to obtain the industry similarity DEJacE, which is then normalized to obtain the recognition rate of each industry.
[0149] In step 6, the similarity between texts is calculated. Taking the text to be identified of enterprise B and the text of enterprise A in the identified industry as an example, the Jaccard similarity EEJacO of the two text bags of words is calculated, and they are further split into sub-bags of words according to the sentence number, and the sentence similarity EESJacO is calculated respectively. Combining the sentence similarity, sentence position and text similarity, the industry similarity EEJacE is calculated, and through normalization processing, the recognition rate of enterprise B in each industry is obtained, which provides a basis for the final industry attribution.
[0150] Steps 7 and 8 respectively complete the text industry identification and enterprise industry identification. Through the comprehensive analysis of steps 5 and 6, the text information of the enterprise to be identified is multi-dimensionally identified, and the final industry affiliation and its corresponding recognition rate are output. For example, enterprise A is identified as "Infrastructure Consult Service" with a recognition rate of 0.92; enterprise B is also identified as "Infrastructure Consult Service" with a recognition rate of 0.78, etc. This method achieves accurate identification of the industry affiliation of overseas enterprises through multi-dimensional analysis and comprehensive evaluation, and improves the automation and intelligence level of industry classification.
[0151] The present invention also includes implementation schemes of a computer device, a computer-readable storage medium, and an information data processing terminal. The computer device executes a pre-stored computer program through a memory and a processor to implement each step of the industry attribution determination method based on bag-of-words similarity analysis. The computer-readable storage medium stores the above program, and the information data processing terminal integrates all functional modules of the method to form a complete industry attribution determination system, which is convenient for rapid deployment and efficient operation in practical applications.
[0152] In summary, the present invention solves the shortcomings of traditional industry attribution determination methods in terms of accuracy, efficiency and multi-dimensional monitoring through systematic steps and innovative technical means, and realizes precise and intelligent analysis of the industry attribution of overseas companies. It has broad application prospects and significant technical advantages.
[0153] Relevant evidence of the technical effects achieved by the embodiments of the present invention.
[0154] 1. The industry recognition rate of overseas enterprises has been improved:
[0155] By using the present invention, the degree of automation in the industry classification of overseas enterprises has been greatly improved, achieving efficient and accurate classification of more enterprises, and significantly increasing the number of overseas enterprises that can be identified by industry through automated classification services. Figure 3 )
[0156] 2. Improvement of industry recognition accuracy:
[0157] After actual testing, the technical solution of the present invention has shown a high recognition rate and accuracy in identifying the industries to which overseas companies belong. Compared with traditional methods, this technology can avoid the deviation caused by language translation when processing multilingual texts and complex industry classifications, thereby classifying them into the correct industry.
[0158] (As shown in Table 1)
[0159] Table 1 Industry classification results of some US stock companies
[0160]
[0161] 3. Interpretability of results:
[0162] The output results of the present invention are highly interpretable, that is, they can clearly explain why a certain enterprise is classified into a certain industry, which helps users understand and trust the recognition results and improves the accuracy and reliability of decision-making.
[0163] 4. Dictionary expandability:
[0164] The technical solution of the present invention can utilize a flexible design architecture to perform accurate industry identification when constructing a multilingual industry dictionary, so that the dictionary can be automatically updated and expanded, greatly broadening the application scope of the dictionary and making it more global and universal.
[0165] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. It can be understood by a person of ordinary skill in the art that the above-mentioned devices and methods can be implemented using computer executable instructions and / or contained in a processor control code, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the carrier medium. The device and its modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, and can also be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0166] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis, characterized in that: The following steps are involved: Step 1: Obtain enterprise text information, obtain enterprise-related introduction texts from enterprise official websites, enterprise information, and sorted enterprise information websites, and record the total number of texts as N; Step 2: Information preprocessing: preprocessing the acquired text information and dictionary; Step 3: feature extraction: extracting text features and dictionary features from the preprocessed text information and dictionary; Step 4: Calculate the similarity between the text and the dictionary, using the Jaccard algorithm to calculate the overlap between the text word bag E and the industry group word bag D; Step 5: dictionary industry identification: multiply the word feature matrix WS with the industry feature vector to obtain the industry similarity DEJacE, normalize DEJacE, and obtain the dictionary industry recognition rate; Step 6: Calculate the text similarity between the text to be identified and the enterprise text of the identified industry; Step 7: Text industry recognition: output the industries of the identified text word bags EO from high to low according to EEJacE, and the text industry recognition rate is EEJacE; Step 8: Enterprise industry identification: Through the above-mentioned dictionary industry identification and text industry identification, conduct multi-dimensional industry identification of overseas enterprises, and finally output the identified industry and industry identification rate.
2. The method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis according to claim 1, characterized in that: The information preprocessing in step 2 specifically includes: Text preprocessing: first, the text is divided into sentences using "." as the position coordinate, and all sentences are numbered. All sentences are preliminarily processed by removing punctuation marks, converting uppercase and lowercase letters, and removing stop words. Then, the processed sentences are segmented, and each word is stemmed and restored. The processed word form and the position of the original text where the word is located, that is, the sentence number S, are stored. Dictionary preprocessing: All industries in the industry dictionary are represented according to the industry level. Stop words are removed, text is segmented, stems are extracted, and word forms are restored for each industry name in the industry dictionary. The processed word form ind, the industry level L where the word is located, and the number of word components num in the industry are stored.
3. The method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis according to claim 1, characterized in that: The feature extraction in step 3 specifically includes: Extract text features and establish a text word bag E. Use bag-of-words to represent the preprocessed text words. The total number of words in the word bag is M. The repetition value of word w in the document is R (R>0), and n is the number of documents containing word w. Use the TF-IDF algorithm IDF(w)=|ln(R / M)|*(n / N) to calculate the feature value IDF of each word; establish a sentence feature matrix WS for all words in the text according to the original sentence order. The feature matrix value WS of two adjacent words is 1, and the feature matrix value of non-adjacent words is 0; Extract dictionary features, build a dictionary bag-of-words, and divide the dictionary into several industry groups according to the inclusion relationship of the industry dictionary hierarchy L. The industry group where the first-level industry is located includes all its sub-level industries, the industry group where the second-level industry is located includes all its sub-level industries, and so on until the industry L-1 level. After removing word segmentation, stemming, lemmatization, and duplicate values for each industry group, it is represented using bag-of-words to build the industry group bag-of-words D.
4. The method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis according to claim 1, characterized in that: The calculation of the similarity between the text and the dictionary in step four specifically includes: First, count the intersection inter and union uni of the two bags-of-words, calculate their ratio inter / uni, denoted as DEJacO. Then, remove the industry group bag-of-words D with DEJacO = 0, and sort the industry group bag-of-words D from high to low according to DEJacO to form the industry groups to be recognized.
5. The method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis according to claim 1, characterized in that: The establishment of the industry feature vector in step five includes: For each industry group bag-of-words D in the industry groups to be recognized, list the industries IND that contain the word in the industry group in turn according to the words in the intersection inter in step four. When the industry IND appears once, its Rp(IND) value is incremented by one, where the initial Rp(IND) = 0. Calculate the industry overlap OLR through the formula OLR(IND) = Rp(IND) / num, and calculate the industry feature value according to the four elements of the word feature value IDF, industry overlap OLR, word position, and industry level. Different weights W(S) are given according to the word position. The earlier the statement, the greater the weight, and the weights decrease in order according to the statement order. Different weights W(L) are given according to different industry levels. The lower the level (the larger L), the greater the weight, and the weights increase in order according to the level of subdivision. The industry feature value is IDF(w)*OLR(IND)*W(S)*W(L), and the industry feature vector is established according to the industry feature value.
6. The method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis according to claim 1, characterized in that: In step six, use the Jaccard algorithm to calculate the overlap between the text bag-of-words E and the recognized text bag-of-words EO. Count the intersection inter and union uni of the two bags-of-words, calculate their ratio inter / uni, denoted as EEJacO. Set a similarity threshold TH for EEJacO, and remove the text bag-of-words E0 with EEJacO < TH. Then, split the text bag-of-words and the text bag-of-words of the recognized industry into text sub-bags according to the statement number S, and calculate the statement similarity in turn according to the number S, denoted as EESJacO. Calculate the industry similarity EEJacE = EESJacO*W(S)*EEJacO according to the three elements of the statement similarity EESJacO, statement position S, and text similarity EEJacO.
7. A system for determining the industry to which overseas enterprises belong based on bag-of-words similarity analysis according to any one of claims 1 to 6, characterized in that: Including: Enterprise text information acquisition module, which obtains the enterprise-related introduction text from the enterprise official website, enterprise news, and organized enterprise information websites, and records the total number of texts as N. Information preprocessing module, which preprocesses the obtained text information and the dictionary. Feature extraction module, which extracts text features and dictionary features from the preprocessed text information and dictionary. The module for calculating the similarity between text and dictionary uses the Jaccard algorithm to calculate the overlap between the text word bag E and the industry word bag D. The dictionary industry recognition module multiplies the word feature matrix WS with the industry feature vector to obtain the industry similarity DEJacE, normalizes DEJacE, and obtains the dictionary industry recognition rate; A text-to-text similarity calculation module calculates the text similarity between the text to be identified and the enterprise text of the identified industry; The text industry recognition module outputs the industry of the identified text word bag EO from high to low according to EEJacE, and its text industry recognition rate is EEJacE; The enterprise industry identification module conducts multi-dimensional industry identification of overseas enterprises through the above-mentioned dictionary industry identification and text industry identification, and finally outputs the identified industry and industry identification rate.
8. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and when the computer program is executed by the processor, the processor executes the steps of the method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis as described in any one of claims 1 to 6.
9. A computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the processor executes the steps of the method for determining the industry to which an overseas enterprise belongs based on bag-of-words similarity analysis as described in any one of claims 1 to 6.
10. An information data processing terminal, characterized in that: The information data processing terminal includes the system for determining the industry to which overseas enterprises belong based on bag-of-words similarity analysis as described in claim 7.