Multi-language industry dictionary establishing method and system based on large model
By introducing large-model technology in the construction of industry dictionary, using breadth search and deep search to identify industry node relationships, calculate similarity and inclusion, the problems of low efficiency, low accuracy, and insufficient multilingual support in the existing technology are solved, and efficient, accurate, and multilingual support industry dictionary construction and dynamic updates are achieved.
Patent Information
- Application Number
- CN202411868078.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-18
- Publication Date
- 2025-05-16
AI Technical Summary
There are problems in the construction process of existing industry dictionaries, which are time-consuming, inefficient, low accuracy, incomplete dictionary coverage, and difficult to unify multilingual support.
The multilingual industry dictionary establishment method based on large models is adopted, and breadth search and in-depth search are carried out through large models, inclusion, peer level, subordinate relationships between industry nodes are identified, industry similarity and inclusion degree are calculated, industry dictionaries are built with multilingual support, and industry dictionaries are verified and updated regularly through enterprise and market.
It improves the efficiency and accuracy of industry dictionary construction, enhances the coverage and multilingual support capabilities of dictionaries, realizes independent updates and maintenance of dictionaries, and adapts to industry changes and language evolution.
Smart Images

Figure CN120012767A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to but is not limited to the technical field of industry classification dictionaries, and in particular relates to a method and system for establishing a multilingual industry dictionary based on a large model. Background Art
[0002] Industry classification dictionaries are usually based on a complete set of industry classification standards or systems. This set of standards covers various industry fields, including manufacturing, service, finance, etc. Industry classification dictionaries are a professional tool for enterprises to identify their industries, aiming to help enterprises clarify the industry field they belong to, so as to better position and conduct business. Industry classification dictionaries usually contain classification standards and classification vocabulary for multiple industries. These vocabulary and classification standards can be used to guide enterprises to segment their businesses and markets, helping enterprises to more accurately understand the industry environment and competitive situation in which they are located. Therefore, it is not only an important tool for enterprises to conduct market analysis and research, but also an important reference for enterprises to formulate strategies and make decisions.
[0003] The construction of industry classification dictionaries usually adopts rule-based or statistical-based methods. Among them, the rule-based method mainly builds the vocabulary and classification rules of the dictionary manually according to the industry classification standards; while the statistical-based method uses a large amount of text data to automatically extract vocabulary features and classification information through machine learning and natural language processing technology.
[0004] At present, the construction of industry classification dictionaries requires the use of natural language processing technology, including lexical analysis, syntactic analysis, semantic understanding, etc. These technologies can help extract key information from texts, understand the semantics and context of words and phrases, and thus accurately classify industries. Big data and cloud computing capabilities are also required. The construction of industry classification dictionaries requires processing large amounts of data and performing complex calculations. Big data technology and cloud computing platforms can provide powerful storage and computing capabilities, support the analysis and processing of massive data, and thus improve the efficiency and accuracy of industry classification. At the same time, the construction of industry classification dictionaries also requires the combination of domain knowledge and expert guidance. The domain knowledge base contains expertise and terminology in a specific field, while expert guidance can provide valuable experience and suggestions for the construction of dictionaries.
[0005] One method for establishing an industry dictionary is based on user search behavior logs, that is, utilizing the user's industry knowledge, through the user's search terms and the corresponding clicked search results, and forming a fuzzy dictionary by analyzing the click probability, search frequency, search term splitting, and establishing fuzzy indexes.
[0006] Another type of industry dictionary is generated based on industry terms and corresponding document sets. Industry relevance analysis is performed on candidate terms, and the chi-square test or information gain algorithm is used to calculate the relevance of each candidate term to the industry category to which it belongs. Based on the size of the relevance and the co-occurrence analysis and association mining of the candidate terms, industry vocabulary is generated and finally an industry dictionary is established.
[0007] However, when building industry dictionaries, existing methods need to take into account the professionalism, diversity, dynamics and complexity of different industries, and require a lot of time and resources for technology integration, data preprocessing, model training and tuning, etc., which is difficult to implement and takes a long time to implement. Secondly, since the dictionary needs to include all important words and phrases in all industries as much as possible, including some professional terms, abbreviations, slang, etc., this requires a lot of manpower, time and resources, and is a huge project. Then, considering that the dictionary needs to adapt to the dynamics and changes of the industry, the words and phrases in the industry will change over time, such as the emergence of new terms, phrases or concepts, or the changes in the meaning of the original words and phrases. Therefore, the dictionary needs to be constantly updated and maintained to maintain its timeliness and accuracy.
[0008] In addition, due to language limitations, dictionaries established in the past cannot meet the needs of multilingual users.
[0009] In view of the above analysis, the technical problems that urgently need to be solved in the existing technology are: the current industry dictionary construction process is time-consuming, inefficient, low in accuracy, incomplete dictionary coverage, and difficulty in unifying multiple languages. Summary of the invention
[0010] In view of the problems existing in the prior art, the present invention provides a method and system for establishing a multilingual industry dictionary based on a large model.
[0011] The present invention is implemented as follows: a method for establishing a multilingual industry dictionary based on a large model, comprising the following steps:
[0012] Step 1: Industry node search: perform a broad search through a large model to obtain industry node information related to products and services;
[0013] Step 2: Identify industry relationships. Use the existing industry classifications in the United States or other countries as a starting point and use a large model to search for industry names.
[0014] Step 3: Establish industry classification. Use the big model to identify the inclusion, same-level, and subordinate relationships between industry node names, mark the identified industry node names and include them in the industry dictionary; nodes not included in the dictionary are classified as "unclassified industry nodes";
[0015] Step 4: Calculate industry relationships, calculate industry similarity M and industry inclusion C for all unclassified industry nodes and the initially constructed industry classification;
[0016] Step 5: Breadth search, starting from the second-level industry, for each level of industry nodes, detect adjacent industry nodes level by level in a breadth-first manner;
[0017] Step 6: Deep search, starting from the first industry node at the first level, vertically explore the industry chain in a depth-first manner;
[0018] Step 7: Establish an industry dictionary. According to the industry name, find relevant chemical formulas, academic names, abbreviations, synonyms, antonyms, slang, other names, etc. to improve the industry name and build an industry dictionary;
[0019] Step 8: Industry dictionary verification, verify from both the enterprise and market aspects;
[0020] Step 9: Industry dictionary update, regularly update and maintain industry classification dictionaries.
[0021] Further, the following steps are included:
[0022] (1) Industry node search: Use the big model to perform semantic analysis on the input product or service keywords, obtain industry node information related to the input keywords through breadth-first search, and extract the name of the industry node and its context;
[0023] (2) Industry relationship classification: Based on the preset standard industry classification system (such as NAICS, ISIC, etc.), the large model is used to identify the inclusion relationship, peer relationship or subordinate relationship between industry nodes, mark the classified industry nodes, and classify the unclassified industry nodes as "unclassified industry nodes";
[0024] (3) Industry dictionary establishment: Through semantic expansion, a related vocabulary of industry nodes is constructed, including but not limited to chemical formulas, abbreviations, synonyms, antonyms, slang and academic terms, and a multilingual dictionary of industry classification is generated.
[0025] Furthermore, the industry relationship classification in step (2) includes the following sub-steps:
[0026] (1) The semantic similarity M between unclassified industry nodes and classified industry nodes is calculated using the large model, and the similarity between industry nodes is determined by the cosine similarity of the embedded vectors;
[0027] (2) Using the large model, the industry inclusion degree C of unclassified industry nodes and classified industry nodes is calculated, and based on the semantic context, whether the industry nodes have an inclusion relationship is determined;
[0028] (3) Based on the comprehensive scores of similarity M and inclusion C, the unclassified industry nodes are classified into the best matching industry classification or retained as special industry nodes for manual review.
[0029] Furthermore, the industry dictionary establishment step also includes the following contents:
[0030] (1) Through the depth-first search algorithm, identify the vertical relationship between nodes at each level in the industry chain and build upstream and downstream industry associations;
[0031] (2) For each industry node, common terms, industry buzzwords, and market expressions are automatically extracted through the contextual semantics generated by the large model;
[0032] (3) Verify the accuracy of the industry dictionary through enterprise and market data, and dynamically update the dictionary based on mismatches or newly added industry nodes.
[0033] Furthermore, the industry node search in step 1 includes:
[0034] First, for each product, query its production process, raw materials, technology, equipment, intermediate products and other key information one by one according to the product name; for each production link, further query the relevant work flow, sub-process, raw materials, technology, equipment, intermediate products, and continue to iterate until all relevant information is obtained; this process will obtain detailed data including raw materials, technology, equipment and products, and integrate them into the product node data set;
[0035] Similarly, starting from the service, key information such as service objects, service content, service process, required technology, required equipment, application scenarios, etc. are queried one by one according to the service name; for each service link, the relevant service objects, service content, work flow, required technology, required equipment, application scenarios are queried, and iterate continuously until all relevant information is obtained; this process will obtain data such as service objects, service content, technology, equipment and application scenarios, and integrate them into the service node data set.
[0036] Furthermore, the industry relationship identification in step 2 includes:
[0037] Set the 13 major industries as the first level, find the sub-industry nodes under each industry to form the second level; according to the second-level name, further find the same-level industry nodes under each industry classification, and add all the same-level industry nodes to the second level; continue to find sub-industry nodes and inclusion relationship nodes according to the second-level name, build the third level of industry classification, and so on; until the last level is found, store the industry name and industry level in the industry relationship data set.
[0038] Furthermore, in step 4, the industry matching degree is calculated using a hypertext matching algorithm, taking into account the industry name matching degree M(name) and the actual matching degree M(real), and the industry matching degree M=W M(name) *M(name)+W M(real) *M(real); Considering the industry name inclusion C(name) and the actual inclusion C(real), the industry inclusion C=W C(name) *M(name)+W C(real) *M(real).
[0039] Furthermore, the industry name matching degree M(name) is calculated using the cosine similarity method; the similarity between two industry names is measured by comparing their vector representations, using the formula The cosine similarity value ranges from 0 to 1, with 1 indicating complete similarity and 0 indicating complete dissimilarity.
[0040] The actual matching degree of the industry, M (real), judges the truest essence of a product or service from multiple perspectives, such as essence and purpose. The industry is described by defining attributes (A) of multiple dimensions, including technical complexity, product type, application scenario, performance parameters, customizability, etc. Each dimension has specific quantitative indicators, such as technical complexity: using a technical difficulty index, such as a score from 1 to 10, with 10 representing the highest difficulty. Application scenario: using scenario coding or classification. Performance parameters: including speed, efficiency, capacity, etc., quantified using specific values or scores. By scoring each dimension, the final result is calculated. Get the actual match.
[0041] Furthermore, the breadth search in step 5 includes:
[0042] For each detected node, industry nodes greater than the threshold are classified as nodes of the same level according to the industry matching degree M; this process is iterated until all nodes of the same level are detected. The large model is used to detect whether unclassified industry nodes can be added to the current industry level, and the nodes that can be added are included in the current industry level. Check whether the same-level industry nodes of each industry node at each level are already included in the industry classification of that level.
[0043] Furthermore, the depth search in step 6 includes:
[0044] For each detected node, according to the industry inclusion degree C, the industry nodes greater than the threshold are classified as sub-industry nodes. This process is iterated until all sub-industry nodes are detected. The large model is used to detect whether the unclassified industry nodes can be added to the sub-industry level, and the nodes that can be added are included in the sub-industry level. Check whether the sub-industry nodes of each industry node in each industry chain are already included in the next level of industry classification.
[0045] Furthermore, the industry dictionary verification in step eight includes:
[0046] Verify the industry names associated with the enterprise through the big model, and check whether the enterprise's main business, raw materials, products, and services are fully included in the industry dictionary. Verify whether all industries included in the report are in the industry dictionary based on the industry report.
[0047] Further, the updating of the industry dictionary in step nine includes:
[0048] Find new product names through corporate announcements, news and other information, calculate industry matching and inclusion, and add new node names to the industry dictionary; calculate industry matching and inclusion based on multilingual industry relationships, and add new node names to the industry dictionary; monitor the latest developments and trends in the industry, use the inference capabilities of large models to predict and classify new words, phrases or concepts, and regularly update and maintain industry classification dictionaries to ensure the real-time and updateability of the dictionary and adapt to industry changes and language evolution.
[0049] Another object of the present invention is to provide a method for establishing a multilingual industry dictionary based on a large model and a system for establishing a multilingual industry dictionary based on a large model, comprising:
[0050] Industry node search module, which conducts breadth search through large models to obtain industry node information related to products and services;
[0051] The industry relationship identification module uses the existing industry classifications in the United States or other countries as a starting point and uses a large model to search for industry names;
[0052] The industry classification establishment module uses a large model to identify the inclusion, same-level, and subordinate relationships between industry node names, mark the identified industry node names and include them in the industry dictionary; nodes not included in the dictionary are classified as "unclassified industry nodes";
[0053] The industry relationship calculation module calculates the industry similarity M and industry inclusion C of all unclassified industry nodes and the initially constructed industry classification;
[0054] The breadth search module starts from the second-level industry and detects adjacent industry nodes at each level in a breadth-first manner;
[0055] The deep search module starts from the first industry node at the first level and vertically explores the industry chain in a depth-first manner;
[0056] Industry dictionary building module, based on the industry name, finds relevant chemical formulas, academic names, abbreviations, synonyms, antonyms, slang, other names, etc., to improve the industry name and build an industry dictionary;
[0057] Industry dictionary verification module, which verifies from both enterprise and market perspectives;
[0058] Industry dictionary update module, regularly updates and maintains industry classification dictionaries.
[0059] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:
[0060] First, the present invention proposes a method for establishing a multilingual industry classification dictionary based on a large model, establishes a hypertext matching algorithm, and establishes a multilingual industry classification dictionary through a large model combined with breadth search and depth search. Through breadth search, terms and expressions with similar semantics to the query words or phrases can be obtained, and through depth search, terms and expressions with similar and subordinate relationships of the query words can be found. Compared with traditional dictionary construction methods, this method has the following technical advantages:
[0061] 1. Introducing big models into dictionary construction, the big model's natural massive background knowledge improves the coverage of the constructed dictionary, and the reasoning ability of the big model improves the accuracy of industry dictionary discovery.
[0062] 2. An innovative hypertext matching calculation method is proposed, which integrates name matching and substance matching, comprehensively considers multiple dimensions in text matching, accurately weighs the relationship between industry names and substance, and significantly improves the classification accuracy of industry nodes.
[0063] 3. The deep search, broad search and large model are organically combined to ensure the richness of the dictionary from two levels.
[0064] 4. The highly automated industry dictionary construction method is adopted to improve the efficiency of industry dictionary construction.
[0065] In summary, this method uses a large model combined with breadth search and depth search to establish a multilingual industry classification dictionary, which can comprehensively consider the professionalism, diversity, dynamics and complexity of different industries, and can greatly improve the efficiency and accuracy of dictionary establishment, solve the problems of long construction time and high construction difficulty in the past, and save a lot of manpower, time and resources. Taking into account the need for dictionaries to adapt to the dynamics and variability of the industry, the dictionary can be updated and maintained autonomously to adapt to industry changes and language evolution, and maintain its timeliness and accuracy. In addition, due to language limitations, the dictionaries established in the past cannot meet the needs of multilingual users. The large model can solve the problem of multilingual support, support the industry classification needs of different language backgrounds, and meet the needs of users with different language backgrounds.
[0066] Second, as auxiliary evidence of the inventiveness of the claims of the present invention, it is also reflected in the following important aspects:
[0067] (1) The expected benefits and commercial value of the technical solution of the present invention after transformation are:
[0068] By improving the efficiency and accuracy of the automated construction of industry classification dictionaries, the present invention can help overseas companies more accurately locate their own industries, optimize market strategies, and thus enhance market competitiveness. In addition, the multi-language support feature enables the present invention to be widely used in fields such as international trade and cross-border e-commerce, meeting the industry classification needs of users with different language backgrounds around the world, and further broadening the market application space. Therefore, the implementation of the present invention will bring considerable economic benefits to enterprises and promote the rapid development of related industries.
[0069] (2) The technical solution of the present invention fills the technical gap in the industry at home and abroad:
[0070] The technical solution of the present invention fills the technical gap in the method of establishing a multilingual industry classification dictionary based on a large model at home and abroad. Traditional dictionary construction methods are often limited by language types, construction efficiency and accuracy, and it is difficult to meet the growing globalization and multilingual needs. The present invention, by introducing a large model and combining breadth search and depth search strategies, not only improves the coverage and accuracy of the dictionary, but also realizes support for multiple languages, providing a new and efficient solution for the establishment of industry classification dictionaries. This innovative technical solution has undoubtedly injected new vitality into the technological development of related industries at home and abroad, and promoted the advancement and upgrading of industry technology.
[0071] (3) The technical solution of the present invention solves the technical problems that people have been eager to solve but have never been able to solve successfully:
[0072] The technical solution of the present invention successfully overcomes the technical problem that people have long been eager to solve but have never made a breakthrough progress - that is, how to efficiently and accurately build an industry classification dictionary that supports multiple languages and can adapt to dynamic changes in the industry. Traditional methods are limited by language barriers, long construction cycles, and difficult updates and maintenance, making it difficult to meet the rapidly changing market demands and the growing demand for multilingual services. The present invention not only greatly improves the efficiency and accuracy of dictionary construction, but also realizes the autonomous update and multilingual support of the dictionary through the introduction of large models and the combination of deep and wide searches, which truly solves this technical problem and opens up a new path for the future development of industry classification dictionaries.
[0073] (4) The technical solution of the present invention overcomes technical prejudice:
[0074] The technical solution of the present invention overcomes the previous prejudice in the technical field, that is, it is believed that the construction of industry classification dictionaries must rely on a lot of manual intervention and traditional methods, and it is difficult to achieve multi-language support and efficient updates at the same time. By introducing large model technology, combined with innovative hypertext matching algorithms and depth-breadth search strategies, the present invention not only significantly improves the automation and accuracy of dictionary construction, but also successfully achieves comprehensive support for multiple languages, as well as autonomous updating and maintenance of dictionaries. This technical solution breaks the limitations of traditional concepts and proves that with the support of big data and artificial intelligence technology, the construction of industry classification dictionaries can be more efficient, intelligent and flexible, injecting new vitality and innovation into the technological development of related fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] Figure 1 is a flow chart of a method for establishing a multilingual industry dictionary based on a large model provided by an embodiment of the present invention;
[0076] Figure 2 is a schematic diagram of a deep search provided by an embodiment of the present invention;
[0077] Figure 3 It is a structural diagram of a system for establishing a multilingual industry dictionary based on a large model provided by an embodiment of the present invention.
[0078] Figure 4 It is a part of the industry classification that expands the coverage of the English industry dictionary provided by the embodiment of the present invention DETAILED DESCRIPTION
[0079] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0080] The embodiment of the present invention uses a large language model (such as GPT, etc.) to parse and expand the keywords of products or services using a breadth search algorithm. The model obtains industry node information related to the input keyword through semantic understanding, and identifies the industry name and sub-industry name related to it by analyzing the contextual relationship of the industry keyword. After obtaining the basic industry node, the semantic association between industry nodes is further parsed to form a preliminary industry hierarchical structure.
[0081] Based on the standardized industry classification systems of various countries (such as the North American Industry Classification System NAICS, ISIC, etc.), with the large model as the core, the relationship between industry nodes and standard industry classification is expanded and mined. The model uses natural language processing technology to identify the inclusion relationship, peer relationship or subordinate relationship between industry nodes, mark the classified industry nodes, and mark the unmatched nodes as "unclassified industry nodes". This process is combined with semantic similarity analysis in the model to ensure the accuracy of node classification.
[0082] For unclassified industry nodes, the embedding vector of the large model is used to calculate the similarity and inclusion indicators:
[0083] Similarity M: Generate semantic vectors of unclassified nodes and existing industry nodes through the model and calculate cosine similarity to determine their similarity;
[0084] Inclusion C: Identify whether there is a subordinate or inclusive relationship through partial matching of industry node names and their contextual semantics. Based on the combined scores of M and C, unclassified nodes are included in the closest industry classification to further improve the industry classification.
[0085] Starting from the second level of industry classification, the industry extension of adjacent nodes is detected using breadth-first search to ensure comprehensive coverage of horizontal relationships. Subsequently, for the vertical depth of the industry chain, the upstream and downstream relationships of each level of the industry are analyzed using depth-first search to clarify its vertical level in the industry chain. This step ensures the integrity of the horizontal and vertical dimensions of the industry classification system.
[0086] After initially completing the industry classification, the big model is called to expand the associated vocabulary of each industry node. Specifically, it includes:
[0087] Chemical formula or technical term: match academic expressions related to the industry;
[0088] Abbreviations and synonyms: Common names for expanding industry nodes;
[0089] Slang and synonyms: Common but informal expressions used in the supplement market.
[0090] Through multiple rounds of contextual training, the large model generates a structured industry dictionary and links it to the industry hierarchy.
[0091] Once established, industry dictionaries are validated through enterprise applications and market data:
[0092] Enterprise verification: Randomly sample enterprise description information, automatically match the industry to which it belongs and verify the correctness of the classification;
[0093] Market validation: Analyze the consistency between industry buzzwords, market trend data and dictionaries to identify missing nodes or incorrect classifications.
[0094] In line with the dynamic changes in industry development, a regular update mechanism for the dictionary is set up to ensure that it continues to adapt to emerging industries and market changes, and ultimately form a dynamically optimized multilingual industry classification tool.
[0095] like Figure 1 As shown, an embodiment of the present invention provides a method for establishing a multilingual industry dictionary based on a large model, comprising the following steps:
[0096] Step 1: Industry node search: perform a broad search through a large model to obtain industry node information related to products and services.
[0097] First, for each product, we query its production process, raw materials, technology, equipment, intermediate products and other key information one by one according to the product name. For each production link, we further query the relevant work flow, subdivided process, raw materials, technology, equipment, intermediate products, and continue to iterate until all relevant information is obtained. This process will obtain detailed data including raw materials, technology, equipment and products, and integrate them into the product node data set.
[0098] Similarly, starting from the service, key information such as service object, service content, service process, required technology, required equipment, application scenario, etc. is queried one by one according to the service name. For each service link, the relevant service object, service content, process, required technology, required equipment, application scenario, and continuous iteration are performed until all relevant information is obtained. This process will obtain data such as service object, service content, technology, equipment, and application scenario, and integrate them into the service node data set.
[0099] Step 2: Identify industry relationships. Based on the existing industry classifications in the United States or other countries, use the big model to search for industry names. Set the 13 major industries as the first level, find the sub-industry nodes under each industry, and form the second level. According to the second-level name, further find the same-level industry nodes under each industry classification, and add all the same-level industry nodes to the second level. Continue to find sub-industry nodes and inclusion relationship nodes according to the second-level name, build the third level of industry classification, and so on. Until the last level is found, store the industry name and industry level in the industry relationship dataset.
[0100] Step 3: Establish industry classification, identify the inclusion, same level, subordinate and other relationships between industry node names through the big model, mark the identified industry node names and include them in the industry dictionary. Nodes not included in the dictionary are classified as "unclassified industry nodes".
[0101] Step 4: Calculate industry relationships. Calculate industry similarity M and industry inclusion C for all unclassified industry nodes and the initially constructed industry classification. Propose and use a hypertext matching algorithm to calculate industry matching, considering industry name matching M (name) and real matching M (real). Industry matching M = W M(name) *M(name)+W M(real) *M(real). Considering the industry name inclusion C(name) and the actual inclusion C(real), the industry inclusion C=W C(name) *M(name)+W C(real) *M(real).
[0102] Among them, the industry name matching degree M(name) is calculated using the cosine similarity method. The similarity between two industry names is measured by comparing their vector representations, using the formula The cosine similarity value ranges from 0 to 1, with 1 indicating complete similarity and 0 indicating complete dissimilarity.
[0103] The actual matching degree of the industry, M (real), judges the true essence of the product or service from multiple perspectives such as essence and purpose. The industry is described by defining attributes (A) of multiple dimensions, such as technical complexity, product type, application scenario, performance parameters, customizability, etc. Each dimension has specific quantitative indicators, such as technical complexity: using a technical difficulty index, such as a score from 1 to 10, with 10 representing the highest difficulty. Application scenario: using scenario coding or classification. Performance parameters: including speed, efficiency, capacity, etc., quantified using specific values or scores. By scoring each dimension, the final result is calculated. Get the actual match.
[0104] Step 5: Breadth search, starting from the second-level industry, detect adjacent industry nodes for each level of industry nodes in a breadth-first manner. For each detected node, industry nodes greater than the threshold are classified as the same-level industry nodes according to the industry matching degree M. This process is iterated until all nodes of the same level have been detected. Use the large model to detect whether unclassified industry nodes can be added to the current industry level, and include the nodes that can be added into the current industry level. Check whether the same-level industry nodes of each industry node at each level are already included in the industry classification of that level.
[0105] Step 6: Deep search, starting from the first industry node of the first level, vertically explore the industry chain in a depth-first manner. For each detected node, according to the industry inclusion degree C, the industry nodes greater than the threshold are classified as sub-industry nodes. This process is iterated until all sub-industry nodes are detected. Use the large model to detect whether the unclassified industry nodes can be added to the sub-industry level, and include the nodes that can be added into the sub-industry level. Check whether the sub-industry nodes of each industry node in each industry chain are already included in the industry classification of the next level.
[0106] Step 7: Establish an industry dictionary. According to the industry name, search for relevant chemical formulas, academic names, abbreviations, synonyms, antonyms, slang, other names, etc. to improve the industry name and build an industry dictionary.
[0107] Step 8: Industry dictionary verification, from both the enterprise and market aspects. Use the big model to verify the industry name associated with the enterprise, and check whether the enterprise's main business, raw materials, products, and services are fully included in the industry dictionary. Verify whether all industries included in the report are in the industry dictionary based on the industry report.
[0108] Step 9: Update the industry dictionary. Find new product names through corporate announcements, news and other information, calculate industry matching and inclusion, and add new node names to the industry dictionary. Calculate industry matching and inclusion based on multilingual industry relationships and add new node names to the industry dictionary. By monitoring the latest developments and trends in the industry, use the inference capabilities of large models to predict and classify new words, phrases or concepts, and regularly update and maintain the industry classification dictionary to ensure the real-time and updateability of the dictionary and adapt to industry changes and language evolution.
[0109] like Figure 3 As shown, the embodiment of the present invention provides a method for establishing a multi-language industry dictionary based on a large model and a system for establishing a multi-language industry dictionary based on a large model, including:
[0110] Industry node search module, which conducts breadth search through large models to obtain industry node information related to products and services;
[0111] The industry relationship identification module uses the existing industry classifications in the United States or other countries as a starting point and uses a large model to search for industry names;
[0112] The industry classification establishment module uses a large model to identify the inclusion, same-level, and subordinate relationships between industry node names, mark the identified industry node names and include them in the industry dictionary; nodes not included in the dictionary are classified as "unclassified industry nodes";
[0113] The industry relationship calculation module calculates the industry similarity M and industry inclusion C of all unclassified industry nodes and the initially constructed industry classification;
[0114] The breadth search module starts from the second-level industry and detects adjacent industry nodes at each level in a breadth-first manner;
[0115] The deep search module starts from the first industry node at the first level and vertically explores the industry chain in a depth-first manner;
[0116] Industry dictionary building module, based on the industry name, finds relevant chemical formulas, academic names, abbreviations, synonyms, antonyms, slang, other names, etc., to improve the industry name and build an industry dictionary;
[0117] Industry dictionary verification module, which verifies from both enterprise and market perspectives;
[0118] Industry dictionary update module, regularly updates and maintains industry classification dictionaries.
[0119] Example 1
[0120] A method for establishing a multilingual industry dictionary based on a large model includes the following steps:
[0121] Step 1: Industry node search: perform a broad search through a large model to obtain industry node information related to products and services.
[0122] Taking "electronic products" as an example, query the production process of "electronic products", including assembly, testing, packaging and other links; query related raw materials, such as chips, batteries, screens, etc.; obtain technical information, such as semiconductor technology, display technology, battery technology; record equipment information, such as assembly lines, testing equipment, etc.; find intermediate products, such as circuit boards, etc.
[0123] Taking "assembly" as an example, we identify the workflows related to assembly: assembly, assembly; in the "assembly" link, we iterate and search to obtain sub-processes: component assembly, circuit board connection, etc.; raw materials: chips, screens, batteries; technology: welding technology, circuit connection technology; equipment: assembly robots, conveyor belts; intermediate products: assembled finished products. Continuous iteration will obtain detailed data including raw materials, technology, equipment, and products, and integrate them into the product node data set.
[0124] Taking "cloud storage" as an example, query the service objects of "cloud storage": enterprises, individuals, etc.; service content: file storage, data backup; service process: user registration, file upload; technical information: virtualization technology, distributed system technology; equipment: servers, network equipment; application scenarios: online file storage, cloud application deployment. Continuous iteration will obtain data such as service objects, service content, technology, equipment and application scenarios, and integrate them into the service node data set.
[0125] Step 2: Identify industry relationships, starting from the 13 major industries in the United States, to form the first level of industry classification. Perform an iterative search for "electronic manufacturing" under "manufacturing" to form the second level; iteratively search for industry nodes at the same level to obtain "electronic component manufacturing", "semiconductor manufacturing", etc., and add them to the second level; iteratively search for sub-industry nodes and inclusion relationship nodes according to the level name to obtain "circuit board manufacturing", "electronic equipment manufacturing", etc., to build the third level; until all levels are built, the industry name and level are stored in the industry relationship data set.
[0126] Step 3: Establish industry classification, identify the inclusion, same level, and subordinate relationships between industry node names. The industries related to "electronic manufacturing" include electronic component manufacturing, semiconductor manufacturing, circuit board manufacturing, and electronic equipment manufacturing. Among them, "electronic component manufacturing" and "electronic manufacturing" are in a same-level relationship and are marked as "identified"; "semiconductor manufacturing" and "electronic manufacturing" are in a same-level relationship and are marked as "identified"; "circuit board manufacturing" and "electronic manufacturing" are in a child-level relationship and are marked as "identified"; "electronic equipment manufacturing" and "electronic manufacturing" are in a child-level relationship and are marked as "identified"; "circuit board manufacturing" and "electronic equipment manufacturing" are in a same-level relationship. Include the identified industry node names in the industry dictionary and establish corresponding data entries. For nodes that are not included in the dictionary, such as "rigid printed circuit board (RPCB)", they are classified as "unclassified industry nodes", and this node will be processed or further identified in subsequent updates.
[0127] Step 4: Calculate industry relationships. Calculate the industry similarity M and industry inclusion C of the unclassified industry node "Rigid Printed Circuit Board (RPCB)" and the initially constructed industry classification "Multilayer Circuit Board".
[0128] Calculate the industry name matching degree M(name), vector 刚性印刷电路板(RPCB) =[0.8,0.5,0.2], vector 多层电路板 =[0.7,0.6,0.3], and the final industry name matching degree = 0.84.
[0129] Calculate the real matching degree M(real), the technical complexity score of rigid printed circuit board (RPCB) is 0.8, the product type score is 0.9, the application scenario score is 0.75, the performance parameter score is 0.85..., the technical complexity score of multilayer circuit board is 0.7, the product type score is 0.85, the application scenario score is 0.85, the performance parameter score is 0.9..., and finally get the industry real matching degree = 0.77. Considering the industry name matching degree M(name) = 0.84 and the real matching degree M(real) = 0.77, calculate the industry matching degree M = 0.6*0.84+0.4*0.77 = 0.81.
[0130] The industry inclusion C is calculated for all unclassified industry nodes "rigid printed circuit boards (RPCB)" and the initially constructed industry classification "printed circuit boards (PCB)". Considering the industry name inclusion C(name) = 0.9 and the actual inclusion C(real) = 0.8, the industry inclusion C = 0.3*0.9+0.7*0.8 = 0.83.
[0131] Step 5: Breadth search, starting from the second-level industry, detect adjacent industry nodes for each level of industry nodes in a breadth-first manner. For each detected node, taking "rigid printed circuit board (RPCB)" as an example, according to the industry matching degree M = 0.81, which exceeds the threshold of 0.75, it means that "rigid printed circuit board (RPCB)" and "multi-layer circuit board" have a certain similarity in industry name, and they can be considered to be classified as industry nodes of the same level. Through the large model detection, the unclassified industry node "rigid printed circuit board (RPCB)" can be added to the current industry level. Detect whether the same-level industry node "single-layer circuit board" of each level of industry node "multi-layer circuit board" is included in the industry classification of that level.
[0132] Step 6: Deep search, starting from the first industry node of the first level, vertically explore the industry chain in a depth-first manner. For each detected node, taking "rigid printed circuit board (RPCB)" as an example, according to the industry inclusion degree C = 0.83, it exceeds the threshold of 0.75, indicating that "rigid printed circuit board (RPCB)" and "printed circuit board (PCB)" have a certain inclusion relationship in the industry name, and it can be considered to be classified as a "printed circuit board (PCB)" sub-industry node. Through the large model detection, the unclassified industry node "rigid printed circuit board (RPCB)" can be added to the sub-industry level. Check whether the sub-industry node "flexible printed circuit board (FPC)" of the industry node "intelligent printed circuit board (PCB)" in each industry chain is included in the next level of industry classification.
[0133] Step 7: Establish an industry dictionary. According to the industry name "printed circuit board", search for related chemical formulas: None, academic name: None, abbreviation: "PCB (Printed Circuit Board)", synonyms: "circuit board", "Printed Wiring Board (PWB)", synonyms: "electronic board" and "circuit board", slang: "electronic substrate", other names: "electronic circuit board", etc., to complete the industry name and build an industry dictionary.
[0134] Step 8: Industry dictionary verification, from both the enterprise and market aspects. Use a large model to verify the industry names associated with the enterprise "Honeywell International" "aerospace, construction technology, chemicals, materials, manufacturing, etc.", and check whether the enterprise's main business "aviation and aerospace systems, equipment and services, etc.", raw materials "aluminum alloys, titanium alloys, magnesium alloys", products "flight management systems, navigation systems", and services "providing energy field technologies and solutions" are fully included in the industry dictionary. According to the industry report "China Precious Metal Catalyst Industry Research Report 2023", verify whether the industries included in the report "such as palladium chloride, rhodium chloride, palladium acetate, etc." are all in the industry dictionary.
[0135] Step 9: Update the industry dictionary. Find new product names through corporate announcements, news and other information, calculate industry matching and inclusion, and add new node names to the industry dictionary. Calculate industry matching and inclusion based on multilingual industry relationships and add new node names to the industry dictionary. By monitoring the latest developments and trends in the industry, use the inference capabilities of large models to predict and classify new words, phrases or concepts, and regularly update and maintain the industry classification dictionary to ensure the real-time and updateability of the dictionary and adapt to industry changes and language evolution.
[0136] The present invention proposes a multilingual industry dictionary establishment system based on a large model, which aims to achieve comprehensive and accurate classification and dictionary construction of various industries around the world through advanced artificial intelligence technology. The core of the system relies on the powerful data processing and natural language understanding capabilities of the large model, combined with the collaborative work of multiple modules, to automatically collect, identify, classify and verify industry information, and finally generate a multilingual supported, structured and dynamically updated industry dictionary. Specifically, the system first uses the large model to perform a breadth search through the industry node search module, comprehensively obtains industry node information related to the target enterprise's products and services, and provides a rich data foundation for subsequent relationship identification and classification establishment.
[0137] In the industry relationship identification module, the system uses the existing industry classifications in the United States or other countries as the starting point, and uses the big model to conduct in-depth search and analysis of industry names to identify complex relationships such as inclusion, peers, and subordination between industry nodes. In this way, the system can build a multi-level, structured industry relationship network to provide a basis for the industry classification establishment module. The industry classification establishment module further uses the big model to identify the relationship between industry node names, and incorporates the identified industry nodes into the industry dictionary. Nodes that have not yet been classified are classified as "unclassified industry nodes" to ensure the integrity of the dictionary and the scalability of the system.
[0138] In the industry relationship calculation module, the system calculates the industry similarity and industry inclusion of all unclassified industry nodes and the initially constructed industry classification. By quantitatively analyzing the similarity and inclusion relationship between unclassified nodes and existing classifications, the system can effectively and accurately attribute unclassified nodes to appropriate industry classifications. In addition, the breadth search module and the depth search module adopt breadth-first and depth-first search strategies respectively, and detect and explore industry nodes step by step from different levels and directions to ensure the comprehensiveness and accuracy of the industry relationship network.
[0139] The industry dictionary building module automatically searches for relevant chemical formulas, academic names, abbreviations, synonyms, antonyms, slang and other names based on the industry name, further improving the diversity and accuracy of the industry name. By integrating multi-source data, the system builds an industry dictionary covering multi-language and multi-dimensional information, providing users with rich and accurate industry classification references. At the same time, the industry dictionary verification module verifies the constructed industry dictionary from both the enterprise and market aspects to ensure the actual application effect of the dictionary and the reliability of the data.
[0140] In order to maintain the dynamic update and continuous optimization of the industry dictionary, the present invention also designs an industry dictionary update module. This module regularly maintains and updates the industry dictionary, automatically adds new industry nodes or adjusts existing categories based on the latest market dynamics, industry development trends and user feedback, and ensures that the dictionary always keeps pace with the actual market environment. Through this continuous update mechanism, the system can provide high-quality industry dictionary services for a long time to meet the ever-changing needs of users.
[0141] The large-model-based multilingual industry dictionary establishment system of the present invention has broad application prospects in multiple fields. First, in terms of market research and analysis, enterprises and research institutions can use the system to quickly obtain accurate industry classification information, conduct market segmentation and competition analysis, and improve market insight and decision-making efficiency. Secondly, in the field of enterprise intelligent management, the industry dictionary generated by the system can be integrated into the enterprise's management software to help the enterprise accurately locate its position in the industry and optimize resource allocation and strategic planning. In addition, the multilingual support feature makes the system particularly important in multinational companies and international trade, and can help companies accurately identify and classify industry information in different language environments, and promote the expansion and collaboration of global business.
[0142] In the field of information retrieval and knowledge management, the structured industry dictionary generated by the system can serve as the basis of the knowledge base, support efficient information retrieval and knowledge discovery, and improve the intelligent level of information management. The combination of artificial intelligence and data mining technology makes the system valuable in big data analysis and machine learning applications, and can provide solid data support for data-driven industry analysis and forecasting. At the same time, governments and public institutions can also use the system for industry statistics and policy formulation to improve the scientificity and accuracy of government management.
[0143] Related products include industry dictionary construction software platforms, industry data analysis tools, enterprise intelligent decision support systems, etc. based on the present invention. These products can integrate the core functions of the industry dictionary system, provide a user-friendly interface and powerful data processing capabilities, and meet the needs of different users in terms of industry classification, data analysis, and intelligent decision-making. In addition, the present invention can also combine cloud computing and edge computing technologies to develop more flexible and efficient industry dictionary service solutions, further expanding its scope of application and market influence.
[0144] Through the industry dictionary building system of the present invention, users can achieve accurate classification and rapid identification of various industries around the world, greatly improving the efficiency and accuracy of industry information management. The system's multi-language support and dynamic update functions ensure its continued adaptability and competitiveness in an international and changing market environment. In short, the present invention provides an efficient, intelligent, and comprehensive technical solution for the construction of industry dictionaries, which has significant technical innovation and broad market application prospects.
[0145] The multilingual industry dictionary building system based on the large model of the present invention has the following significant advantages and innovations:
[0146] 1. Efficient data processing capabilities: Relying on the powerful natural language processing capabilities of the large model, the system can quickly and accurately extract and identify industry node information from massive text data, greatly improving the efficiency and accuracy of industry dictionary construction.
[0147] 2. Multi-language support: The system supports the establishment of multi-language industry dictionaries to meet the needs of global enterprises and multinational organizations for industry information in different language environments, enhancing the applicability and internationalization level of the system.
[0148] 3. Automated relationship identification and classification: Through the industry relationship identification module and classification establishment module, the system can automatically identify the complex relationships between industry nodes, achieve accurate industry classification, and reduce errors caused by human intervention and subjective judgment.
[0149] 4. Dynamic update mechanism: The industry dictionary update module ensures the real-time and dynamic nature of the dictionary content, can promptly reflect the latest changes in the market and industry, and maintain the efficiency and practicality of the dictionary.
[0150] 5. Multi-dimensional feature extraction and verification: Combining text feature extraction and dictionary feature extraction, the system constructs a comprehensive industry feature vector and verifies it through both enterprise and market verification to ensure the accuracy and reliability of the industry dictionary.
[0151] 6. Intelligent application support: The system can be integrated into a variety of intelligent application scenarios, such as market analysis, enterprise management, knowledge base construction, etc., to provide intelligent industry information support and improve user work efficiency and decision-making quality.
[0152] The present invention innovatively solves the shortcomings of traditional industry classification methods in data processing efficiency, language support, relationship identification and dynamic updating by constructing a multilingual industry dictionary establishment system based on a large model. Through the collaborative work of multiple modules and the use of the deep learning and natural language processing capabilities of the large model, the system has achieved comprehensive and accurate classification and dictionary construction of various industries around the world, significantly improving the level of intelligent industry information management. The broad application prospects cover multiple fields such as market research, enterprise intelligent management, multinational business support, information retrieval and knowledge management, and have significant technical innovation and broad market application potential. In the future, with the further development and optimization of large model technology, the industry dictionary establishment system of the present invention will continue to improve its performance and functions, provide more efficient and accurate industry information support for global enterprises and institutions, and promote the in-depth development of intelligent industry management and decision-making.
[0153] Relevant evidence of the technical effects achieved by the embodiments of the present invention.
[0154] 1. Industry coverage has been significantly improved:
[0155] Through comparative analysis of a large amount of experimental data, after using the present invention, the number of industry classifications in the industry classification dictionary has increased from the original 5412 to 5798, and the coverage of the dictionary has been significantly improved, which is nearly 30% higher than the traditional method. This fully demonstrates the significant effect of the present invention in improving the quality of industry classification dictionaries. (e.g. Figure 4 )
[0156] 2. Improvement of automation:
[0157] The implementation of the present invention significantly improves the automation level of industry classification dictionary construction. From data preprocessing to dictionary generation, the entire process is almost fully automated, greatly reducing manual intervention and shortening the construction cycle. According to statistics, after using the present invention, the construction efficiency has increased by more than 50%, which is of great significance for improving business efficiency and reducing costs.
[0158] 3. Implementation of multi-language support:
[0159] The present invention has successfully achieved comprehensive support for multiple languages, including but not limited to English, Chinese, French, etc. Through multi-language environment testing, the accuracy and practicality of the present invention in different language environments have been verified. It provides great convenience for multinational companies and multi-language users, and enhances the market competitiveness and application value of the present invention.
[0160] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. It can be understood by a person of ordinary skill in the art that the above-mentioned devices and methods can be implemented using computer executable instructions and / or contained in a processor control code, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the carrier medium. The device and its modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, and can also be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.
[0161] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modifications, equivalent substitutions and improvements made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.
Claims
1. A method for establishing a multilingual industry dictionary based on a large model, characterized in that: The following steps are involved: Step 1: Industry node search: perform a broad search through a large model to obtain industry node information related to products and services; Step 2: Identify industry relationships. Use the existing industry classifications in the United States or other countries as a starting point and use a large model to search for industry names. Step 3: Establish industry classification. Use the big model to identify the inclusion, same-level, and subordinate relationships between industry node names, mark the identified industry node names and include them in the industry dictionary; nodes not included in the dictionary are classified as "unclassified industry nodes"; Step 4: Calculate industry relationships, calculate industry similarity M and industry inclusion C for all unclassified industry nodes and the initially constructed industry classification; Step 5: Breadth search, starting from the second-level industry, for each level of industry nodes, detect adjacent industry nodes level by level in a breadth-first manner; Step 6: Deep search, starting from the first industry node at the first level, vertically explore the industry chain in a depth-first manner; Step 7: Establish an industry dictionary. According to the industry name, find relevant chemical formulas, academic names, abbreviations, synonyms, antonyms, and slang to improve the industry name and build an industry dictionary; Step 8: Industry dictionary verification, verify from both the enterprise and market aspects; Step 9: Industry dictionary update, regularly update and maintain industry classification dictionaries.
2. The method for establishing a multilingual industry dictionary based on a large model as claimed in claim 1, characterized in that: The industry node search in step 1 includes: (1) For each product, query its production process, raw materials, technology, equipment, intermediate products and other key information one by one according to the product name; for each production link, further query the relevant work flow, sub-process, raw materials, technology, equipment, intermediate products, and continue to iterate until all relevant information is obtained; obtain detailed data including raw materials, technology, equipment and products, and integrate them into the product node data set; (2) Starting from the service, query the key information such as service object, service content, service process, required technology, required equipment, and application scenario one by one according to the service name; for each service link, query the relevant service object, service content, work process, required technology, required equipment, and application scenario, and continue to iterate until all relevant information is obtained; obtain data such as service object, service content, technology, equipment, and application scenario, and integrate them into the service node data set.
3. The method for establishing a multilingual industry dictionary based on a large model as claimed in claim 1, characterized in that: The industry relationship identification in step 2 includes: Set the 13 major industries as the first level, find the sub-industry nodes under each industry to form the second level; according to the second-level name, further find the same-level industry nodes under each industry classification, and add all the same-level industry nodes to the second level; continue to find sub-industry nodes and inclusion relationship nodes according to the second-level name, build the third level of industry classification, and so on; until the last level is found, store the industry name and industry level in the industry relationship data set.
4. The method for establishing a multilingual industry dictionary based on a large model as claimed in claim 1, characterized in that: In step 4, the industry matching degree is calculated using a hypertext matching algorithm, taking into account the industry name matching degree M(name) and the actual matching degree M(real), and the industry matching degree M=W M(name) *M(name)+W M(real) *M(real); Considering the industry name inclusion C(name) and the actual inclusion C(real), the industry inclusion C=W C(name) *M(name)+W C(real) *M(real).
5. The method for establishing a multilingual industry dictionary based on a large model as claimed in claim 4, characterized in that: The industry name matching degree M(name) is calculated using the cosine similarity method; the similarity between two industry names is measured by comparing their vector representations, using the formula The cosine similarity value ranges from 0 to 1, where 1 indicates complete similarity and 0 indicates complete difference; The real matching degree of the industry, M(real), is to judge the true nature of the product or service from multiple perspectives such as essence and purpose. The industry is described by defining multiple dimensions of attributes (A), including technical complexity, product type, application scenario, performance parameters, and customizability. Each dimension has specific quantitative indicators. By scoring each dimension, the final result is calculated. Get the actual match.
6. The method for establishing a multilingual industry dictionary based on a large model as claimed in claim 1, characterized in that: The breadth search in step 5 includes: For each detected node, industry nodes with a matching degree greater than the threshold are classified as industry nodes of the same level according to the industry matching degree M. This process is iterated until all nodes of the same level have been detected. The large model is used to detect whether unclassified industry nodes can be added to the current industry level, and the nodes that can be added are included in the current industry level. The same-level industry nodes of each industry node at each level are detected to determine whether they are already included in the industry classification of that level.
7. The method for establishing a multilingual industry dictionary based on a large model as claimed in claim 1, characterized in that: The depth search in step 6 includes: For each detected node, industry nodes greater than the threshold are classified as sub-industry nodes according to the industry inclusion degree C. This process is iterated until all sub-industry nodes have been detected. The large model is used to detect whether unclassified industry nodes can be added to the sub-industry level, and the nodes that can be added are included in the sub-industry level. The sub-industry nodes of each industry node in each industry chain are detected to see whether they are already included in the next level of industry classification.
8. The method for establishing a multilingual industry dictionary based on a large model as claimed in claim 1, characterized in that: The industry dictionary verification in step eight includes: Use the big model to verify the industry names associated with the enterprise, and check whether the enterprise's main business, raw materials, products, and services are fully included in the industry dictionary; verify based on the industry report whether all industries included in the report are in the industry dictionary.
9. The method for establishing a multilingual industry dictionary based on a large model as claimed in claim 1, characterized in that: The industry dictionary update in step nine includes: Find new product names through corporate announcements and news information, calculate industry matching and inclusion, and add new node names to the industry dictionary; calculate industry matching and inclusion based on multilingual industry relationships, and add new node names to the industry dictionary; monitor the latest developments and trends in the industry, use the inference capabilities of large models to predict and classify new words, phrases or concepts, and regularly update and maintain industry classification dictionaries to ensure the real-time and updateability of the dictionary and adapt to industry changes and language evolution.
10. A system for establishing a multi-lingual industry dictionary based on a large model according to the method for establishing a multi-lingual industry dictionary based on a large model as claimed in any one of claims 1 to 9, characterized in that: include: Industry node search module, which conducts breadth search through large models to obtain industry node information related to products and services; The industry relationship identification module uses the existing industry classifications in the United States or other countries as a starting point and uses a large model to search for industry names; The industry classification establishment module uses a large model to identify the inclusion, same-level, and subordinate relationships between industry node names, mark the identified industry node names and include them in the industry dictionary; nodes not included in the dictionary are classified as "unclassified industry nodes"; The industry relationship calculation module calculates the industry similarity M and industry inclusion C of all unclassified industry nodes and the initially constructed industry classification; The breadth search module starts from the second-level industry and detects adjacent industry nodes at each level in a breadth-first manner; The deep search module starts from the first industry node at the first level and vertically explores the industry chain in a depth-first manner; Industry dictionary building module, based on the industry name, finds relevant chemical formulas, academic names, abbreviations, synonyms, antonyms, and slang to improve the industry name and build an industry dictionary; Industry dictionary verification module, which verifies from both enterprise and market perspectives; Industry dictionary update module, regularly updates and maintains industry classification dictionaries.