Ocean and related industry classification and identification method and system

Through the combination of machine learning algorithms and keywords, a marine-related enterprise recognition model was established, which solved the problems of large workload, long time and low accuracy in enterprise classification, and achieved high accuracy and interpretability of enterprise classification, taking into account manual annotation and national standards, improving the accuracy and generalization ability of identification.

CN120296173AActive Publication Date: 2025-07-11GUANGDONG PROVINCIAL MARINE DEV PLANNING RES CENT +1

Patent Information

Application Number
CN202510796999.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-11
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

The existing technology has problems such as large workload, long time consuming, low accuracy and subjective factors in enterprise classification. Especially when faced with a large number of enterprise samples and emerging industries, it is difficult for existing methods to ensure the accuracy and consistency of classification.

Method used

Using machine learning algorithm combined with keywords, a marine-related enterprise identification model is established by establishing an enterprise credit information database, using a random forest model and a TF-IDF algorithm, combining the enterprise name, business scope and national economic industry classification feature fields, and identifying it through relative Hamming distance and feature word replacement, taking into account manual annotation samples and national standard documents, and adjusting the identification rules to improve accuracy.

Benefits of technology

It realizes high-accuracy enterprise classification, has interpretability, can display classification basis, solves the problems of uneven manual labeling samples and inconsistent labeling results, and improves the accuracy and generalization ability of identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296173A_ABST
    Figure CN120296173A_ABST
Patent Text Reader

Abstract

The invention provides an ocean and related industry classification and identification method and system. The method comprises the following steps: S1, establishing an enterprise credit information database; s2, acquiring a standard file, and performing text preprocessing on the standard file and the manual identification sample; s3, establishing a sea-related enterprise identification model according to an artificial identification sample and the standard file; s4, identifying the to-be-identified enterprise by using the sea-related enterprise identification model to obtain a model identification sample, and calculating the identification accuracy and the extra identification ratio of each marine industry by comparing the model identification sample with the artificial identification sample; s5, adjusting the sea-related enterprise identification model, and identifying the to-be-identified enterprise by using the adjusted sea-related enterprise identification model to obtain an enterprise classification result; and S6, calculating an evaluation score of each identified marine industry for the enterprises in the enterprise classification result. The method can improve the accuracy of the recognition result, has interpretability, and can display the classification basis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing. Specifically, it designs a method for classifying and identifying the ocean and related industries based on the public credit information of ocean-related business entities, and particularly relates to a method and system for classifying and identifying the ocean and related industries. Background Art

[0002] Regarding the problem of enterprise classification, in the current practical work in this technical field, the method of manual investigation and identification is mostly adopted, which has a large workload, takes a long time, and is affected by subjective factors, resulting in biased identification results. The more texts there are, the less advisable it is. When analyzing the relationship between enterprises and industries, it is time-consuming and laborious to manually classify enterprises into corresponding industries, especially when facing a large number of enterprise samples to be classified and emerging industries lacking relevant classification experience. And the method of automatically classifying by simply specifying keywords cannot guarantee the accuracy of classification. The selected keywords are not comprehensively covered, and misjudgments often occur.

[0003] In the existing related technologies, such as the invention patent with the publication number CN112686043A, it discloses a method for classifying emerging industries to which an enterprise belongs based on word vectors. The invention obtains the input emerging industry and obtains relevant information on the Internet according to its name; uses the Textrank algorithm according to the relevant information of the emerging industry to obtain its candidate keywords; uses the K-means algorithm to cluster the candidate keywords to obtain the clustering keywords of the emerging industry; obtains the business scope of the enterprise from the official website and obtains the enterprise business word library according to the business scope; expands the clustering keywords of the emerging industry according to the enterprise business word library to obtain the keyword library of the emerging industry; obtains the inverse document frequency weight of the words according to the enterprise business word library; obtains the basic evaluation score, comprehensive evaluation score, and enterprise classification score in sequence according to the business scope of the enterprise to be classified and the keyword library of the emerging industry; and obtains the classification result of the emerging industry to which the enterprise belongs according to the enterprise classification score. This invention only applies to one feature of the enterprise business scope and does not use relevant national standards.

[0004] The invention patent with the publication number CN110019769A discloses an intelligent enterprise classification algorithm, including: a text preprocessing process and a classification algorithm. The text preprocessing process includes feature selection, word segmentation, stop word removal, and selection of representative words; the classification algorithm is a machine learning algorithm that needs to use existing correctly classified data to train the algorithm to obtain a reliable classifier to classify new descriptive texts. When the sampling of the manually labeled samples is uneven, the labeled results are inconsistent with the standards, and when there are no samples of a certain classification, the identification results will be biased. Summary of the Invention

[0005] In view of the deficiencies in the prior art, the present invention provides a method and system for classifying and identifying marine and related industries.

[0006] According to the method and system for classifying and identifying marine and related industries provided by the present invention, the solution is as follows: In the first aspect, a method for classifying and identifying marine and related industries is provided, and the method includes: Step S1: Establish an enterprise credit information database, divide the enterprises therein into manually identified samples and enterprises to be identified, and select multiple feature fields such as enterprise name, business scope, and national economic industry classification in the enterprise credit information database; Step S2: Obtain a standard document, and perform text preprocessing on the standard document and the manually identified samples respectively; Step S3: According to the manually identified samples and the standard document, learn the feature combinations and association rules of enterprises in each marine industry in the three feature fields of enterprise name, business scope, and national economic industry classification, and establish a model for identifying marine-related enterprises; Step S4: Use the model for identifying marine-related enterprises to identify the enterprises to be identified, divide them into n marine industries, obtain model identification samples, and calculate the identification accuracy rate and additional identification ratio of each marine industry by comparing the model identification samples with the manually identified samples; Step S5: Adjust the model for identifying marine-related enterprises according to the identification accuracy rate and additional identification ratio, use the adjusted model for identifying marine-related enterprises to identify the enterprises to be identified, divide them into n marine industries, and obtain enterprise classification results; Step S6: Calculate the evaluation scores for each identified marine industry for the enterprises in the enterprise classification results.

[0007] Preferably, the step S2 includes: Step S2.1: Perform word segmentation extraction on the standard document, and screen out standard marine industry feature words; Step S2.2: Perform word segmentation extraction on the business scope and enterprise name of the manually identified samples, and use the TF-IDF algorithm to screen out sample marine industry feature words from the enterprise keyword library.

[0008] Preferably, the step S3 includes: Step S3.1: Combine the standard marine industry feature words and the sample marine industry feature words to form marine industry feature words; Step S3.2: Take each feature word in the marine industry feature words as a variable, and convert the business scope text of each enterprise in the manual recognition sample into a vector containing only 0 and 1 according to the presence or absence of the feature word. Among them, 0 represents that a certain feature word is not contained in the business scope text of the enterprise, and 1 represents that a certain feature word is contained. After the above conversion for all samples, a one-hot matrix is obtained; Step S3.3: Use the random forest model, take the marine industry classification result of the manual recognition sample as the dependent variable, and take each feature word in the constructed one-hot matrix as the independent variable for training to obtain a classification model for judging the marine industry by feature words; Step S3.4: Export the weights of the independent variables in the random forest model, which are called the feature weights of each feature word. According to the magnitude of the feature weights of the feature words, multiple recognition rules are formulated for each marine industry to form an identification model for marine-related enterprises.

[0009] Preferably, the said step S4 includes: Step S4.1: Segment the text in the enterprise to be recognized, compare the obtained feature words with the marine industry feature words, calculate the relative Hamming distance, and perform fuzzy recognition of the enterprise name and business classification by setting the threshold of the relative Hamming distance; When the Hamming distance between the feature words of the enterprise to be recognized and the marine industry feature words is relatively small compared to the length of the marine industry feature words, select the feature words with the relative Hamming distance, and judge whether the two feature words can be considered the same feature word. If so, replace them with the marine industry feature words when inputting the identification model for marine-related enterprises; Step S4.2: By comparing the model recognition samples and the manual recognition samples, calculate the recognition accuracy rate and the additional recognition ratio of each marine and related industry.

[0010] Preferably, the said step S4.2 includes: The recognition accuracy rate is the proportion of the number of enterprises that are manually recognized and are marine industries to the number of enterprises that are manually recognized as marine industries, which is used to evaluate the capture ability of the identification model for marine-related enterprises for known marine-related enterprises. The calculation method is as follows: The additional recognition ratio is the ratio of the number of enterprises recognized as marine industries by the model to the number of enterprises recognized as marine industries by manual recognition, which is used to evaluate the generalization ability of the identification model for marine-related enterprises in the real scenario. The calculation method is as follows: The thresholds for the recognition accuracy rate and the additional recognition ratio of each ocean and related industries are set differently, ensuring that the recognition accuracy rate of each ocean and related industries is higher than 70%, and the additional recognition ratio is set according to the ratio of the enterprises to be recognized and the manually recognized samples. Preferably, in step S5, adjusting the ocean-related enterprise recognition model includes: Step S5.1: Formulate exclusion rules, including performing word segmentation extraction on the enterprise name keywords of the model recognition samples, using the TF-IDF algorithm and manual supplementation to screen out the ocean industry exclusion words, excluding enterprises that are not ocean-related and whose main business is not this ocean industry, and controlling the additional recognition ratio; Step S5.2: Adjust the recognition rules, including adjusting the set threshold in the TF-IDF algorithm for screening feature words, increasing or decreasing the ocean industry feature words, and at the same time adjusting the feature weights, and increasing or decreasing the recognition rules for each ocean industry; Step S5.3: Conduct manual random inspections, and increase or decrease the ocean industry exclusion words and the ocean industry feature words according to the results of the manual random inspections; Step S5.4: Screen and supplement the ocean industry exclusion words and the ocean industry feature words, adjust the weights of the feature fields, adjust the intersection and union logic in the recognition rules, and increase the recognition rules to adjust the ocean-related enterprise recognition model to obtain the adjusted ocean-related enterprise recognition model; At the same time, evaluate the model performance by calculating the recognition accuracy rate and the additional recognition ratio of the adjusted ocean-related enterprise recognition model.

[0011] Preferably, step S6 includes: Step S6.1: Use the descriptions of each ocean industry in the standard document to form an ocean industry description word library; Step S6.2: Calculate the evaluation scores of the enterprises in the enterprise classification results for each recognized ocean industry according to the text similarity between the ocean industry description word library and the business scope of the enterprises in the enterprise classification results; If an enterprise is classified into multiple ocean and related industries by the ocean-related enterprise recognition model, then judge which ocean industry the enterprise belongs to according to the evaluation scores of the enterprise for each recognized ocean industry.

[0012] In a second aspect, an ocean and related industries classification and recognition system is provided, and the system includes: Module M1: Establish an enterprise credit information database, divide the enterprises therein into manually recognized samples and enterprises to be recognized, and select multiple feature fields such as enterprise name, business scope, and national economic industry classification in the enterprise credit information database; Module M2: Obtain the standard document, and perform text preprocessing on the standard document and the manually recognized samples respectively; Module M3: According to the manually identified samples and the standard documents, learn the characteristic combinations and association rules of enterprises in each marine industry in the three characteristic fields of enterprise name, business scope, and national economic industry classification, and establish an identification model for marine-related enterprises; Module M4: Use the identification model for marine-related enterprises to identify the enterprise to be identified, classify it into n marine industries to obtain model identification samples, and calculate the identification accuracy rate and additional identification ratio of each marine industry by comparing the model identification samples with the manually identified samples; Module M5: Adjust the identification model for marine-related enterprises according to the identification accuracy rate and additional identification ratio, use the adjusted identification model for marine-related enterprises to identify the enterprise to be identified, classify it into n marine industries to obtain the enterprise classification result; Module M6: Calculate the evaluation scores for each identified marine industry for the enterprises in the enterprise classification result.

[0013] Preferably, the module M2 includes: Module M2.1: Perform word segmentation extraction on the standard documents and screen out standard marine industry characteristic words; Module M2.2: Perform word segmentation extraction on the business scope and enterprise name of the manually identified samples, and use the TF-IDF algorithm to screen out sample marine industry characteristic words from the enterprise keyword library; The module M3 includes: Module M3.1: Combine the standard marine industry characteristic words and the sample marine industry characteristic words to form marine industry characteristic words; Module M3.2: Take each characteristic word in the marine industry characteristic words as a variable, and convert the business scope text of each enterprise in the manually identified samples into a vector containing only 0 and 1 according to whether the characteristic word exists, where 0 represents that the business scope text of the enterprise does not contain a certain characteristic word, and 1 represents that it contains a certain characteristic word. After the above conversion for all samples, obtain a one hot matrix; Module M3.3: Use the random forest model, take the marine industry classification result of the manually identified samples as the dependent variable, and take each characteristic word in the constructed one hot matrix as the independent variable for training to obtain a classification model for judging the marine industry by characteristic words; Module M3.4: Export the weights of the independent variables in the random forest model, which are called the characteristic weights of each characteristic word. According to the magnitudes of the characteristic weights of the characteristic words, formulate multiple identification rules for each marine industry to form an identification model for marine-related enterprises.

[0014] Preferably, the module M4 includes: Module M4.1: Segment the text in the enterprise to be recognized, compare the obtained feature words with the marine industry feature words, calculate the relative Hamming distance, and perform fuzzy recognition of the enterprise name and business classification by setting the threshold of the relative Hamming distance; When the Hamming distance between the feature words of the enterprise to be recognized and the marine industry feature words is relatively small compared to the length of the marine industry feature words, select the feature words with the relative Hamming distance, and determine whether the two feature words can be considered the same feature words. If so, use the marine industry feature words for replacement when inputting into the marine-related enterprise recognition model; Module M4.2: Calculate the recognition accuracy rate and additional recognition ratio of each marine and related industry by comparing the model recognition samples and manual recognition samples; The said Module M4.2 includes: The recognition accuracy rate is the proportion of the number of enterprises that are manually recognized and are marine industries to the number of enterprises that are manually recognized as marine industries, which is used to evaluate the capturing ability of the marine-related enterprise recognition model for known marine-related enterprises. The calculation method is as follows: The additional recognition ratio is the ratio of the number of enterprises recognized by the model as marine industries to the number of enterprises recognized by manual as marine industries, which is used to evaluate the generalization ability of the marine-related enterprise recognition model in the real scenario. The calculation method is as follows: The thresholds set for the recognition accuracy rate and additional recognition ratio of each marine and related industry are different, ensuring that the recognition accuracy rate of each marine and related industry is higher than 70%, and the additional recognition ratio is set according to the ratio of the enterprises to be recognized and the manual recognition samples; In the said Module M5, adjusting the marine-related enterprise recognition model includes: Module M5.1: Formulate exclusion rules, including extracting enterprise name keyword segments from the model recognition samples, using the TF-IDF algorithm and manual supplementation to screen out marine industry exclusion words, excluding enterprises that are not marine-related and whose main business is not this marine industry, and controlling the additional recognition ratio; Module M5.2: Adjust the recognition rules, including adjusting the threshold in the TF-IDF algorithm for screening feature words, increasing or decreasing the marine industry feature words, and at the same time adjusting the feature weights, and increasing or decreasing the recognition rules for each marine industry; Module M5.3: Conduct manual spot checks, and increase or decrease the marine industry exclusion words and the marine industry feature words according to the results of the manual spot checks; Module M5.4: Screen and supplement the exclusion words and characteristic words of the marine industry, adjust the weights of the characteristic fields, adjust the intersection and union logic in the recognition rules, and add recognition rules to adjust the recognition model of marine-related enterprises to obtain an adjusted recognition model of marine-related enterprises; At the same time, evaluate the model performance by calculating the recognition accuracy rate and the additional recognition ratio of the adjusted recognition model of marine-related enterprises; The said module M6 includes: Module M6.1: Use the descriptions of each marine industry in the standard document to form a marine industry description thesaurus; Module M6.2: Calculate the evaluation scores of the enterprises in the enterprise classification result for each recognized marine industry according to the text similarity between the marine industry description thesaurus and the business scope of the enterprises in the enterprise classification result; If an enterprise is classified into multiple marine and related industries by the recognition model of marine-related enterprises, then judge which marine industry the enterprise belongs to according to the evaluation scores of the enterprise for each recognized marine industry.

[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. By adopting the method of combining machine learning algorithms with keywords, the present invention solves the problem of weak interpretability of machine learning algorithms and can show the basis for the determination of enterprise classification; 2. By adopting both manually labeled samples and relevant national standard documents, the present invention solves the problems of uneven sampling of manually labeled samples and inconsistent labeling results and standards.

[0016] Other beneficial effects of the present invention will be elaborated in the specific implementation manners through the introduction of specific technical features and technical solutions. Those skilled in the art should be able to understand the beneficial technical effects brought by the said technical features and technical solutions through these introductions. Brief Description of the Drawings

[0017] By reading the detailed description of the non-restrictive embodiments with reference to the following drawings, other features, purposes and advantages of the present invention will become more obvious: Figure 1 It is a schematic diagram of the overall process of the present invention. Detailed Description of the Invention

[0018] The present invention will be described in detail below with reference to specific embodiments. The following embodiments will help those skilled in the art to further understand the present invention, but do not limit the present invention in any form. It should be noted that those of ordinary skill in the art can make several changes and improvements without departing from the concept of the present invention. These all belong to the protection scope of the present invention.

[0019] An embodiment of the present invention provides a method for classifying and identifying marine and related industries. By using a machine learning algorithm combined with keywords, enterprises are automatically classified into twenty-eight marine and related industry classifications and non-marine-related enterprises. Taking into account relevant national standards and manually labeled samples, multiple features such as enterprise name, business classification, and national economic industry classification are used. The recognition result has a high accuracy, is interpretable, can display the classification basis, and can solve the problems of biased manually labeled samples and incomplete classification results. Refer to Figure 1 As shown in Step S1: Establish an enterprise credit information database, divide the enterprises in it into manually identified samples and enterprises to be identified, and select multiple feature fields including enterprise name, business scope, and national economic industry classification.

[0020] Step S2: Obtain standard documents, and perform text preprocessing on the standard documents and the manually identified samples respectively.

[0021] Specifically, the standard documents in this step include "Classification of Marine and Related Industries" (GB / T 20794-2021) and "Technical Regulations for Identifying Marine-related Units" (draft for comments). The standard documents are segmented and extracted, and standard marine industry feature words are screened out by using a large language model (LLM) for understanding and manual screening. The business scope and enterprise name of the enterprises in the manually identified samples are segmented and extracted, and sample marine industry feature words are screened out from the enterprise keyword library by using the TF-IDF algorithm.

[0022] Step S3: According to the manually identified samples and standard documents, learn the feature combinations and association rules of enterprises in each marine industry in the three feature fields of enterprise name, business scope, and national economic industry classification, and establish a marine-related enterprise identification model. This step specifically includes: Step S3.1: Combine the standard marine industry feature words and the sample marine industry feature words to form marine industry feature words; Step S3.2: Take each feature word in the marine industry feature words as a variable, and convert the business scope text of each enterprise in the manually identified samples into a vector containing only 0 and 1 according to whether the feature word exists. Among them, 0 represents that a certain feature word does not exist in the business scope text of the enterprise, and 1 represents that it exists. After the above conversion for all samples, a one hot matrix is obtained; Step S3.3: Use a random forest model, take the marine industry classification results of the manually identified samples as the dependent variable, and take each feature word in the constructed one hot matrix as the independent variable for training to obtain a classification model that can judge the marine industry through feature words; Step S3.4: Export the weights of independent variables in the random forest model, which are called the feature weights of each feature word. According to the magnitudes of the feature weights of the feature words, multiple recognition rules are formulated for each marine industry to form a marine enterprise recognition model.

[0023] Step S4: Use the marine enterprise recognition model to identify the enterprise to be identified, classify it into n marine industries to obtain a model recognition sample, and calculate the recognition accuracy rate and additional recognition ratio of each marine industry by comparing the model recognition sample with the manual recognition sample.

[0024] Step S4.1: Segment the text in the enterprise to be identified, compare the obtained feature words with the marine industry feature words, calculate the relative Hamming distance, and perform fuzzy recognition on the enterprise name and business classification by setting the threshold of the relative Hamming distance; When the Hamming distance between the feature words of the enterprise to be identified and the marine industry feature words is relatively small compared to the length of the marine industry feature words, select the feature words with the relative Hamming distance, and determine whether the two feature words can be considered the same feature word. If so, replace them with the marine industry feature words when inputting into the marine enterprise recognition model; Step S4.2: Calculate the recognition accuracy rate and additional recognition ratio of each marine and related industry by comparing the model recognition sample with the manual recognition sample.

[0025] The said Step S4.2 includes: The recognition accuracy rate is the proportion of the number of enterprises that are manually recognized and are marine industries to the number of enterprises that are manually recognized as marine industries, which is used to evaluate the capturing ability of the marine enterprise recognition model for known marine enterprises. The calculation method is as follows: The additional recognition ratio is the ratio of the number of enterprises that are recognized by the model as marine industries to the number of enterprises that are manually recognized as marine industries, which is used to evaluate the generalization ability of the marine enterprise recognition model in the real scenario. The calculation method is as follows: The thresholds set for the recognition accuracy rate and additional recognition ratio of each marine and related industry are different, ensuring that the recognition accuracy rate of each marine and related industry is higher than 70%, and the additional recognition ratio is set according to the ratio of the enterprise to be identified to the manual recognition sample.

[0026] Step S5: Adjust the marine enterprise recognition model, use the adjusted marine enterprise recognition model to identify the enterprise to be identified, classify it into n marine industries to obtain the enterprise classification result.

[0027] Specifically, the adjustment of the sea-related enterprise recognition model in step S5 includes: Step S5.1: Formulate exclusion rules, including performing word segmentation extraction on the enterprise name keywords of the model recognition samples, using the TF-IDF algorithm and manual supplementation to screen out sea industry exclusion words, excluding enterprises that are not sea-related and whose main business is not this sea industry, and controlling the additional recognition ratio; Step S5.2: Adjust the recognition rules, including adjusting the set threshold in the TF-IDF algorithm for screening feature words, increasing or decreasing the sea industry feature words, and at the same time adjusting the feature weights to increase or decrease the recognition rules for each sea industry; Step S5.3: Conduct manual random inspections, and based on the results of the manual random inspections, increase or decrease the sea industry exclusion words and the sea industry feature words; Step S5.4: Screen and supplement the sea industry exclusion words and the sea industry feature words, adjust the weights of the feature fields, adjust the intersection and union logic in the recognition rules, and increase the recognition rules to adjust the sea-related enterprise recognition model to obtain an adjusted sea-related enterprise recognition model; At the same time, evaluate the model performance by calculating the recognition accuracy rate and the additional recognition ratio of the adjusted sea-related enterprise recognition model.

[0028] Step S6: Calculate the evaluation scores for each recognized sea industry for the enterprises in the enterprise classification results.

[0029] Specifically, step S6 includes: Step S6.1: Use the descriptions of each sea industry in the standard documents to form a sea industry description word library; Step S6.2: Calculate the evaluation scores for the enterprises in the enterprise classification results for each recognized sea industry according to the text similarity between the sea industry description word library and the business scope of the enterprises in the enterprise classification results. If an enterprise is classified into multiple sea and related industries by the sea-related enterprise recognition model, then judge which sea industry the enterprise belongs to according to the evaluation scores of the enterprise for each recognized sea industry.

[0030] The present invention also provides a sea and related industry classification and recognition system, and the sea and related industry classification and recognition system can be implemented by executing the process steps of the sea and related industry classification and recognition method, that is, those skilled in the art can understand the sea and related industry classification and recognition method as the preferred implementation manner of the sea and related industry classification and recognition system. This system specifically includes the following content: Module M1: Establish an enterprise credit information database, divide the enterprises in it into manual recognition samples and enterprises to be recognized, and select multiple feature fields such as enterprise name, business scope, and national economic industry classification.

[0031] Module M2: Obtain standard documents, and perform text preprocessing on the standard documents and the manually identified samples respectively.

[0032] Specifically, the standard documents in this module include "Classification of Marine and Related Industries" (GB / T 20794-2021) and "Technical Regulations for the Identification of Marine-related Entities" (Draft for Comment). Word segmentation and extraction are performed on the standard documents, and standard marine industry characteristic words are screened out by using large language models for understanding and manual screening. Word segmentation and extraction are performed on the business scope and enterprise name in the manually identified samples, and sample marine industry characteristic words are screened out from the enterprise keyword library by using the TF-IDF algorithm.

[0033] Module M3: According to the manually identified samples and standard documents, learn the characteristic combinations and association rules of enterprises in each marine industry in three characteristic fields: enterprise name, business scope, and national economic industry classification, and establish an identification model for marine-related enterprises. This module specifically includes: Module M3.1: Combine the standard marine industry characteristic words and the sample marine industry characteristic words into marine industry characteristic words; Module M3.2: Take each characteristic word in the marine industry characteristic words as a variable, and convert the business scope text of each enterprise in the manually identified samples into a vector containing only 0 and 1 according to the presence or absence of the characteristic word. Among them, 0 represents that a certain characteristic word is not contained in the business scope text of the enterprise, and 1 represents that it contains a certain characteristic word. After the above conversion for all samples, a one-hot matrix is obtained; Module M3.3: Use the random forest model to train with the marine industry classification results of the manually identified samples as the dependent variable and each characteristic word in the constructed one-hot matrix as the independent variable to obtain a classification model that can judge the marine industry through characteristic words; Module M3.4: Export the weights of the independent variables in the random forest model, which are called the characteristic weights of each characteristic word. According to the magnitude of the characteristic weights of the characteristic words, multiple identification rules are formulated for each marine industry to form an identification model for marine-related enterprises.

[0034] Module M4: Use the identification model for marine-related enterprises to identify the enterprises to be identified, classify them into n marine industries to obtain model identification samples, and calculate the identification accuracy rate and additional identification ratio of each marine industry by comparing the model identification samples and the manually identified samples.

[0035] Module M4.1: Perform word segmentation on the text in the enterprises to be identified, compare the obtained characteristic words with the marine industry characteristic words, calculate the relative Hamming distance, and perform fuzzy identification of the enterprise name and business classification by setting the threshold of the relative Hamming distance; When the Hamming distance between the enterprise feature word to be recognized and the marine industry feature word is relatively small compared to the length of the marine industry feature word, select the feature word with the relative Hamming distance, and determine whether the two feature words can be considered the same. If so, replace the enterprise feature word to be recognized with the marine industry feature word in the input marine-related enterprise recognition model. Module M4.2: Calculate the recognition accuracy rate and additional recognition ratio of each marine and related industry by comparing the model recognition samples and manual recognition samples.

[0036] The said module M4.2 includes: The recognition accuracy rate is the proportion of the number of enterprises that are manually recognized and are marine industries to the number of enterprises that are manually recognized as marine industries, which is used to evaluate the capture ability of the marine-related enterprise recognition model for known marine-related enterprises. The calculation method is as follows: The additional recognition ratio is the ratio of the number of enterprises recognized by the model as marine industries to the number of enterprises recognized by manual as marine industries, which is used to evaluate the generalization ability of the marine-related enterprise recognition model in the real scenario. The calculation method is as follows: The thresholds set for the recognition accuracy rate and additional recognition ratio of each marine and related industry are different, ensuring that the recognition accuracy rate of each marine and related industry is higher than 70%, and the additional recognition ratio is set according to the ratio of the enterprises to be recognized and the manual recognition samples.

[0037] Module M5: Adjust the marine-related enterprise recognition model, use the adjusted marine-related enterprise recognition model to recognize the enterprises to be recognized, and classify them into n marine industries to obtain the enterprise classification results.

[0038] Specifically, the adjustment of the marine-related enterprise recognition model in this module M5 includes: Module M5.1: Formulate exclusion rules, including performing word segmentation extraction on the keywords of the enterprise names in the model recognition samples, using the TF-IDF algorithm and manual supplementation to screen out the marine industry exclusion words, excluding the enterprises that are not marine-related and whose main business is not this marine industry, and controlling the additional recognition ratio; Module M5.2: Adjust the recognition rules, including adjusting the set threshold in the TF-IDF algorithm for screening feature words, increasing or decreasing the marine industry feature words, and at the same time adjusting the feature weights, increasing or decreasing the recognition rules for each marine industry; Module M5.3: Conduct manual spot checks, and increase or decrease the marine industry exclusion words and the marine industry feature words according to the results of the manual spot checks; Module M5.4: Screen and supplement the exclusion words and characteristic words of the marine industry, adjust the weights of the characteristic fields, adjust the intersection and union logic in the recognition rules, and add recognition rules to adjust the marine enterprise recognition model to obtain an adjusted marine enterprise recognition model; At the same time, evaluate the model performance by calculating the recognition accuracy rate and the additional recognition ratio of the adjusted marine enterprise recognition model.

[0039] Module M6: Calculate the evaluation scores for each recognized marine industry for the enterprises in the enterprise classification results.

[0040] Specifically, this module M6 includes: Module M6.1: Use the descriptions of each marine industry in the standard document to form a marine industry description word library; Module M6.2: Calculate the evaluation scores for each recognized marine industry for the enterprises in the enterprise classification results according to the text similarity between the marine industry description word library and the business scope of the enterprises in the enterprise classification results. If an enterprise is classified into multiple marine and related industries by the marine enterprise recognition model, then judge which marine industry the enterprise belongs to according to the evaluation scores of the enterprise for each recognized marine industry.

[0041] Next, a more specific description of the present invention will be given.

[0042] A method for classifying and recognizing marine and related industries provided by the present invention has a high recognition result accuracy rate, is interpretable, can display the classification basis, and can solve the problems of biased manually labeled samples and incomplete classification results. Specifically, it includes the following contents: Step S1: Establish an enterprise credit information database, and divide the enterprises in it into "manually identified samples" and "enterprises to be identified". And select multiple "characteristic fields" such as enterprise name, business scope, and national economic industry classification in the enterprise credit information database.

[0043] Step S2: Obtain the standard documents, namely "Classification of Marine and Related Industries" (GB / T 20794-2021) and "Technical Regulations for Identifying Marine-related Units" (draft for comments), and use a large language model (LLM) to understand the relevant content of the introduction and division of each marine industry in the documents, and summarize descriptive statements that conform to the characteristics of each marine industry. Subsequently, a series of characteristic words that can best represent each industry are manually screened from the descriptive statements to form "standard marine industry characteristic words".

[0044] Divide each marine industry category in the "manually identified samples" into the corresponding "corpus". Tokenize the text in the "corpus" to obtain the number of occurrences of each feature word in each "corpus". According to the effectiveness of the feature word semantics, manually establish a list of feature words that are not helpful for industrial classification, such as "project", "general", "is", punctuation marks and other "stop words", and form a "stop word list". Exclude the "stop words" in the "corpus" according to the "stop word list".

[0045] Use the TF-IDF algorithm to screen the feature words. The specific steps include calculating the TF-IDF value of each feature word, sorting the feature words according to the calculated TF-IDF value, and screening out the feature words with higher TF-IDF values by setting a threshold to form the "sample marine industry feature words".

[0046] TF-IDF (term frequency–inverse document frequency) is a commonly used weighting technique for information retrieval and text mining. TF-IDF is a statistical method used to evaluate the importance of a word for a document set or a single document in a corpus. The importance of a word increases proportionally with the number of times it appears in a document, but at the same time decreases inversely with the frequency of its appearance in the corpus. The main idea of TF-IDF is that if a certain word appears frequently (high TF) in an article and rarely appears in other articles, then this word or phrase is considered to have good category discrimination ability and is suitable for classification.

[0047] The TF-IDF calculation formula is as follows: TF-IDF is the product of TF and IDF; TF represents term frequency; IDF represents inverse document frequency.

[0048] Among them, represents the th feature word in the th corpus the number of occurrences; represents the th feature word in the th corpus the frequency of occurrence; represents the total number of feature words in the th corpus; represents the Number of feature words in the th corpus appearance times

[0049] Among them, represents the inverse document frequency of the th feature word; represents the number of all corpora, represents the number of corpora containing the feature word

[0050] Step S3: According to the "manually identified samples", "Classification of Marine and Related Industries" and "Technical Regulations for Identifying Marine-related Units", learn the feature combinations and association rules of enterprises in each marine industry in the three feature fields of enterprise name, business scope, and national economic industry classification, and establish an "identification model for marine-related enterprises". This step specifically includes: Step S3.1: Combine the "standard marine industry feature words" and "sample marine industry feature words" to form "marine industry feature words"; Step S3.2: Take each feature word in the "marine industry feature words" as a variable, and convert the business scope text of each enterprise in the "manually identified samples" into a vector containing only 0 and 1 according to whether the feature word exists. 0 represents that the business scope text of the enterprise does not contain a certain feature word, and 1 represents that it contains a certain feature word. After the above conversion for all samples, a matrix containing only 0 and 1 can be obtained, which is called a one hot matrix; Step S3.3: Use the "random forest model", take the marine industry classification results of the "manually identified samples" as the dependent variable, and each feature word in the constructed one hot matrix as the independent variable for training to obtain a classification model that can judge the marine industry through feature words; "Random forest" is an ensemble learning model composed of multiple decision trees. Each decision tree is trained by randomly selecting samples and randomly selecting subsets of features, and finally the results are aggregated by voting (classification) or averaging (regression) to improve the prediction accuracy and generalization ability.

[0051] Step S3.4: Export the weights of the independent variables in the "random forest model", which are called the "feature weights" of each feature word. According to the magnitude of the "feature weights" of the feature words, each marine and related industry formulates multiple "identification rules" to form an "identification model for marine-related enterprises".

[0052] ​There are differences in the intensity of the correlation between different "feature fields" and the authority of the determination basis for whether an enterprise belongs to the marine industry. Moreover, there are semantic conflicts among the "feature words" of different "feature fields" (for example, the "enterprise name" contains "aquaculture" but the "National Economic Industry Classification" belongs to the construction industry). In order to weigh the determination bases of multiple "feature fields", through differential weight assignment and priority rules, the most direct and authoritative "feature field" is preferentially adopted, and then other "feature fields" are gradually combined for supplementary verification, so as to improve the accuracy and efficiency of the recognition results.

[0053] The "recognition rules" include the weights of three "feature fields", namely enterprise name, business scope, and national economic industry classification. Each "feature field" contains multiple "marine industry feature words", as well as the priority and union and intersection logical relationships among different "marine industry feature words".

[0054] For example: The "recognition rules" for marine fishery include: 1. The "National Economic Industry Classification" belongs to "marine aquaculture" or "marine fishing". 2. The "enterprise name" contains "marine aquaculture" or "marine fishery" and other "marine industry feature words" and the "National Economic Industry Classification" belongs to "agriculture, forestry, animal husbandry, and fishery", etc.

[0055] Step S4: Use the "marine-related enterprise recognition model" to recognize the "enterprise to be recognized" and classify it into 28 two-digit major categories of marine and related industries (each marine industry (marine and related industries) mentioned in the present invention refers to the two-digit major categories of marine and related industries in the "Classification of Marine and Related Industries" (GB / T 20794-2021), a total of 28), to obtain the "model recognition sample".

[0056] This step specifically includes: segmenting the text in the "enterprise to be recognized" in the same way as the "manual recognition sample", comparing the obtained feature words with the "marine industry feature words", and calculating the "relative Hamming distance". By setting the threshold of the "relative Hamming distance", the enterprise name and business classification are vaguely recognized.

[0057] The Hamming distance is named after Richard Wesley Hamming. In information theory, the Hamming distance between two strings of the same length is the number of different characters in the corresponding positions of the two strings. In other words, it is the number of characters that need to be replaced to transform one string into another string.

[0058] In actual calculation, the present invention introduces the concept of "relative Hamming distance", and the formula is as follows: When the Hamming distance between the "enterprise to be identified" feature word and the "marine industry feature word" is small, and the length of the "marine industry feature word" is large, the "relative Hamming distance" is small. Select the feature words with a small "relative Hamming distance", and manually judge whether the two feature words can be considered the same feature word. If so, then use the "marine industry feature word" to replace it when inputting into the "marine-related enterprise identification model".

[0059] For example: when the "marine industry feature word" is "aquatic product processing", words such as "aquatic product processing", "processing of aquatic products", "processing of aquatic products", and "production, sales, and processing of aquatic products" are all replaced by "aquatic product processing", so as to perform fuzzy identification on enterprise names and business classifications.

[0060] By comparing the "model identification samples" and the "manual identification samples", calculate the "recognition accuracy rate" and "extra recognition ratio" of each marine and related industry.

[0061] Because a large number of unlabeled samples need to be predicted, traditional supervised learning metrics (such as precision and recall) cannot be directly used. Therefore, the "recognition accuracy rate" and "extra recognition ratio" metrics are introduced to evaluate the model performance.

[0062] The "recognition accuracy rate" is the ratio of the number of enterprises that are manually identified and are in the marine industry to the number of enterprises that are manually identified as in the marine industry. This metric is used to evaluate the capture ability of the "marine-related enterprise identification model" for known marine-related enterprises. When the "recognition accuracy rate" approaches 100%, it indicates that the model covers the positive examples manually labeled sufficiently and the risk of missed detection is low. By evaluating this metric, the consistency between the model's prediction of unlabeled data and manual judgment can be ensured: The "extra recognition ratio" is the ratio of the number of enterprises identified by the model as in the marine industry to the number of enterprises manually identified as in the marine industry. This metric is used to evaluate the generalization ability of the "marine-related enterprise identification model" in real scenarios. If this value is too high, it indicates that the model's decision-making logic is too broad and may misclassify non-marine-related enterprises; if this value is too low, it may miss potential marine-related enterprises that meet the national standards but are not manually labeled. By evaluating this metric, the recognition coverage rate and accuracy can be dynamically balanced during model iteration, especially when dealing with a large amount of unlabeled data, to avoid systematic biases caused by overfitting to manually labeled samples or rigid rules: The thresholds for the "recognition accuracy rate" and "additional recognition ratio" are set differently for each ocean and related industries. Generally, ensure that the "recognition accuracy rate" of each ocean and related industries is higher than 70%. The "additional recognition ratio" is set according to the ratio of "enterprises to be recognized" and "manually recognized samples", and the initial threshold for each ocean and related industries is generally around 10. Subsequent adjustments will be made according to the results of manual spot checks for each ocean and related industries.

[0063] Step S5: Fine-tune and refine the "ocean-related enterprise recognition model", use the adjusted "ocean-related enterprise recognition model" to recognize the enterprises to be recognized, and classify them into the two-digit code categories of 28 ocean and related industries to obtain the "enterprise classification result".

[0064] Specifically, the adjustment of the ocean-related enterprise recognition model in this step S5 includes: Step S5.1: Formulate an "exclusion rule" to exclude enterprises that are clearly not ocean-related and whose main business is not this ocean industry, and control the "additional recognition ratio"; It includes performing word segmentation extraction of enterprise name keywords on the "model recognition samples", using the TF-IDF algorithm and manual supplementation to screen out the "ocean industry exclusion words".

[0065] For example, the "exclusion rule" for "marine fishery" includes: the "enterprise name" contains "feature words" such as "construction" and "advertising" that indicate that the main business of the enterprise is clearly not related to marine fishery.

[0066] Step S5.2: Adjust the "recognition rule", including adjusting the TF-IDF threshold of the candidate feature words, increasing or decreasing the "ocean industry feature words", and at the same time adjusting the "feature weight" to increase or decrease the recognition rule for each ocean industry.

[0067] For example, if the "recognition accuracy rate" is lower than the threshold or the "additional recognition ratio" is too low, then adjust the TF-IDF threshold of the candidate feature words to add more "ocean industry feature words", or modify the intersection and union logic in the "recognition rule" to improve the generalization ability of the model so as to recognize more potential ocean-related enterprises.

[0068] If the "additional recognition ratio" is too high, then adjust the TF-IDF threshold of the candidate feature words to reduce the "ocean industry feature words" or add more "ocean industry exclusion words".

[0069] Step S5.3: Conduct manual spot checks, and according to the results of the manual spot checks, increase or decrease the "ocean industry exclusion words" and "ocean industry feature words"; Analyze the results of manual sampling inspection, and manually judge according to the standard document whether the main business of this enterprise is this marine industry, so as to judge whether the model recognition is correct. If the model misses a sea-related enterprise, add "marine industry characteristic words" to adjust the "recognition rules"; if there is a misjudgment, add "marine industry exclusion words" or modify the "marine industry characteristic words" to adjust the "recognition rules".

[0070] Step S5.4: Adjust the "sea-related enterprise recognition model" to obtain the "adjusted sea-related enterprise recognition model".

[0071] Fine-tune and refine the "sea-related enterprise recognition model" by manually screening and supplementing the "marine industry exclusion words" and "marine industry characteristic words", adjusting the weights of the "characteristic fields", adjusting the intersection and union logic in the "recognition rules", and adding the "recognition rules". At the same time, calculate the "recognition accuracy rate" and "extra recognition ratio" after the model adjustment to evaluate the model performance.

[0072] Step S6: Calculate the evaluation scores for each recognized marine industry for the enterprises in the "enterprise classification result".

[0073] Specifically, this step S6 includes: Step S6.1: Use the descriptions of each marine industry in the "Classification of Marine and Related Industries" to form a "marine industry description word library". Step S6.2: Calculate the text similarity between the "marine industry description word library" and the business scope of the enterprises in the "enterprise classification result" to obtain the "evaluation scores" of the enterprises for each recognized marine industry.

[0074] If an enterprise is classified into multiple marine and related industries by the "sea-related enterprise recognition model", then according to the "evaluation scores" of the enterprise for each recognized marine industry, it can be judged that the enterprise belongs to the marine industry with a higher "evaluation score".

[0075] The text similarity is calculated using cosine similarity. The specific calculation method is to segment the texts of the "marine industry description word library" and the business scope of the enterprises in the "enterprise classification result" to obtain two word lists: A=\left [ {{t}_{1},{t}_{2},...{t}_{i}} \right ] B=\left [ {{t}_{1},{t}_{2},...{t}_{j}} \right ] Merge and de-duplicate the two word lists to obtain all the words in the input sample. T(A,B) = T(A) + T(B) = [t₁, t₂,... tₖ] Calculate the number of occurrences of the k-th word in A and B as the feature vector:

[0076] Finally, substitute into the formula to calculate the cosine similarity: Using the above method, each enterprise can obtain the "cosine similarity" of the enterprise to each marine industry. Divide the cosine similarity value of the specific identified marine industry of the enterprise by the sum of the cosine similarity values of all the identified marine industries of the enterprise to obtain the "evaluation score" of the specific identified marine industry of the enterprise.

[0077] The embodiments of the present invention provide a method and system for classifying and identifying marine and related industries. By adopting a method of combining machine learning algorithms with keywords, the problem of weak interpretability of machine learning algorithms is solved, and the basis for the determination of enterprise classification can be displayed. By taking into account both manually labeled samples and relevant national standard documents, the problems of uneven sampling of manually labeled samples and inconsistency between the labeling results and the standards are solved.

[0078] Those skilled in the art know that in addition to implementing the system and its various devices, modules, and units provided by the present invention in the form of pure computer-readable program code, the method steps can be logically programmed to enable the system and its various devices, modules, and units provided by the present invention to be implemented in the form of logic gates, switches, application-specific integrated circuits, programmable logic controllers, and embedded microcontrollers, etc., to achieve the same functions. Therefore, the system and its various devices, modules, and units provided by the present invention can be regarded as a hardware component, and the devices, modules, and units included therein for implementing various functions can also be regarded as the structure within the hardware component; the devices, modules, and units for implementing various functions can also be regarded as both software modules for implementing the method and the structure within the hardware component.

[0079] The specific embodiments of the present invention have been described above. It should be understood that the present invention is not limited to the above specific embodiments, and those skilled in the art can make various changes or modifications within the scope of the claims, which does not affect the essence of the present invention. Without conflict, the embodiments of the present application and the features in the embodiments can be combined arbitrarily with each other.

Claims

1. A method for classifying and identifying marine and related industries, characterized in that, Including: Step S1: Establish an enterprise credit information database, divide the enterprises in it into manually identified samples and enterprises to be identified, and select multiple feature fields such as enterprise name, business scope, and national economic industry classification in the enterprise credit information database; Step S2: Obtain a standard document, and perform text preprocessing on the standard document and the manually identified samples respectively; Step S3: According to the manually identified samples and the standard document, learn the feature combinations and association rules of enterprises in each marine industry in the three feature fields of enterprise name, business scope, and national economic industry classification, and establish a marine enterprise identification model; Step S4: Use the marine enterprise identification model to identify the enterprises to be identified, divide them into n marine industries to obtain model identification samples, and calculate the identification accuracy rate and additional identification ratio of each marine industry by comparing the model identification samples and the manually identified samples; Step S5: Adjust the marine enterprise identification model according to the identification accuracy rate and additional identification ratio, use the adjusted marine enterprise identification model to identify the enterprises to be identified, divide them into n marine industries to obtain enterprise classification results; Step S6: Calculate the evaluation scores for each identified marine industry for the enterprises in the enterprise classification results.

2. The method for classifying and identifying marine and related industries according to claim 1, wherein The said Step S2 includes: Step S2.1: Perform word segmentation extraction on the standard document and screen out standard marine industry feature words; Step S2.2: Perform word segmentation extraction on the business scope and enterprise name of the enterprises in the manually identified samples, and use the TF-IDF algorithm to screen out sample marine industry feature words from the enterprise keyword library.

3. The method for classifying and identifying marine and related industries according to claim 1, wherein The said Step S3 includes: Step S3.1: Combine the standard marine industry feature words and the sample marine industry feature words to form marine industry feature words; Step S3.2: Take each feature word in the marine industry feature words as a variable, and convert the business scope text of each enterprise in the manually identified samples into a vector containing only 0 and 1 according to whether the feature word exists, where 0 represents that the feature word does not exist in the business scope text of the enterprise, and 1 represents that it exists. After the above conversion for all samples, obtain a one hot matrix; Step S3.3: Use a random forest model, take the marine industry classification results of the manually identified samples as the dependent variable, and take each feature word in the constructed one hot matrix as the independent variable for training to obtain a classification model for judging marine industries by feature words; Step S3.4: Export the weights of the independent variables in the random forest model, which are called the feature weights of each feature word, and formulate multiple identification rules for each marine industry according to the magnitude of the feature weights of the feature words to form a marine enterprise identification model.

4. The marine and related industry classification and identification method according to claim 1, characterized in that, The said Step S4 includes: Step S4.1: Segment the text in the enterprises to be identified, compare the obtained feature words with the marine industry feature words, calculate the relative Hamming distance, and perform fuzzy identification of the enterprise name and business classification by setting the threshold of the relative Hamming distance; When the Hamming distance between the enterprise feature word to be recognized and the marine industry feature word is relatively small compared to the length of the marine industry feature word, select the feature word with the relative Hamming distance, and determine whether the two feature words can be considered the same feature word. If so, replace the enterprise feature word to be recognized with the marine industry feature word when inputting it into the marine-related enterprise recognition model. Step S4.2: Calculate the recognition accuracy rate and additional recognition ratio of each marine and related industry by comparing the model recognition samples and the manual recognition samples.

5. The marine and related industry classification and identification method according to claim 4, wherein The step S4.2 includes: The recognition accuracy rate is for manual recognition and is The number of enterprises in the marine industry accounts for the manual recognition as The proportion of the number of enterprises in the marine industry, which is used to evaluate the capture ability of the sea-related enterprise recognition model for known sea-related enterprises. The calculation method is as follows: The additional recognition ratio is the ratio of the number of enterprises in the marine industry recognized by the model to the number of enterprises in the marine industry recognized manually, which is used to evaluate the generalization ability of the sea-related enterprise recognition model in the real scenario. The calculation method is as follows: The thresholds set for the recognition accuracy rate and additional recognition ratio of each marine and related industry are different, ensuring that the recognition accuracy rate of each marine and related industry is higher than 70%, and the additional recognition ratio is set according to the ratio of the enterprise to be recognized and the manual recognition samples.

6. The method for classifying and identifying marine and related industries according to claim 3, wherein In the step S5, adjusting the marine-related enterprise recognition model includes: Step S5.1: Formulate an exclusion rule, including performing word segmentation extraction on the enterprise name keywords of the model recognition samples, using the TF-IDF algorithm and manual supplementation to screen out the marine industry exclusion words, excluding enterprises that are not marine-related and whose main business is not this marine industry, and controlling the additional recognition ratio. Step S5.2: Adjust the recognition rule, including adjusting the threshold in the TF-IDF algorithm for screening feature words, increasing or decreasing the marine industry feature words, and at the same time adjusting the feature weights, and increasing or decreasing the recognition rules for each marine industry. Step S5.3: Conduct manual random inspections, and increase or decrease the marine industry exclusion words and the marine industry feature words according to the results of the manual random inspections. Step S5.4: Screen and supplement the marine industry exclusion words and the marine industry feature words, adjust the weights of the feature fields, adjust the union and intersection logic in the recognition rules, and increase the recognition rules to adjust the marine-related enterprise recognition model to obtain an adjusted marine-related enterprise recognition model. At the same time, evaluate the model performance by calculating the recognition accuracy rate and additional recognition ratio of the adjusted marine-related enterprise recognition model.

7. The method for classifying and identifying marine and related industries according to claim 1, characterized in that, The step S6 includes: Step S6.1: Use the descriptions of each marine industry in the standard document to form a marine industry description word library. Step S6.2: Calculate the evaluation scores of the enterprises in the enterprise classification results for each recognized marine industry according to the text similarity between the marine industry description word library and the business scope of the enterprises in the enterprise classification results. If an enterprise is classified into multiple marine and related industries by the marine-related enterprise recognition model, then according to the evaluation scores of the enterprise for each recognized marine industry, determine that the enterprise belongs to the marine industry with a higher evaluation score.

8. An ocean and related industries classification and identification system, characterized in that, Including: Module M1: Establish an enterprise credit information database, divide the enterprises in it into manual recognition samples and enterprises to be recognized, and select multiple feature fields such as enterprise name, business scope, and national economic industry classification in the enterprise credit information database. Module M2: Obtain the standard document and perform text preprocessing on the standard document and the manual recognition samples respectively. Module M3: Learn the feature combinations and association rules of the enterprises in each marine industry in the three feature fields of enterprise name, business scope, and national economic industry classification according to the manual recognition samples and the standard document, and establish a marine-related enterprise recognition model. Module M4: Use the offshore enterprise identification model to identify the enterprise to be identified, classify it into n marine industries to obtain a model identification sample, and calculate the identification accuracy rate and additional identification ratio of each marine industry by comparing the model identification sample and the manual identification sample; Module M5: Adjust the offshore enterprise identification model according to the identification accuracy rate and additional identification ratio, use the adjusted offshore enterprise identification model to identify the enterprise to be identified, classify it into n marine industries to obtain an enterprise classification result; Module M6: Calculate the evaluation score for each identified marine industry for the enterprises in the enterprise classification result.

9. The marine and related industries classification and identification system according to claim 8, characterized in that, The module M2 includes: Module M2.1: Perform word segmentation extraction on the standard document and screen out standard marine industry feature words; Module M2.2: Perform word segmentation extraction on the business scope and enterprise name of the manual identification sample, and use the TF-IDF algorithm to screen out sample marine industry feature words from the enterprise keyword library; The module M3 includes: Module M3.1: Combine the standard marine industry feature words and the sample marine industry feature words into marine industry feature words; Module M3.2: Take each feature word in the marine industry feature words as a variable, and convert the business scope text of each enterprise in the manual identification sample into a vector containing only 0 and 1 according to the presence or absence of the feature word, where 0 represents that the business scope text of the enterprise does not contain a certain feature word, and 1 represents that it contains a certain feature word. After performing the above conversion on all samples, obtain a one hot matrix; Module M3.3: Use the random forest model to train with the marine industry classification result of the manual identification sample as the dependent variable and each feature word in the constructed one hot matrix as the independent variable to obtain a classification model for judging the marine industry by feature words; Module M3.4: Export the weights of the independent variables in the random forest model, which are called the feature weights of each feature word. According to the size of the feature weights of the feature words, formulate multiple identification rules for each marine industry to form an offshore enterprise identification model.

10. The marine and related industry classification and identification system according to claim 9, characterized in that The module M4 includes: Module M4.1: Segment the text in the enterprise to be identified, compare the obtained feature words with the marine industry feature words, calculate the relative Hamming distance, and perform fuzzy identification of the enterprise name and business classification by setting the threshold of the relative Hamming distance; When the Hamming distance between the feature words of the enterprise to be identified and the marine industry feature words is relatively small compared to the length of the marine industry feature words, select the feature words with the relative Hamming distance, and judge whether the two feature words can be considered the same feature word. If so, replace them with the marine industry feature words when inputting the offshore enterprise identification model; Module M4.2: Calculate the identification accuracy rate and additional identification ratio of each marine and related industry by comparing the model identification sample and the manual identification sample; The module M4.2 includes: The identification accuracy rate is the ratio of the number of enterprises that are manually identified and belong to marine industry A to the number of enterprises that are manually identified as belonging to marine industry A, and is used to evaluate the capture ability of the offshore enterprise identification model for known offshore enterprises. The calculation method is as follows: The recognition accuracy rate is for manual recognition and is The number of enterprises in the marine industry accounts for the manual recognition as The proportion of the number of enterprises in the marine industry, which is used to evaluate the capture ability of the marine-related enterprise recognition model for known marine-related enterprises. The calculation method is as follows: The additional recognition ratio is the ratio of the number of enterprises recognized by the model as the number of enterprises in the marine industry to the number of enterprises in the marine industry identified manually and is used to evaluate the generalization ability of the offshore enterprise recognition model in real scenarios. The calculation method is as follows: The thresholds for the recognition accuracy rate and the additional recognition ratio of each ocean and related industry are set differently, ensuring that the recognition accuracy rate of each ocean and related industry is higher than 70%, and the additional recognition ratio is set according to the ratio of the enterprises to be recognized and the manually recognized samples; In the module M5, adjusting the ocean-related enterprise recognition model includes: Module M5.1: Formulating exclusion rules, including performing word segmentation extraction on the enterprise name keywords of the model recognition samples, using the TF-IDF algorithm and manual supplementation to screen out the ocean industry exclusion words, excluding enterprises that are not ocean-related and whose main business is not this ocean industry, and controlling the additional recognition ratio; Module M5.2: Adjusting the recognition rules, including adjusting the threshold in the TF-IDF algorithm for screening feature words, increasing or decreasing the ocean industry feature words, and at the same time adjusting the feature weights, and increasing or decreasing the recognition rules for each ocean industry; Module M5.3: Conducting manual spot checks, and increasing or decreasing the ocean industry exclusion words and the ocean industry feature words according to the results of the manual spot checks; Module M5.4: Adjusting the ocean industry exclusion words and the ocean industry feature words by screening and supplementing, adjusting the weights of the feature fields, adjusting the intersection and union logics in the recognition rules, and adding recognition rules, so as to adjust the ocean-related enterprise recognition model and obtain the adjusted ocean-related enterprise recognition model; At the same time, evaluate the model performance by calculating the recognition accuracy rate and the additional recognition ratio of the adjusted ocean-related enterprise recognition model; The module M6 includes: Module M6.1: Using the descriptions of each ocean industry in the standard document to form an ocean industry description word library; Module M6.2: Calculate the evaluation scores of the enterprises in the enterprise classification results for each recognized ocean industry according to the text similarity between the ocean industry description word library and the business scope of the enterprises in the enterprise classification results; If an enterprise is classified into multiple ocean and related industries by the ocean-related enterprise recognition model, then judge which ocean industry the enterprise belongs to according to the evaluation scores of the enterprise for each recognized ocean industry.

Citation Information

Patent Citations

  • Chinese enterprise name entity accurate identification secondary matching method

    CN111597304A

  • Enterprise industry classification identification and characteristic pollutant identification method and device

    CN111914090A

  • Enterprise industry classification method based on domain ontology and system

    CN112182223A

  • Word vector-based enterprise emerging industry classification method

    CN112686043A

  • Enterprise industry recognition system and recognition method based on text similarity

    CN114090736A

Cited By

  • Marine-related enterprise management method and system based on intelligent label

    CN120725630A

  • Intelligent tag-based management method and system for maritime-related enterprises

    CN120725630B