Enterprise Industry Classification Using Semantic Lexicon Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for identifying the industry classification of enterprises using point of information data are prone to errors due to insufficient semantic lexicon capacity, easy overfitting, low computing speed, and low efficiency, which affects the accuracy of soil and groundwater pollution risk management.
Innovation Solution
A method and device that acquire information point data, determine feature words and values using a preset semantic lexicon and industry summary information, and utilize a Gaussian Naive Bayes model with adjusted alpha smoothing parameters to construct an industry classification prediction model, ensuring accurate industry classification and characteristic pollutant identification.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional artificial methods are used to judge industry classification, then identification accuracy is ensured, but manpower and time consumption increases
Solution Approach 1:
The system enables automatic industry classification through self-service mechanisms where the computer automatically extracts feature words from point of information data, calculates feature values, and determines industry classification without human intervention, thereby maintaining high accuracy while eliminating manual labor and time consumption
Solution Approach 2:
The patent replaces the mechanical artificial judgment system with an automated computational system that uses feature word extraction, feature value calculation, and classification algorithms to automatically determine industry classification, substituting human cognitive processes with machine-based automated processing
2Extent of automation
If point of information data is used to determine industry classification, then automation is achieved, but identification accuracy decreases due to errors in extracting feature words
Solution Approach 1:
The patent optimizes identification accuracy by dynamically adjusting parameters including alpha smoothing parameters for probability calculation, feature value weighting parameters, and threshold parameters for classification, allowing the system to adapt to different data characteristics and improve extraction accuracy while maintaining automation
Solution Approach 2:
The system incorporates feedback mechanisms where classification results are validated against known data, and the model is iteratively optimized by adjusting feature extraction parameters and classification thresholds based on performance metrics, thereby improving accuracy while maintaining automated operation
3Productivity
If existing text classification algorithms are used, then processing speed is achieved, but semantic lexicon capacity is insufficient leading to poor decision support
Solution Approach 1:
The patent performs preliminary action by pre-building a comprehensive semantic lexicon with industry-specific feature words and their relationships before classification tasks, and pre-calculating feature value weightings, thereby enabling both fast processing and reliable decision support without compromising computing speed
Data Source
AI summary
Disclosed is a method and device for identifying an industry classification of an enterprise and characteristic pollutants of the enterprise, wherein the method for identifying an industry classification of an enterprise comprises: acquiring information point data of a target enterprise; determining feature words of the information point data and feature values of the feature words according to a preset semantic lexicon, preset industry summary information and the information point data; and determining the industry classification to which the target enterprise belongs according to a preset industry classification prediction model and the feature values. Through implementing the present invention, the obtained feature values can effectively avoid interference of meaningless words, such that the industry classification to which the target enterprise belongs obtained from identification is more accurate.


