Enterprise Industry Classification Using Semantic Lexicon Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for identifying the industry classification of enterprises using point of information data are prone to errors due to insufficient semantic lexicon capacity, easy overfitting, low computing speed, and low efficiency, which affects the accuracy of soil and groundwater pollution risk management.

Innovation Solution

A method and device that acquire information point data, determine feature words and values using a preset semantic lexicon and industry summary information, and utilize a Gaussian Naive Bayes model with adjusted alpha smoothing parameters to construct an industry classification prediction model, ensuring accurate industry classification and characteristic pollutant identification.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional artificial methods are used to judge industry classification, then identification accuracy is ensured, but manpower and time consumption increases

Engineering Contradiction:
Improveidentification accuracyVSAvoidtime consumption
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system enables automatic industry classification through self-service mechanisms where the computer automatically extracts feature words from point of information data, calculates feature values, and determines industry classification without human intervention, thereby maintaining high accuracy while eliminating manual labor and time consumption

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the mechanical artificial judgment system with an automated computational system that uses feature word extraction, feature value calculation, and classification algorithms to automatically determine industry classification, substituting human cognitive processes with machine-based automated processing

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Extent of automation

If point of information data is used to determine industry classification, then automation is achieved, but identification accuracy decreases due to errors in extracting feature words

Engineering Contradiction:
Improveautomation levelVSAvoididentification accuracy
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent optimizes identification accuracy by dynamically adjusting parameters including alpha smoothing parameters for probability calculation, feature value weighting parameters, and threshold parameters for classification, allowing the system to adapt to different data characteristics and improve extraction accuracy while maintaining automation

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system incorporates feedback mechanisms where classification results are validated against known data, and the model is iteratively optimized by adjusting feature extraction parameters and classification thresholds based on performance metrics, thereby improving accuracy while maintaining automated operation

Inventive Principle:
Principle #23Feedback

3Productivity

If existing text classification algorithms are used, then processing speed is achieved, but semantic lexicon capacity is insufficient leading to poor decision support

Engineering Contradiction:
Improvecomputing speedVSAvoiddecision support effectiveness
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent performs preliminary action by pre-building a comprehensive semantic lexicon with industry-specific feature words and their relationships before classification tasks, and pre-calculating feature value weightings, thereby enabling both fast processing and reliable decision support without compromising computing speed

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20220147023A1Method and device for identifying industry classification of enterprise and particular pollutants of enterprise
Publication Date: 2022.05.12 CHINESE ACAD OF ENVIRONMENTAL PLANNING
  • US20220147023A1 patent drawing
  • US20220147023A1 patent drawing
  • US20220147023A1 patent drawing

AI summary

Disclosed is a method and device for identifying an industry classification of an enterprise and characteristic pollutants of the enterprise, wherein the method for identifying an industry classification of an enterprise comprises: acquiring information point data of a target enterprise; determining feature words of the information point data and feature values of the feature words according to a preset semantic lexicon, preset industry summary information and the information point data; and determining the industry classification to which the target enterprise belongs according to a preset industry classification prediction model and the feature values. Through implementing the present invention, the obtained feature values can effectively avoid interference of meaningless words, such that the industry classification to which the target enterprise belongs obtained from identification is more accurate.