Intelligent industry classification and coding method and device in social investigation

Through the method of hierarchical classification and multimodal semantic fusion, the efficiency and accuracy problems of industry coding in social surveys are solved, and efficient and accurate industry text classification is achieved to adapt to the characteristics of different industries.

CN120705319AActive Publication Date: 2025-09-26NANJING QILIANG INFORMATION TECHNOLOGY CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202511224013.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-29
Publication Date
2025-09-26
Estimated Expiration
2045-08-29

AI Technical Summary

Technical Problem

Existing technologies have problems with low efficiency, strong subjectivity and insufficient accuracy in the industry coding process in social surveys. Especially when facing emerging industries and complex descriptions, manual coding is inefficient and automated methods are difficult to achieve high-precision classification.

Method used

A hierarchical classification strategy is adopted to generate a set of middle-class candidates through the pre-trained BERT model. A small-category semantic matching dataset is constructed by combining international standard corpus and self-built vocabulary. Keyword matching, shallow semantics and deep semantic models are integrated for fusion calculation, and the final result is calibrated through business rules.

Benefits of technology

It achieves efficient and accurate industry text classification, reduces subjective differences, improves the classification accuracy and robustness of complex texts, and adapts to the characteristics of different industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705319A_ABST
    Figure CN120705319A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent industry classification and coding method and device in social investigation, and belongs to the technical field of natural language processing, machine learning and data classification. And the accuracy and efficiency of industry coding are remarkably improved by combining hierarchical classification and semantic matching technologies. According to the technical scheme, the method comprises the following steps: performing industry class judgment on a to-be-coded industry text based on a BERT model, and generating first five class candidates; for each middle class, in combination with an international standard and a self-established corpus, performing subclass semantic similarity calculation through keyword matching, shallow semantic matching and a deep semantic model; and finally, semantic weight and business analysis are integrated, and subclass codes with the highest recommendation degree are output. The device comprises a middle class judgment module, a small class data processing module, a multi-modal semantic matching module and a coding decision module. The method has the beneficial effects that through hierarchical classification strategy and multi-modal semantic fusion, the problems of fuzzy text classification and high manual dependence degree in the complex industry are solved, and the automation level of large-scale data processing is remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of natural language processing, machine learning, and data classification, and is particularly suitable for the automated classification and coding of unstructured industry texts in social surveys. The present invention specifically relates to an intelligent industry classification and coding method and device for social surveys. Background Art

[0002] In the field of social surveys and data analysis, industry coding is a key link in converting unstructured text into structured information. Its accuracy and efficiency directly affect the reliability and timeliness of subsequent research. Currently, industry coding mainly relies on manual operations or automated methods based on machine learning, but both have significant technical bottlenecks. Manual coding requires professionals to match texts one by one according to industry classification standards. When faced with emerging industries or complex descriptions, coders need to repeatedly review materials, which is inefficient and highly subjective, resulting in significant differences in the classification results of different coders for the same text. Although automated methods reduce manual intervention, when traditional end-to-end models directly predict subcategory codes, the complexity of the model increases sharply due to the need to process the high-dimensional space of thousands of subcategories. In addition, the semantic overlap between subcategories can easily cause confusion, making it difficult to achieve high-precision classification.

[0003] Existing automation solutions often use a single deep learning model (such as BERT and LSTM) for end-to-end prediction. This limitation lies in their over-reliance on global semantic representations and their neglect of the importance of explicit features and business rules. For example, relying solely on semantic similarity for niche scenarios such as "battery manufacturing" and "battery recycling" can easily lead to misjudgments. Furthermore, some technologies attempt to narrow the search scope through hierarchical classification, but their designs have significant flaws: the mid-class prediction stage often relies on a single result. Once the first-level classification is incorrect, subsequent sub-class matching will completely deviate from the correct path. For example, "intelligent driving system research and development" may be incorrectly classified as "manufacturing," resulting in the inability to subsequently match the correct code under "technology services."

[0004] An analysis of existing stratification methods reveals that they often utilize rule templates or single semantic models in the subcategory matching stage, lacking the deep integration of multi-dimensional features. While rule templates can capture explicit terminology, they lack generalization capabilities given the diversity and dynamic nature of industry terminology. While single semantic models can parse contextual semantics, they struggle to distinguish between nuanced business scenarios. These methods, failing to integrate the multimodal features of keyword matching, shallow semantics, and deep semantics, limit their classification accuracy for complex text and are unable to adapt to the specific needs of diverse industries.

[0005] To address the above issues, an intelligent encoding solution that takes into account both efficiency and accuracy is urgently needed. Summary of the Invention

[0006] The main purpose of the present invention is to overcome the defects of the existing technology, and to provide an intelligent industry coding method and device for social survey scenarios, so as to at least solve the above technical problems existing in the current industry coding. The present invention effectively breaks through the bottleneck of the existing technology through the innovative design of hierarchical classification and multimodal semantic fusion: generating multiple candidate sets in the mid-class prediction stage to avoid global deviations caused by single errors; integrating keyword matching, shallow semantic analysis and deep semantic parsing in small class matching, and adapting to different industry needs through dynamic weight adjustment. In addition, the business rule library is introduced to perform secondary calibration on the semantic matching results, which significantly improves the classification robustness of fuzzy text. Compared with the existing technology, this solution achieves accurate parsing of complex industry descriptions while maintaining efficient processing capabilities, providing a more reliable solution for the automated coding of large-scale social survey data.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] The first aspect of the present invention is to provide an intelligent industry coding method for social surveys. The method comprises: Based on the pre-trained BERT model, the industry category of the input industry text is judged and the top five category candidate sets are generated; For each middle-class candidate, we build a small-class semantic matching dataset by combining international standard corpus with our own industry vocabulary. Generate semantic similarity weights between industry text and subcategory data through the fusion calculation of keyword matching, shallow semantic model, and deep semantic model; Based on semantic weight and business rule analysis, the industry subcategory codes with the highest recommendation are sorted and output.

[0009] In the above method, the step of generating a candidate set of industry classes based on the pre-trained BERT model specifically includes: Construct a medium-dimension training corpus based on international industry classification standards and a self-built industry question database; The BERT model is used to train the corpus to generate an industry-class classifier; The industry text to be coded is input into the classifier, the probability distribution of each class is output, and the top five classes with the highest probability are selected as the candidate set.

[0010] In the above method, the step of constructing a sub-category semantic matching dataset specifically includes: For each candidate in the middle category, extract the description text and keywords of the corresponding subcategory from international standard documents and self-built industry vocabulary; Preprocessing the subcategory text, including word segmentation, stop word removal, and standardized expression; Generate a structured subcategory semantic matching dataset, including subcategory codes, text descriptions, and keyword labels.

[0011] In the above method, the implementation of the multimodal semantic matching algorithm includes the following sub-steps: Keyword matching: Extract keywords from industry texts and subcategory corpora based on the TF-IDF algorithm and calculate the intersection weight; Shallow semantic matching: Convert industry text and subcategory descriptions into TF-IDF vectors and calculate similarity scores using cosine similarity. Deep semantic matching: Use the pre-trained Sentence-BERT model to generate sentence vectors and calculate the cosine similarity score at the semantic level; The above three matching scores are fused according to the preset weights to generate a comprehensive semantic similarity weight.

[0012] In the above method, the step of determining the target code based on the semantic weight and business rule analysis includes: Dynamically adjust the weight distribution ratio of keywords, shallow semantics, and deep semantics based on business scenario requirements; Sort the comprehensive similarity scores of each subcategory candidate in descending order and select the top three candidate codes; Combined with manual review rules or business priorities, the final target industry subcategory code is determined from the candidate codes.

[0013] In the above method, the weight distribution ratio is: keyword matching weight 30%, shallow semantic matching weight 20%, deep semantic matching weight 50%.

[0014] The second aspect of the present invention is to provide an intelligent industry classification and coding device, characterized by comprising: Middle-class classification module: loads the pre-trained BERT model to perform middle-class probability prediction and candidate screening for industry texts; Subcategory corpus management module: stores international standard and self-built subcategory corpus data, and supports dynamic update and retrieval; Multimodal matching engine: Integrates keyword matching, shallow semantic matching, and deep semantic matching algorithms to output comprehensive similarity weights; Coding decision module: Generates and outputs the target industry code based on the weight fusion results and business rule library.

[0015] In the above method, the multimodal matching engine further comprises: Keyword extraction unit, which extracts core terms based on TF-IDF and regular expressions; Shallow semantic computing unit, calling the Scikit-learn library to complete TF-IDF vectorization and similarity calculation; The deep semantic modeling unit loads the Sentence-BERT model to generate sentence vectors and calculate the cosine distance.

[0016] The third aspect of the present invention is to provide a computer-readable storage medium, which stores a computer program. When the computer program is executed, it implements the intelligent industry classification and coding method in social surveys described in the present invention.

[0017] A fourth aspect of the present invention is to provide an electronic device, comprising: processor; Memory; The memory stores instructions that can be executed by the processor, and when the instructions are executed, the intelligent industry classification and coding method in the social survey described in the present invention is implemented.

[0018] Compared with the prior art, the intelligent industry classification and coding method and device for social surveys described in the present invention has the following technical features and beneficial effects: (1) Significantly improve classification accuracy and robustness Solve the problem of semantic ambiguity: Use a hierarchical classification strategy to avoid global encoding bias caused by single-class prediction errors.

[0019] Multimodal semantic fusion: Combining keyword matching, shallow semantic analysis, and deep semantic models, it solves the problem of misjudgment of complex industry texts by traditional single models.

[0020] Secondary calibration of business rules: Manual review rules and dynamic weight adjustment mechanism further correct semantic matching results and enhance the fault tolerance of ambiguous text.

[0021] (2) Significantly improve coding efficiency Reduce manual reliance: Automated processes replace traditional manual matching, especially for emerging industries and complex descriptions, reducing subjective differences and time costs.

[0022] Layered processing reduces complexity: a two-stage process of processing medium-sized classes first and small-sized classes second avoids the performance bottleneck of end-to-end models directly processing high-dimensional small-sized classes.

[0023] (3) Enhance industry adaptability Dynamic corpus support: Self-built industry lexicon continuously incorporates emerging terms to address the problem of insufficient generalization of traditional rule templates.

[0024] Flexible weight configuration: Dynamically adjust multimodal weights according to business scenarios to adapt to different industry characteristics.

[0025] Manual review interface: Results with a confidence level <80% are automatically pushed to manual review, balancing automation and reliability.

[0026] (4) Standardization and scalability Compatible with international standards: Corpus construction is based on international industry classification standards to ensure coding standardization.

[0027] Modular device design: Modules such as medium-class classification and multimodal matching engine support independent upgrades to facilitate technology iteration. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the present invention will be described in detail below in conjunction with the accompanying drawings and detailed embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. Among them: Figure 1 A flowchart of a method for classifying and coding intelligent industries in a social survey provided by an embodiment of the present invention; Figure 2 A multimodal semantic matching architecture diagram provided by an embodiment of the present invention; Figure 3 A schematic diagram of coupling dynamic weight adjustment and rule base provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0029] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention. Figure 1-Figure 3 Further explanation of the classification and coding methods and devices of intelligent industries in social surveys.

[0030] Example 1: Method Example

[0031] This embodiment provides a method for coding intelligent industries in social surveys, such as Figure 1 As shown, the embodiment includes the following steps: Step 1: Data Collection and Organization: Obtain a large amount of industry text data from social surveys. This data is generated by businesses and individuals when reporting industry information and is diverse and unstructured. Perform a preliminary screening of the collected raw data to remove obvious errors, duplications, or irrelevant data records. For example, records containing garbled characters or purely advertising content unrelated to the industry should be removed.

[0032] Step 2: Prepare middle-category data: Based on national standards, conduct a comprehensive cleanup of industry-related middle-category data. Carefully check the data for accuracy and completeness, correct any errors, and supplement any missing content. At the same time, using our own long-maintained industry vocabulary, we supplement the data with emerging vocabulary, professional terms, and nicknames within the industry. For example, for the "information technology" industry, we include emerging terms such as "big data" and "artificial intelligence," as well as nicknames such as "AI" and "big data technology." After this process, we construct a rich and accurate middle-category dimension corpus, as shown below: Table 1

[0033] Step 3: Model Training: Utilizing the prepared mid-category corpus data, the BERT model is trained on a targeted basis. During training, appropriate training parameters, such as the learning rate and number of training rounds, are set to optimize model performance. Through multiple iterations of training, the model fully learns the characteristics and patterns of mid-category text from different industries, enabling it to accurately identify and distinguish mid-category texts across various industries. The result is a fully trained model that can be used for industry mid-category judgment.

[0034] Step 4: Determine the middle class: Taking the example of "a company specializing in the research, development, and production of new solar panels," this text is fed into the trained BERT model. The model analyzes and calculates the input text, outputting the probability distribution of each middle class. From this output, the top five middle classes with the highest probability are selected as the candidate middle class set. In this example, the top five possible middle classes include "manufacturing," "science and technology services," "electricity, heat, gas, and water production and supply," "wholesale and retail," and "scientific research and technical services." Examples are shown below: Table 2

[0035] Step 5: Subcategory Data Processing: For each candidate middle category, such as "Manufacturing" and "Technology Services," detailed information about the corresponding subcategory is extracted from national standard documents and internally maintained industry data. For "Manufacturing," descriptive text and keywords for subcategories such as "Battery Manufacturing," "New Energy Equipment Manufacturing," and "Electrical Machinery and Equipment Manufacturing" are extracted. For "Technology Services," relevant data for subcategories such as "Technology R&D Services," "Technology Promotion Services," and "Information Technology Consulting Services" are extracted.

[0036] Preprocess these subclass texts, including word segmentation to split sentences into individual words; remove stop words such as "de", "le", "zai", etc. which have no practical meaning; unify term expressions and standardize terms with different expressions but the same meaning. For example, unify "solar panel" and "photovoltaic panel" into "solar panel". After processing, generate a structured subclass semantic matching dataset, and each dataset contains subclass codes, detailed text descriptions, and keyword tags. The example is as follows: Table 3

[0037] Step 6, semantic judgment: Step 6.1, keyword matching: Based on the TF - IDF algorithm, extract the keywords of the enterprise description text and the subclass corpus respectively. For the text of "an enterprise focusing on the research and development and production of new solar panels", extract keywords such as "solar panel", "research and development", "production", "new", etc. Then calculate the intersection weights of these keywords and the keywords in each subclass corpus. If the subclass corpus of "new energy equipment manufacturing" contains keywords such as "solar panel" and "production", the intersection weight will increase accordingly.

[0038] Step 6.2, shallow semantic matching: Convert the enterprise description text and the subclass description into TF - IDF vectors, and calculate the similarity score between the two through cosine similarity. For example, the cosine similarity between the subclass description of "new energy equipment manufacturing" and the enterprise text in the TF - IDF vector space is relatively high, indicating that they are relatively similar in shallow semantics, that is, from the perspective of the distribution and frequency of words, they have a certain degree of relevance.

[0039] Step 6.3, deep semantic matching: Use the pre - trained Sentence - BERT model to generate sentence vectors for the enterprise description text and the subclass description, and then calculate the cosine similarity score at the semantic level. This model can understand the deep semantic relationships in the text. For example, it can recognize the semantic similarity between "solar panel" and "photovoltaic panel". The score obtained in this way can more accurately reflect the semantic association between texts.

[0040] Step 6.4, comprehensive weight calculation: The three matching scores are combined according to preset weights. Assuming a keyword matching weight of 30%, a shallow semantic matching weight of 20%, and a deep semantic matching weight of 50%, a weighted calculation is performed to determine the comprehensive semantic similarity weight between each subcategory and the company description text. For example, the "New Energy Equipment Manufacturing" subcategory has a keyword matching score of 80, a shallow semantic matching score of 70, and a deep semantic matching score of 90. The comprehensive semantic similarity weight is 80 × 30% + 70 × 20% + 90 × 50% = 83.

[0041] The final result example is as follows: Table 4

[0042] Step 7, subcategory judgment: Combining the comprehensive weights derived from semantic analysis with business rule analysis, we rank the candidate subcategories. Considering that the company explicitly mentions "R&D and production," its business focus is on the actual manufacturing of products. The "New Energy Equipment Manufacturing" subcategory not only scores highly in semantic matching but also aligns with the company's business priorities. Therefore, the "New Energy Equipment Manufacturing" subcategory data is ranked first, identified as the closest subcategory data, and the corresponding industry subcategory code is output. In this example, the corresponding industry subcategory code for "New Energy Equipment Manufacturing" might be "3865" (assuming this code represents this subcategory in the coding system), completing the classification and coding of the company's industry text.

[0043] Example 2: Device Example

[0044] This embodiment provides an intelligent industry coding device for social surveys, such as Figure 3 As shown, it includes the following modules: The text preprocessing module uses regular expression matching and Jieba word segmentation tools to clean the input text, remove redundant characters, and generate standardized text; The medium-class classification module, powered by an NVIDIA T4 GPU, loads a pre-trained BERT model and outputs the top five medium-class candidates. The multimodal matching engine module uses the TF-IDF algorithm to extract keywords and the deep semantic model Sentence-BERT to perform keyword, shallow semantic, and deep semantic matching to generate a comprehensive score. The dynamic decision module adjusts the weight according to business rules, outputs the final code, and supports dynamic updates of the API; The manual review interface module pushes results with a confidence level of <80% to the manual review queue.

[0045] Example 3: Computer Program Product Example

[0046] This embodiment is a computer program product that can execute the above-described method embodiment. The program instructions contained in the computer program product, when executed, can perform the steps of the above-described method embodiment. The computer program product can be written in one or more programming languages, such as C, C++, Java, Python, etc. The program code can be executed in whole or in part, and can be executed on a local computer device, a remote computing device, or a server.

[0047] Embodiment 4: Computer-readable storage medium embodiment

[0048] This embodiment is a computer-readable storage medium. The program instructions stored therein, when executed, are capable of executing the steps of the above-described method embodiments. The computer-readable storage medium may be a single readable medium or a combination of multiple readable media. The readable medium may be a readable signal medium or a readable storage medium. Examples of readable signal media include, but are not limited to, infrared, magnetic, and semiconductor media; and examples of readable storage media include, but are not limited to, hard disks, portable disks, card-type memories (such as SD cards), read-only memory (ROM), and random access memory (RAM).

[0049] The above implementation is only a preferred embodiment of the intelligent industry coding method and device in the social survey of the present invention, and does not limit the form and scope of the present invention. Any disassembly, recombination, equivalent structure and equivalent process made using the contents of the description and drawings of the present invention are also included in the patent protection scope of the present invention.

[0050] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. A method for intelligent industry classification and coding in social surveys, characterized by: For social survey scenarios, the steps are as follows: S1. Based on the pre-trained BERT model, the industry category of the input industry text is judged and the top five category candidate sets are generated; S2. For each middle-class candidate, we combine international standard corpus and industry vocabulary to build a small-class semantic matching dataset; S3. Generate semantic similarity weights between industry text and subcategory data through keyword matching, shallow semantic model, and deep semantic model fusion calculation; S4. Based on semantic weight and business rule analysis, sort and output the industry subcategory codes with the highest recommendation.

2. The intelligent industry classification and coding method in social survey according to claim 1 is characterized in that: The specific steps of step S1 are as follows: S11. Construct a training corpus of medium-level dimensions based on international industry classification standards and industry question database; S12. Using the BERT model to train the corpus to generate an industry-class classifier; S13. Input the industry text to be coded into the classifier, output the probability distribution of each class, and select the top five classes with the highest probability as the candidate set.

3. The intelligent industry classification and coding method in social survey according to claim 1 is characterized in that: In step S2, the step of constructing a sub-category semantic matching dataset specifically includes: S21. For each candidate middle category, extract the description text and keywords of the corresponding subcategory from international standard documents and industry thesaurus; S22, preprocessing the subcategory text, including word segmentation, stop word removal and standardized expression; S23. Generate a structured subcategory semantic matching dataset, including subcategory codes, text descriptions, and keyword tags.

4. The intelligent industry classification and coding method in social survey according to claim 1 is characterized in that: In step S3, the implementation of the multimodal semantic matching algorithm includes the following sub-steps: Keyword matching: Extract keywords from industry texts and subcategory corpora based on the TF-IDF algorithm and calculate the intersection weight; Shallow semantic matching: Convert industry text and subcategory descriptions into TF-IDF vectors and calculate similarity scores using cosine similarity. Deep semantic matching: Use the pre-trained Sentence-BERT model to generate sentence vectors and calculate the cosine similarity score at the semantic level; The above three matching scores are fused according to the preset weights to generate a comprehensive semantic similarity weight.

5. The intelligent industry classification and coding method in social survey according to claim 1 is characterized in that: In step S4, the step of determining the target code based on the semantic weight and business rule analysis includes: S41. Dynamically adjust the weight distribution ratio of keywords, shallow semantics, and deep semantics according to business scenario requirements; S42, sorting the comprehensive similarity scores of each sub-category candidate in descending order, and selecting the top three candidate codes; S43. Determine the final target industry subcategory code from the candidate codes based on manual review rules or business priorities.

6. The intelligent industry classification and coding method in social survey according to claim 4 is characterized in that: The weight distribution ratio is: keyword matching weight 30%, shallow semantic matching weight 20%, deep semantic matching weight 50%.

7. An intelligent industry classification and coding device for social surveys, characterized in that: include: Middle-class classification module: loads the pre-trained BERT model to perform middle-class probability prediction and candidate screening for industry texts; Subcategory corpus management module: stores international standard and self-built subcategory corpus data, and supports dynamic update and retrieval; Multimodal matching engine module: integrates keyword matching, shallow semantic matching, and deep semantic matching algorithms, and outputs comprehensive similarity weights; Coding decision module: Generates and outputs the target industry code based on the weight fusion results and business rule library.

8. The intelligent industry classification and coding device for social surveys according to claim 7 is characterized in that: The multimodal matching engine module includes: Keyword extraction unit, which extracts core terms based on TF-IDF and regular expressions; Shallow semantic computing unit, calling the Scikit-learn library to complete TF-IDF vectorization and similarity calculation; The deep semantic modeling unit loads the Sentence-BERT model to generate sentence vectors and calculate the cosine distance.

9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which, when executed, implements the intelligent industry classification and coding method in social surveys described in any one of claims 1 to 6.

10. An electronic device, characterized in that: include: processor; Memory; The memory stores instructions that can be executed by the processor, and when the instructions are executed, they implement the intelligent industry classification and coding method in social surveys described in any one of claims 1-6.

Citation Information

Patent Citations

  • Automatic question answering method and device based on multi-semantic matching, equipment and medium

    CN114090747A

  • Industry category prediction method, electronic equipment and computer storage medium

    CN115577838A

  • Coding method, device and system and storage medium

    CN116502603A

  • Industrial information classification method and device, computer equipment and storage medium

    CN116975743A

  • Taxpayer industry classification method based on noise label learning

    WO2022178919A1