Method and system applied to industry identification of small and micro enterprises
By integrating multi-dimensional data and design combination models, a small and micro enterprise industry classification label system is built, and the accuracy and efficiency of credit risk management in the existing technology is solved, and high-precision industry identification and automated management are achieved.
Patent Information
- Application Number
- CN202510384863.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-08
AI Technical Summary
The existing technology cannot accurately reflect the actual operating conditions of small and micro enterprises, resulting in inaccuracy and inefficiency of credit risk management, and relying on manual subjective analysis is prone to errors.
Build an end-to-end small and micro enterprise industry classification identification method based on structured invoice data, integrate industrial and commercial data, credit performance data and supply chain data, design and combine natural language models and timing models to form a refined industry label system, and optimize model prediction capabilities through supervised learning.
It has achieved the accuracy and automation of industry classification of small and micro enterprises, improved the accuracy and efficiency of credit risk management, ensured the objectivity and timeliness of classification results, and exceeded 95%.
Smart Images

Figure CN120278810A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a method and system for identifying small and micro enterprises in different industries. Background Art
[0002] In the current credit risk management of small and micro enterprises, accurate industry classification is the key. However, there is currently a lack of an objective-fact-based industry classification method for small and micro enterprises in the industry. The mainstream technologies mainly focus on the structured recognition scenarios of unstructured invoice data, and do not conduct deeper exploration of structured invoice data in the credit scenario. This results in the inability to accurately reflect the real business behavior of enterprises in the credit scenario, thereby affecting the accuracy of risk management.
[0003] The existing industrial and commercial information and the "National Economic Industry Classification" standard have limitations in practical applications. These information often cannot accurately reflect the actual business situation of enterprises, resulting in the need for subjective cross-analysis and verification of multiple information by manual for enterprise information verification. This process is not only inefficient but also error-prone, and cannot meet the high-precision requirements of credit risk management.
[0004] Specifically, the existing industry classification methods mainly classify enterprises into large, medium, small, micro and other types based on indicators such as the number of employees, operating income, and total assets. However, this method has the following problems in practical applications:
[0005] Incomplete information: Classifying only based on the above indicators cannot comprehensively reflect the actual business situation and industry characteristics of enterprises. For example, some small and micro enterprises may have core technologies or unique competitiveness in specific fields, but these characteristics are difficult to reflect in traditional classification methods.
[0006] Lagging data update: The update of enterprise information often lags behind, resulting in the classification results being unable to reflect the latest business situation of enterprises in a timely manner. This is particularly disadvantageous for credit risk management in a rapidly changing market environment.
[0007] Strong subjectivity: Manual cross-analysis and verification of multiple information are not only inefficient but also easily affected by subjective judgments, resulting in inaccurate classification results.
[0008] To solve these problems, it is necessary to innovatively construct an industry classification label system for small and micro enterprises in the credit scenario. This system should be based on the already structured invoice data and integrate heterogeneous industrial and commercial data. By designing a combined natural language model and a time series model, end-to-end identification of small and micro enterprise industry classification can be achieved, thereby obtaining accurate information for refined risk management in the industry. This method can not only improve the accuracy of classification but also realize the automation of the whole process model, significantly enhancing the efficiency and accuracy of credit risk management. Summary of the Invention
[0009] In view of this, the present invention provides a method and system for identifying industries of small and micro enterprises, so as to solve the problem that the existing technology often fails to accurately reflect the actual business conditions of enterprises, resulting in the need for manual subjective cross-analysis and verification of multiple-party information for enterprise information verification.
[0010] The technical solution adopted by the present invention is as follows:
[0011] A method for identifying industries of small and micro enterprises, including:
[0012] Step 1: Initially group historical enterprises according to the industry labels of the "National Economic Industry Classification", integrate the industrial and commercial data, credit performance data, judicial negative data, and supply chain data of the enterprises to form a multi-dimensional index portrait, and optimize the grouping of the enterprises according to the consistency of the portrait indexes to ensure that the enterprises within each group have similar business characteristics, and name the optimized grouping results with industry labels to form a new industry label system;
[0013] It should be noted that: The number of small and micro enterprises in China exceeds 70 million, and the roles of the industrial chain production division of labor are complex. The industry labels of the "National Economic Industry Classification" mainly target publicly listed enterprises and are difficult to accurately depict the industry types of all small and micro enterprises. Moreover, there is no unified industry classification and identification label type in the banking industry, and most of them are roughly classified according to the business expansion direction of the institution itself and the industrial and commercial registration information.
[0014] The specific steps of Step 1 include the following steps:
[0015] Step 1.1: Initially group historical enterprise customers according to the 396 middle-level and 913 small-level industry labels of the "National Economic Industry Classification";
[0016] It should be noted that: Although the labels of the "National Economic Industry Classification" cannot be directly applied to the small and micro enterprise customers of our bank, the meanings of its finest-grained labels are detailed and specific, and can be used as the basic direction for depicting enterprise types.
[0017] Step 1.2: Collect and integrally associate the credit performance data, enterprise industrial and commercial data, judicial negative data, and supply chain data of all historical enterprise customers to form a multi-dimensional index portrait including credit-granting indexes, risk indexes, in-loan indexes, and supply chain indexes;
[0018] Since there is no existing technology for reference in classification design, constructing portrait indexes to reflect the morphological performance after enterprise classification helps to invent industry labels that conform to the business expansion direction and credit management strategy of our bank. Some of the important professional indexes include: credit-granting rate, quota utilization rate, note term, interest rate, overdue rate, judicial negative proportion, average credit score, proportion of frozen customers, etc.
[0019] Step 1.3: Perform merging and splitting operations on the preliminary grouping in Step 1.1 according to the set criteria;
[0020] In Step 1.3, the said criteria include: the proportion of each customer group is not less than the set percentage threshold; merging of similar industries, merging of non-business industries, splitting of key industries; observing portrait indicators, and requiring indicators such as overdue risk and credit interest rate within the group to be consistent. For example: splitting out a low-risk subset from a high-risk grouping, splitting out a high-risk subset from a low-risk grouping, and considering merging subsets with risk consistency.
[0021] Step 1.4: After regrouping, update the indicator portrait in Step 1.2, and repeat the merging and splitting operations in Step 1.3 until the industry labels meet the requirements of the online usage party;
[0022] Step 1.5: Summarize the naming of industry labels for the final grouping result to complete the design of the enterprise industry label system.
[0023] Step 2: Associate and clean the industrial and commercial data, invoice data of the enterprise to be cleaned, and the categories of "National Economic Industry Classification" to ensure the integrity and accuracy of the data; based on the optimized industry label system in Step 1, combined with the associated and cleaned data, formulate keyword rules for marking to ensure that the rules are closely related to the actual business behavior of the enterprise, and apply the keyword rules for marking to the enterprises that can be hit to form model training samples;
[0024] It should be noted that: The industry labels completed in Step 1 are based on the bottom-up re-integration of the "National Economic Industry Classification", and there are still problems that do not represent the true business situation of the enterprise. Therefore, it is necessary to establish an association between the real production behavior and the real industry label through marking. In addition, supervised learning relies on labeled samples, enabling the model to optimize the prediction ability through the loss function, helping to capture data patterns and avoid overfitting. Therefore, it is necessary to initialize the labels of the training samples.
[0025] The said Step 2 specifically includes the following steps:
[0026] Step 2.1: Associate and clean the industrial and commercial data, invoice data of all historical small and micro enterprises, and the categories of "National Economic Industry Classification";
[0027] In Step 2.1, the industrial and commercial data includes the enterprise name and business scope, the invoice data includes the sales content text and takes the top sales commodity text in descending order of the invoiced amount, and the categories of "National Economic Industry Classification" include the first-level and fourth-level industry categories to ensure the integrity and accuracy of the data.
[0028] Step 2.2: Cross-reference the industrial and commercial data, invoice data, and categories in the "National Economic Industry Classification" in Step 2.1. Based on the optimized industry label system in Step 1, formulate keyword rules for marking. The keyword rules are related to the actual business production activities of the enterprise, where the invoice data is used to reflect the most real business production activities of the enterprise, and the industrial and commercial data and the categories in the "National Economic Industry Classification" are used as supplementary information sources;
[0029] Step 2.3: Apply the keyword rules for marking to the enterprises that can be hit to form model training samples.
[0030] Step 3: Based on the model training samples obtained in Step 2, design the text organizational structure and model structure to complete the model training for identifying industry types.
[0031] The specific steps of Step 3 include the following steps:
[0032] Step 3.1: Organize the marked samples generated in Step 2.3 into the model input text structure in the order of industrial and commercial text, "National Economic Industry Classification" text, and invoice sales text; fixing this structural form can ensure that the model captures the positional relationship between industrial and commercial data and invoice data during the learning process, and pays attention to the marking logic with industrial and commercial text as the basis and invoice text as the key.
[0033] Step 3.2: Based on the organized model input text structure in Step 3.1, use deep learning technology for end-to-end training. The training integrates the industrial and commercial text and invoice text of the enterprise, and combines the pre-marked industry labels to replace the keyword rule system formulated manually, and defines a new paradigm for identifying the industry category of the enterprise;
[0034] It should be noted that: The general NLP model Bert cannot meet the classification accuracy requirements on the data of this project. Therefore, theoretical design is carried out first, that is: the inter-textual checking relationship needs to use the BERT model based on Attention for the conversion between text and tensor, and the sequential relationship of non-homogeneous industrial and commercial and invoice data is suitable to use a time series model to solve the network memory problem. Then, an experimental comparison method is used to explore the optimal solution of the model architecture, such as: Bert+DNN (Deep Neural Network), Bert+Attention, Bert+RNN (Recurrent Neural Network), Bert+LSTM (Long Short-Term Memory). On the one hand, by deepening the network structure, the model can accommodate more information; on the other hand, combined with the characteristics of the text data of this project, the experimental effect evaluation proves that the Bert+LSTM architecture that combines the NLP model and the time series model obtains the highest accuracy effect and meets the model accuracy requirements.
[0035] Step 3.2 specifically includes the following steps:
[0036] Step 3.21: Establish a deep learning model that combines the architectures of Bert and LSTM. Determine to use the BERT model based on Attention to process the implicative relationship between texts, and adopt a time series model to solve the problem of the sequence relationship of non-homogeneous data between industry and commerce and invoices;
[0037] Step 3.22: Input the labeled sample X i text into the deep learning model, align the text lengths and add [CLS] and [SEP] tokens to represent the classification task and sentence segmentation. After the Bert model transforms the Pool layer tensor, it is input into the LSTM model. The output vector of the LSTM model is converted into a probability distribution σ(z i ) through the softmax function, as shown in the following formula:
[0038]
[0039] where z i is the output value of the i-th node, and K is the number of output nodes;
[0040] Step 3.23: Calculate the FocalLoss value of the loss function between the predicted label i corresponding to the maximum probability σ(z ) and the sample label Y i . Update the model parameters through gradient descent to minimize the loss function value, and finally obtain the prediction model model.
[0041] In Step 3.23, the loss function Focal Loss is shown in the following formula:
[0042] FL(p t ) = -(1 - p t ) γ log(p t )
[0043] where p t is the probability that the model predicts the sample to belong to the positive class, and γ is an adjustable hyperparameter. When γ = 0, the Focal Loss degenerates into the standard cross-entropy loss function. The larger the value of γ, the smaller the penalty of the Focal Loss for easily classified samples and the larger the penalty for difficult-to-classify samples. Therefore, the Focal Loss can make the model pay more attention to difficult-to-classify samples, thereby improving the performance of the model.
[0044] Step 3.3: Train and tune the parameters of the model to find the optimal solution.
[0045] Step 3.3 specifically includes the following steps:
[0046] Step 3.31: Divide the dataset into a training set, a validation set, and a test set according to the ratio of 6:2:2;
[0047] Step 3.32: Based on the text information and industry multi-classification labels of the training set, train the model by minimizing the Loss of the training set through the gradient descent algorithm;
[0048] Step 3.33: Monitor the change of Accuracy of the multi-classification labels on the validation set, and stop training when the Accuracy of the validation set reaches the maximum value;
[0049] Step 3.34: Evaluate the Accuracy level of the model results on the test set. The evaluation metrics of the test set are used to horizontally compare the baseline model (Bert) and all experimental group model architectures.
[0050] Step 4: Promote the industry identification of the model to all industrial and commercial small and micro enterprises, and the credit manager formulates corresponding application strategies.
[0051] A system applied to the industry identification of small and micro enterprises, including:
[0052] Label formation module: Initially group historical enterprises according to the industry labels in the "National Economic Industry Classification", integrate the industrial and commercial data, credit performance data, judicial negative data, and supply chain data of the enterprises to form a multi-dimensional index portrait. According to the consistency of the portrait indexes, optimize the grouping of the enterprises to ensure that the enterprises within each group have similar business characteristics, and name the industry labels for the optimized grouping results to form an industry label system;
[0053] Training sample acquisition module: Associate and clean the industrial and commercial data, invoice data, and categories in the "National Economic Industry Classification" of the enterprises to ensure the integrity and accuracy of the data; Based on the optimized industry label system in the label formation module, combine the associated and cleaned data, formulate keyword rules for labeling to ensure that the rules are closely related to the actual business behaviors of the enterprises, and apply the keyword rules for labeling to the enterprises that can be hit to form model training samples;
[0054] Model training module: Based on the model training samples obtained by the training sample acquisition module, design the text organizational structure and model structure to complete the model training that can identify industry types.
[0055] In summary, due to the adoption of the above technical solutions, the beneficial effects of the present invention are:
[0056] 1. The present invention constructs an industry classification label system for small and micro enterprises in the credit scenario, and based on the already structured invoice data, fuses heterogeneous industrial and commercial data. In view of the characteristics of the text information for industry classification, a combined natural language model and a time series model are designed, and an end-to-end industry classification recognition of small and micro enterprises is realized through a new model structure to obtain accurate information for refined risk management of the industry.
[0057] 2. The present invention solves the problem that there is no effective technical means in the industry to identify the true business production industry of enterprises. The "National Economic Industry Classification" reflects the subjective expression of the business owner on the business type of the enterprise, while the business production industry identified through the industrial and commercial enterprise name, business scope, and enterprise sales invoice content has objectivity, authenticity, timeliness, and accuracy.
[0058] 3. The present invention solves the problem that the existing natural language processing models are unable to handle the identification problem of this text data. A brand-new model architecture design and a loss function optimized for class imbalance are adopted, enabling the classification result of the model based on non-homologous and non-contextual industrial and commercial and invoice texts to achieve an accuracy of over 95%.
[0059] 4. The present invention simplifies the complex process of traditional industry classification based on keyword rules. The classification model has learning ability and generalization ability. Through model training with more than one hundred thousand pre-labeled enterprise samples, the classification recognition of all industrial and commercial enterprises with a total of more than 70 million can be achieved. BRIEF DESCRIPTION OF THE DRAWINGS
[0060] The present invention will be described by way of examples with reference to the accompanying drawings, where:
[0061] Figure 1 is a schematic diagram of the text organizational structure in Embodiment 1 of the present invention;
[0062] Figure 2 is a schematic diagram of the combined model structure in Embodiment 1 of the present invention;
[0063] Figure 3 is a flowchart in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0064] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some but not all of the embodiments of the present invention. The components of the embodiments of the present invention usually described and illustrated herein can be arranged and designed in various different configurations.
[0065] Accordingly, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.
[0066] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0067] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.
[0068] In the present invention, unless otherwise clearly specified and limited, the first feature being “on” or “under” the second feature may include direct contact between the first and second features, or may include indirect contact between the first and second features through additional features therebetween. Moreover, the first feature being “above”, “over” and “on top of” the second feature includes the first feature being directly above and obliquely above the second feature, or merely indicating that the horizontal height of the first feature is higher than that of the second feature. The first feature being “under”, “below” and “beneath” the second feature includes the first feature being directly below and obliquely below the second feature, or merely indicating that the horizontal height of the first feature is lower than that of the second feature.
[0069] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments may be combined with each other.
[0070] Embodiment 1
[0071] As Figures 1 - 3 shown, an identification method for small and micro enterprises is disclosed in an embodiment of the present invention, including:
[0072] Step 1: Initially group historical enterprises according to the industry labels in the "National Economic Industry Classification", integrate the industrial and commercial data, credit performance data, judicial negative data and supply chain data of the enterprises to form a multi-dimensional index portrait, optimize the grouping of the enterprises according to the consistency of the portrait indexes, ensure that the enterprises within each group have similarity in business characteristics, and name the optimized grouping results with industry labels to form a new industry label system;
[0073] Among them, it should be noted that: the number of small and micro enterprises in China exceeds 70 million, and the roles of industrial chain production and division of labor are complex. The industry labels in the "National Economic Industry Classification" mainly target publicly listed enterprises and are difficult to accurately depict the industry types of all small and micro enterprises. Moreover, there is no unified industry classification and recognition label type in the banking industry, and most are roughly classified according to the business directions of the institutions themselves and the industrial and commercial registration information.
[0074] Step 1 specifically includes the following steps:
[0075] Step 1.1: Initially group historical enterprise customers according to the 396 middle-level categories and 913 small-level category industry labels in the "National Economic Industry Classification";
[0076] It should be noted that: although the "National Economic Industry Classification" labels cannot be directly applied to the small and micro enterprise customers of our bank, the meanings of its finest-grained labels are detailed and specific, which can be used as the basic direction for depicting enterprise types.
[0077] Step 1.2: Collect, correlate, and integrate the credit performance data, enterprise industrial and commercial data, judicial negative data, and supply chain data of all historical enterprise customers to form a multi-dimensional indicator portrait including credit-granting indicators, risk indicators, in-loan indicators, and supply chain indicators;
[0078] Since there is no existing technology for reference in classification design, constructing portrait indicators to reflect the morphological performance after enterprise classification helps to invent industry labels that conform to the business directions and credit management strategies of our bank. Some of the important professional indicators include: credit-granting rate, quota utilization rate, promissory note term, interest rate, overdue rate, proportion of judicial negatives, average credit score, proportion of frozen customers, etc.
[0079] Step 1.3: Perform merging and splitting operations on the initial grouping in Step 1.1 according to the set criteria;
[0080] In Step 1.3, the criteria include: the proportion of each customer group is not less than the set percentage threshold; merge similar industries, merge non-business industries, and split key industries; observe the portrait indicators, and require indicators such as in-group overdue risk and credit-granting interest rate to be consistent. For example: split out low-risk subsets from high-risk groups, split out high-risk subsets from low-risk groups, and consider merging subsets with risk consistency.
[0081] After regrouping, update the indicator portrait in Step 1.2, and repeat the merging and splitting operations in Step 1.3 until the industry labels meet the requirements of the online usage parties;
[0082] Step 1.5: Summarize the naming of industry labels for the final grouping results to complete the design of the enterprise industry label system.
[0083] Step 2: Associate and clean the industrial and commercial data, invoice data, and categories of "National Economic Industry Classification" of the enterprises to be cleaned to ensure the integrity and accuracy of the data; Based on the optimized industry label system in Step 1, combined with the associated and cleaned data, formulate keyword rules for marking to ensure that the rules are closely related to the actual business operations of the enterprises, and apply the keyword rules for marking to the enterprises that can be hit to form model training samples;
[0084] It should be noted that: The industry labels completed in Step 1 are based on the bottom-up reintegration of the "National Economic Industry Classification", and there are still problems that do not represent the true business conditions of the enterprises. Therefore, it is necessary to establish an association between the real production behavior and the real industry label through marking. In addition, supervised learning relies on labeled samples, enabling the model to optimize its prediction ability through the loss function, helping to capture data patterns and avoid overfitting. Therefore, it is necessary to initialize the labels of the training samples.
[0085] The specific steps of Step 2 include the following steps:
[0086] Step 2.1: Associate and clean the industrial and commercial data, invoice data, and categories of "National Economic Industry Classification" of all historical small and micro enterprises;
[0087] In Step 2.1, the industrial and commercial data includes the enterprise name and business scope, the invoice data includes the sales content text and takes the top sales commodity text in descending order of the invoiced amount, and the categories of "National Economic Industry Classification" include the first-level and fourth-level industry categories to ensure the integrity and accuracy of the data.
[0088] Step 2.2: Cross the industrial and commercial data, invoice data, and categories of "National Economic Industry Classification" in Step 2.1, and based on the optimized industry label system in Step 1, formulate keyword rules for marking. The keyword rules are related to the real business production activities of the enterprises, where the invoice data is used to reflect the most real business production activities of the enterprises, and the industrial and commercial data and the categories of "National Economic Industry Classification" are used as supplementary information sources;
[0089] Step 2.3: Apply the keyword rules for marking to the enterprises that can be hit to form model training samples.
[0090] Step 3: Based on the model training samples obtained in Step 2, design the text organizational structure and model structure, and complete the model training that can identify industry types.
[0091] The specific steps of Step 3 include the following steps:
[0092] Step 3.1: Organize the labeled samples generated in Step 2.3 into the model text structure in the order of industrial and commercial texts, "National Economic Industry Classification" texts, and invoice sales texts; fixing this structural form can ensure that the model captures the positional relationship between industrial and commercial data and invoice data during the learning process, and pays attention to the labeling logic based on industrial and commercial texts and with invoice texts as the key. The text organizational structure is as Figure 1 shown.
[0093] Step 3.2: Based on the organized model text structure in Step 3.1, use deep learning technology for end-to-end training. The training integrates the enterprise's industrial and commercial texts and invoice texts, and combines the pre-labeled industry tags to replace the keyword rule system formulated manually, defining a new paradigm for identifying the enterprise's industry category;
[0094] It should be noted that: The general NLP model Bert cannot meet the classification accuracy requirements for the data in this project. Therefore, theoretical design is carried out first, that is: The logical relationship between texts needs to use the BERT model based on Attention for the conversion between text and tensor, and the sequential relationship of non-homogeneous data of industry and commerce and invoices is suitable to use a time series model to solve the network memory problem. Then, an experimental comparison method is used to explore the optimal solution of the model architecture. For example: Bert+DNN (Deep Neural Network), Bert+Attention, Bert+RNN (Recurrent Neural Network), Bert+LSTM (Long Short-Term Memory). On the one hand, by deepening the network structure, the model can accommodate more information; on the other hand, combined with the characteristics of the text data in this project, the experimental effect evaluation proves that the Bert+LSTM architecture combining the NLP model and the time series model obtains the highest accuracy effect and meets the model accuracy requirements.
[0095] The specific steps of Step 3.2 include the following steps:
[0096] Step 3.21: Establish a deep learning model combining the architectures of Bert and LSTM, determine to use the BERT model based on Attention to process the logical relationship between texts, and use a time series model to solve the sequential relationship problem of non-homogeneous data of industry and commerce and invoices;
[0097] Step 3.22: Input the labeled sample X i text into the deep learning model, align the text lengths and add [CLS] and [SEP] tokens to represent the classification task and sentence segmentation. After the Bert model converts the Pool layer tensor, it is input into the LSTM model, and the output vector of the LSTM model is converted into a probability distribution σ(z i ), as shown in the following formula:
[0098]
[0099] Among them, z i is the output value of the i-th node, and K is the number of output nodes;
[0100] Step 3.23: Calculate the predicted label corresponding to the maximum probability σ(z i ), and the FocalLoss value of the loss function between the predicted label and the sample label Y i . Update the model parameters through gradient descent to minimize the loss function value, and finally obtain the prediction model model. The combined model structure is as Figure 2 shown.
[0101] In Step 3.23, the loss function Focal Loss is shown as follows:
[0102] FL(p t ) = -(1 - p t ) γ log(p t )
[0103] Among them, p t is the probability that the model predicts that the sample belongs to the positive class, γ is an adjustable hyperparameter. When γ = 0, the Focal Loss degenerates into the standard cross-entropy loss function. The larger the value of γ, the smaller the penalty of the Focal Loss for easily classified samples and the larger the penalty for difficult-to-classify samples. Therefore, the Focal Loss can make the model pay more attention to difficult-to-classify samples, thereby improving the performance of the model.
[0104] Step 3.3: Train and tune the model to find the optimal solution.
[0105] The specific steps of the said Step 3.3 are as follows:
[0106] Step 3.31: Divide the data set into a training set, a validation set, and a test set according to the ratio of 6:2:2;
[0107] Step 3.32: Based on the text information of the training set and the multi-classification labels of the industry, train the model by minimizing the training set Loss through the gradient descent algorithm;
[0108] Step 3.33: Monitor the change of Accuracy of the multi-classification label on the validation set, and stop training when the validation set Accuracy reaches the maximum value;
[0109] Step 3.34: Evaluate the Accuracy level of the model results on the test set. The evaluation metrics of the test set are used for horizontal comparison between the baseline model (Bert) and all experimental group model architectures.
[0110] Step 4: Promote the industry identification of the model to all industrial and commercial small and micro enterprises, and the credit manager formulates corresponding application strategies.
[0111] Embodiment 2
[0112] Based on Embodiment 1, this embodiment proposes a system for industry identification of small and micro enterprises, including:
[0113] Label formation module: Initially group historical enterprises according to the industry labels in the "National Economic Industry Classification", integrate the industrial and commercial data, credit performance data, judicial negative data, and supply chain data of the enterprises to form a multi-dimensional index portrait. According to the consistency of the portrait indicators, optimize the grouping of the enterprises to ensure that the enterprises within each group have similar business characteristics, and name the industry labels for the optimized grouping results to form an industry label system;
[0114] Training sample acquisition module: Associate and clean the industrial and commercial data, invoice data, and categories in the "National Economic Industry Classification" of the enterprises to ensure the integrity and accuracy of the data; Based on the optimized industry label system in the label formation module, combine the associated and cleaned data to formulate a tagging keyword rule to ensure that the rule is closely related to the actual business behavior of the enterprises, and apply the tagging keyword rule to tag the enterprises that can be hit to form model training samples;
[0115] Model training module: Based on the model training samples obtained by the training sample acquisition module, design the text organizational structure and model structure to complete the model training that can identify industry types.
[0116] The circuits, electronic components, and modules involved are all existing technologies, which can be fully realized by those skilled in the art without further elaboration. The content protected by the present invention does not involve improvements to software and methods either.
[0117] In this specification, each embodiment is described in a progressive manner. The key points of each embodiment are the differences from other embodiments. The same or similar parts among the embodiments can be referred to each other.
[0118] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for identifying small and micro enterprise industries, characterized in that, Including: Step 1: Initially group historical enterprises according to the industry labels in the "National Economic Industry Classification", integrate the industrial and commercial data, credit performance data, judicial negative data, and supply chain data of the enterprises to form a multi-dimensional index portrait. Optimize the grouping of the enterprises according to the consistency of the portrait indexes to ensure that the enterprises within each group have similar business characteristics. Name the optimized grouping results with industry labels to form a new industry label system; Step 2: Associate and clean the industrial and commercial data, invoice data, and categories in the "National Economic Industry Classification" of the enterprises to ensure the integrity and accuracy of the data; Based on the optimized industry label system in Step 1, combine the associated and cleaned data to formulate a keyword marking rule to ensure that the rule is closely related to the actual business behavior of the enterprises. Apply the keyword marking rule to mark the enterprises that can be hit to form a model training sample; Step 3: Based on the model training sample obtained in Step 2, design the text organizational structure and model structure to complete the model training that can identify the industry type.
2. The method for identifying small and micro enterprise industries according to claim 1, wherein The specific steps of Step 1 include the following steps: Step 1.1: Initially group historical enterprise customers according to the 396 middle-class and 913 small-class industry labels in the "National Economic Industry Classification"; Step 1.2: Collect and associate and integrate the credit performance data, enterprise industrial and commercial data, judicial negative data, and supply chain data of all historical enterprise customers to form a multi-dimensional index portrait including credit-granting indexes, risk indexes, in-loan indexes, and supply chain indexes; Step 1.3: Perform merging and splitting operations on the initial grouping in Step 1.1 according to the set criteria; Step 1.4: After regrouping, update the index portrait in Step 1.2 and repeat the merging and splitting operations in Step 1.3 until the industry labels meet the requirements of the online users; Step 1.5: Summarize the industry label naming of the final grouping results to complete the design of the enterprise industry label system.
3. The method for identifying small and micro enterprise industries according to claim 2, characterized in that, In Step 1.3, the said criteria include: The proportion of each customer group is not less than the set percentage threshold; Merge similar industries, merge non-business industries, and split key industries; Observe the portrait indexes, and require that indexes such as overdue risk and credit-granting interest rate within the group be consistent.
4. The method for identifying small and micro enterprise industries according to claim 1, characterized in that, The specific steps of Step 2 include the following steps: Step 2.1: Associate and clean the industrial and commercial data, invoice data, and categories in the "National Economic Industry Classification" of all historical small and micro enterprises; Step 2.2: Cross the industrial and commercial data, invoice data, and categories in the "National Economic Industry Classification" in Step 2.
1. Based on the optimized industry label system in Step 1, formulate a keyword marking rule. The keyword rule is related to the real business production activities of the enterprises. Among them, the invoice data is used to reflect the most real business production activities of the enterprises, and the industrial and commercial data and the categories in the "National Economic Industry Classification" are used as supplementary information sources; Step 2.3: Apply the keyword marking rule to mark the enterprises that can be hit to form a model training sample.
5. The method for identifying small and micro enterprise industries according to claim 4, wherein, In step 2.1, the industrial and commercial data includes the enterprise name and business scope, the invoice data includes the sales content text and takes the top sales commodity text in descending order of the invoiced amount, and the categories of the "National Economic Industry Classification" include the first-level and fourth-level industry categories to ensure the integrity and accuracy of the data.
6. The method for identifying small and micro enterprise industries according to claim 1, characterized in that, The specific steps of step 3 are as follows: Step 3.1: Organize the labeled samples generated in step 2.3 into the in-model text structure in the order of industrial and commercial text, "National Economic Industry Classification" text, and invoice sales text. Step 3.2: Based on the in-model text structure organized in step 3.1, use deep learning technology for end-to-end training. The training integrates the enterprise's industrial and commercial text and invoice text, combines the pre-labeled industry labels, replaces the keyword rule system formulated manually, and defines a new paradigm for identifying the enterprise's industry category. Step 3.3: Train and tune the parameters of the model to find the optimal solution.
7. A method for identifying small and micro enterprise industries according to claim 6, characterized in that, The specific steps of step 3.2 are as follows: Step 3.21: Establish a deep learning model combining the architectures of Bert and LSTM, determine to use the BERT model based on Attention to process the cross-check relationship between texts, and adopt a time series model to solve the problem of the sequence relationship between non-homogeneous data of industry and commerce and invoices. Step 3.22: Input the labeled sample X i text into the deep learning model, align the text length and add [CLS] and [SEP] tokens to represent the classification task and sentence segmentation. After the Bert model transforms the Pool layer tensor, it is input into the LSTM model. The output vector of the LSTM model is converted into a probability distribution σ(z i ), as shown in the following formula: where z i is the output value of the i-th node, K is the number of output nodes, j represents traversing each label from 1 to K, and e means converting Z to a positive number (>0); Step 3.23: Calculate the predicted label corresponding to the maximum probability σ(z i ) and the sample label Y i to obtain the Focal Loss value of the loss function. Update the model parameters through gradient descent to minimize the loss function value, and finally obtain the prediction model model.
8. The method for identifying small and micro enterprise industries according to claim 6, wherein In step 3.23, the loss function Focal Loss is shown as follows: FL(p t ) = -(1 - p t ) γ log(p t ) where p t is the probability that the model predicts the sample to belong to the positive class, and γ is an adjustable hyperparameter. When γ = 0, Focal Loss degenerates into the standard cross-entropy loss function.
9. The method for identifying small and micro enterprise industries according to claim 6, characterized in that, The specific steps of step 3.3 are as follows: Step 3.31: Divide the data set into a training set, a validation set, and a test set according to the ratio of 6:2:
2. Step 3.32: Based on the text information and industry multi-classification labels of the training set, train the model by minimizing the training set Loss through the gradient descent algorithm. Step 3.33: Monitor the change of Accuracy of the multi-classification labels on the validation set, and stop training when the Accuracy of the validation set reaches the maximum value. Step 3.34: Evaluate the Accuracy level of the model results on the test set. The evaluation indicators of the test set are used to make a horizontal comparison between the baseline model (Bert) and all experimental group model architectures.
10. A system applied to the industry identification of small and micro enterprises, characterized in that, A method for identifying industries of small and micro enterprises as claimed in claims 1-9, comprising: Label formation module: Initially group historical enterprises according to the industry labels of the "National Economic Industry Classification", integrate the enterprise's industrial and commercial data, credit performance data, judicial negative data, and supply chain data to form a multi-dimensional index portrait. According to the consistency of the portrait indicators, optimize the grouping of enterprises to ensure that the enterprises within each group have similar business characteristics, and name the optimized grouping results with industry labels to form an industry label system. Training sample acquisition module: Associate and clean the enterprise's industrial and commercial data, invoice data, and "National Economic Industry Classification" categories to ensure the integrity and accuracy of the data; Based on the optimized industry label system in the label formation module, combine the associated and cleaned data to formulate a keyword rule for marking to ensure that the rule is closely related to the actual business behavior of the enterprise, and apply the keyword rule for marking to the enterprises that can be hit to form model training samples. Model training module: Based on the model training samples obtained by the training sample acquisition module, design the text organizational structure and model structure to complete the model training capable of identifying industry types.