Long-tail distribution data classification method, training method, device, equipment and medium

By using a model containing multiple neural network levels, feature extraction and fusion of multi-label long-tail distribution data is solved, and the problem of poor classification effect of multi-label long-tail distribution data in the existing technology is achieved, achieving more efficient classification effects.

CN120011554APending Publication Date: 2025-05-16CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411455964.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively classify multi-label long-tail distribution data, resulting in poor classification effect.

Method used

A neural network with the ability to classify a minority data, a second neural network with the ability to classify a minority data and a majority data, and a fusion network and a classification network are used to extract and fuse features of multi-label long-tail distribution data, and generate classification features through attention mechanism and maximum pooling operation, and finally classify prediction is performed.

Benefits of technology

The classification efficiency of multi-label long-tail distribution data is improved, the processing ability of a few types of data is enhanced, and the classification effect is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011554A_ABST
    Figure CN120011554A_ABST
Patent Text Reader

Abstract

The invention provides a long-tail distribution data classification method and device, a model training method and device, equipment and a medium. The classification method comprises the steps of obtaining a text of to-be-classified multi-label long-tail distribution data; the text is input into a trained neural network model for classification prediction, the classification of the text is obtained, and the neural network model comprises a first neural network with minority class data classification capability, a second neural network with minority class data and majority class data classification capability, a fusion network and a classification network. The first neural network and the second neural network are respectively connected with the fusion network, and the fusion network is connected with the classification network. According to the method, the multi-label long-tail distribution data is classified through the network in the two stages in the trained neural network model, so that the network convergence speed is increased, more classification information is incorporated, and the classification efficiency of the multi-label long-tail distribution data is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a classification method for long-tail distribution data, a model training method, a device, an electronic device and a computer-readable storage medium. Background Art

[0002] Long-tailed distribution data: It is a common data distribution, or a skewed distribution, which means that a few categories (also called head categories) contain a large number of samples, while most categories (also called tail categories) have only a very small number of samples. In other words, the proportion of samples in most categories is at a low level, while the proportion of samples in a few categories is very high.

[0003] In the related art, for the classification task of multi-label long-tail distribution data, classification is usually carried out by data method, model method or post-processing method. However, if starting from the data method, the data distribution is usually rebalanced by oversampling minority class data and undersampling majority class data. However, this method treats each sample equally, ignoring the different degrees of difficulty in learning different instances (samples) in minority class samples; and the repeated sampling method also changes the original data distribution, which is easy to distort the decision boundary. From the perspective of deep learning models, it is necessary to start from the network structure design, usually separating the representation learning and classification learning stages, and training the feature extraction module and classification module separately. However, this method directly splits the end-to-end learning, which usually causes the problem of poor migration effect. The post-processing method is to directly modify the model output in the inference stage, and improve the processing preference for minority samples by adjusting the weight of the softmax input value. However, this method lacks a comprehensive understanding of the data and usually has poor results.

[0004] Therefore, how to solve the classification problem of multi-label long-tail distribution data is a problem that needs to be solved. Summary of the invention

[0005] The present invention provides a classification method, model training method, device, electronic device and computer-readable storage medium for long-tail distribution data, so as to at least solve the problem that the multi-label long-tail distribution data cannot be effectively classified in the related art, resulting in poor classification effect of the multi-label long-tail distribution data. The technical solution of the present invention is as follows:

[0006] According to a first aspect of an embodiment of the present invention, a classification method for long-tail distribution data is provided, comprising:

[0007] Get the text of multi-label long-tail distribution data to be classified;

[0008] The text is input into a trained neural network model for classification prediction to obtain the category of the text, wherein the neural network model includes: a first neural network with the ability to classify minority class data, a second neural network with the ability to classify minority class data and majority class data, a fusion network and a classification network, the first neural network and the second neural network are respectively connected to the fusion network, and the fusion network is connected to the classification network.

[0009] Optionally, inputting the text into a trained neural network model for classification prediction to obtain the category of the text includes:

[0010] After the text is input into the first neural network in the trained neural network model for processing, it is respectively input into the first linear change matrix and the second linear change matrix for feature extraction to obtain the corresponding first feature matrix and the second feature matrix;

[0011] Input the text into a second neural network in a trained neural network model for feature extraction to obtain a third feature matrix; wherein the first neural network and the second neural network both have the same bidirectional long short-term memory network LSTM structure;

[0012] Using the attention mechanism on the fusion network in the trained neural network model, perform a maximum pooling operation on the first feature matrix, the second feature matrix, and the third feature matrix after feature fusion to obtain a first fusion feature;

[0013] Performing a maximum pooling operation on the third feature matrix output by the second neural network, and adding the maximum pooled features to the first fusion features to obtain classification features;

[0014] The classification features are input into the classification network in the trained neural network model for classification prediction to obtain the category of the text.

[0015] Optionally, the method further comprises: pre-training the neural network model in the following manner:

[0016] Obtain a training set, wherein the training set includes: minority class data and labels of the minority class data; and majority class data and labels of the majority class data;

[0017] Based on the minority class data and the labels of the minority class data, a bidirectional long short-term memory network LSTM is trained for multi-task binary classification using a meta-learning method to obtain a first neural network with a small number of sample classification capabilities;

[0018] Based on the parameters of the first neural network, a second neural network is constructed, wherein the first neural network and the second neural network both adopt the same bidirectional long short-term memory network LSTM structure;

[0019] Inputting the majority class data into the first neural network and the second neural network respectively for multi-label classification training to obtain two features output by the first neural network and one feature output by the second neural network;

[0020] All features output by the first neural network and the second neural network are fused through an attention mechanism, and a maximum pooling operation is performed on the fused features to obtain fused features;

[0021] After performing a maximum pooling operation on the features output by the second neural network, the features are added to the fusion features to obtain classification features;

[0022] Inputting the classification features into a classification network for classification prediction training to obtain training values;

[0023] Using a loss function, the labels of the majority class data and the training values ​​are calculated to obtain a loss value;

[0024] The parameters of the second neural network are updated through a back-propagation mechanism based on the loss value, and after multiple iterations, the second neural network has the ability to classify long-tail distribution multi-labels, thereby obtaining a trained neural network model.

[0025] Optionally, the performing multi-task binary classification training on the minority class data by using a meta-learning method to obtain a first neural network with a small sample classification capability includes:

[0026] Obtain training data, including text and labels corresponding to the text;

[0027] Determining the minority class data text in the training data set according to the set long-tail distribution threshold;

[0028] For the minority class data texts, construct a binary classification data set according to the labels corresponding to the minority class data texts one by one;

[0029] For the binary classification data set, the MAML meta-learning method is used to perform multi-task binary classification training on the bidirectional long short-term memory network LSTM to obtain a first neural network with few-sample classification capability.

[0030] Optionally, the binary classification data set is trained by using a MAML meta-learning method to perform multi-task binary classification training on a bidirectional long short-term memory network LSTM to obtain a first neural network with a small sample classification capability, including:

[0031] Selecting text and corresponding labels of the first batch of samples and text and corresponding labels of the second batch of samples from the binary classification data set;

[0032] Based on the texts and corresponding labels of the selected first batch of samples, calculate a first loss value of a bidirectional long short-term memory network LSTM on the first batch of samples;

[0033] Based on the first loss value on the first batch of samples, the gradient descent method is used to calculate the LSTM parameters after gradient update;

[0034] Based on the texts and corresponding labels of the selected second batch of samples, a second loss value of the LSTM after the LSTM parameters are updated on the second batch of samples is calculated through back propagation;

[0035] Repeat the above calculation steps multiple times, and accumulate the second loss values ​​obtained multiple times to obtain the meta-learning loss value;

[0036] The meta-learning loss value is used to perform a meta-back propagation to obtain the trained parameters of the LSTM, and the LSTM based on the trained LSTM parameters is called the first neural network with few-sample classification capability.

[0037] According to a second aspect of an embodiment of the present invention, a training method for a neural network model is provided, comprising:

[0038] Obtain a training set, wherein the training set includes: minority class data and labels of the minority class data; and majority class data and labels of the majority class data;

[0039] Based on the minority class data and the labels of the minority class data, a bidirectional long short-term memory network LSTM is trained for multi-task binary classification using a meta-learning method to obtain a first neural network with a small number of sample classification capabilities;

[0040] Based on the parameters of the first neural network, a second neural network is constructed, wherein the first neural network and the second neural network both adopt the same bidirectional long short-term memory network LSTM structure;

[0041] Inputting the majority class data into the first neural network and the second neural network respectively for multi-label classification training to obtain two features output by the first neural network and one feature output by the second neural network;

[0042] All features output by the first neural network and the second neural network are fused through an attention mechanism, and a maximum pooling operation is performed on the fused features to obtain fused features;

[0043] After performing a maximum pooling operation on the features output by the second neural network, the features are added to the fusion features to obtain classification features;

[0044] Inputting the classification features into a classification network for classification prediction training to obtain training values;

[0045] Using a loss function, the labels of the majority class data and the training values ​​are calculated to obtain a loss value;

[0046] The parameters of the second neural network are updated through a back-propagation mechanism based on the loss value, and after multiple iterations, the second neural network has the ability to classify long-tail distribution multi-labels, thereby obtaining a trained neural network model.

[0047] According to a third aspect of an embodiment of the present invention, a classification device for long-tail distribution data is provided, comprising:

[0048] An acquisition module is used to obtain the text of multi-label long-tail distribution data to be classified;

[0049] The prediction module is used to input the text into a trained neural network model for classification prediction to obtain the category of the text, wherein the neural network model includes: a first neural network with minority class data classification capability, a second neural network with minority class data and majority class data classification capability, a fusion network and a classification network, the first neural network and the second neural network are respectively connected to the fusion network, and the fusion network is connected to the classification network.

[0050] Optionally, the prediction module includes:

[0051] A first feature extraction module is used to input the text into the first neural network in the trained neural network model for processing, and then input it into the first linear change matrix and the second linear change matrix for feature extraction to obtain the corresponding first feature matrix and the second feature matrix;

[0052] A second feature extraction module is used to input the text into a second neural network in a trained neural network model for feature extraction to obtain a third feature matrix; wherein the first neural network and the second neural network both have the same bidirectional long short-term memory network LSTM structure;

[0053] A feature fusion module, used to use an attention mechanism to perform a maximum pooling operation on the first feature matrix, the second feature matrix and the third feature matrix on a fusion network in a trained neural network model to obtain a first fusion feature;

[0054] A classification feature determination module, used for performing a maximum pooling operation on the third feature matrix output by the second neural network, and adding the maximum pooled feature to the first fusion feature to obtain a classification feature;

[0055] The classification prediction module is used to input the classification features into the classification network in the trained neural network model to perform classification prediction and obtain the category of the text.

[0056] Optionally, the device also includes: a training module, used for pre-training the neural network model.

[0057] Optionally, the training module includes:

[0058] A training set acquisition module is used to acquire a training set, wherein the training set includes: minority class data and labels of the minority class data; and majority class data and labels of the majority class data;

[0059] A first training module is used to perform multi-task binary classification training on a bidirectional long short-term memory network LSTM by using a meta-learning method based on the minority class data and the labels of the minority class data, so as to obtain a first neural network with a small number of sample classification capabilities;

[0060] A construction module, used to construct a second neural network based on the parameters of the first neural network, wherein the first neural network and the second neural network both adopt the same bidirectional long short-term memory network LSTM structure;

[0061] A second training module is used to input the majority class data into the first neural network and the second neural network respectively for multi-label classification training to obtain two features output by the first neural network and one feature output by the second neural network;

[0062] A fusion module, used to fuse all features output by the first neural network and the second neural network through an attention mechanism, and perform a maximum pooling operation on the fused features to obtain fused features;

[0063] A feature determination module, configured to perform a maximum pooling operation on the features output by the second neural network and add the features to the fusion features to obtain classification features;

[0064] A classification prediction module, used for inputting the classification features into a classification network for classification prediction training to obtain training values;

[0065] A loss determination module, used to calculate the labels of the majority class data and the training values ​​using a loss function to obtain a loss value;

[0066] The iterative module is used to update the parameters of the second neural network through a back-propagation mechanism based on the loss value, and after multiple iterations, the second neural network has the long-tail distribution multi-label classification capability to obtain a trained neural network model.

[0067] Optionally, the first training module includes:

[0068] A data acquisition module, used to acquire training data, including text and labels corresponding to the text;

[0069] A text determination module, used to determine the minority class data text in the training data set according to a set long-tail distribution threshold;

[0070] A binary classification data determination module, for constructing a binary classification data set for the minority class data texts one by one according to the labels corresponding to the minority class data texts;

[0071] The binary classification training module is used to perform multi-task binary classification training on the bidirectional long short-term memory network LSTM using the MAML meta-learning method on the binary classification data set to obtain a first neural network with a small number of sample classification capabilities.

[0072] Optionally, the binary classification training module includes:

[0073] A selection module, used to select text and corresponding labels of the first batch of samples, and text and corresponding labels of the second batch of samples from the binary classification data set;

[0074] A first loss calculation module, used to calculate a first loss value of a bidirectional long short-term memory network LSTM on the first batch of samples based on the texts and corresponding labels of the selected first batch of samples;

[0075] A parameter calculation module, used for calculating the LSTM parameters after gradient update by using a gradient descent method based on the first loss value on the first batch of samples;

[0076] A second loss calculation module, used for calculating, based on the texts and corresponding labels of the selected second batch of samples, a second loss value of the LSTM after parameter update on the second batch of samples through back propagation;

[0077] An accumulation module, used to repeat the above calculation steps multiple times, and accumulate the second loss values ​​obtained multiple times to the meta-learning loss;

[0078] A meta-back propagation module is used to perform a meta-back propagation using the meta-learning loss to obtain the trained parameters of the LSTM, and the trained LSTM is called the first neural network with few-sample classification capability.

[0079] According to a fourth aspect of an embodiment of the present invention, there is provided a training device for a neural network model, comprising:

[0080] A training set acquisition module is used to acquire a training set, wherein the training set includes: minority class data and labels of the minority class data; and majority class data and labels of the majority class data;

[0081] A first training module, based on the minority class data and the labels of the minority class data, performs multi-task binary classification training on a bidirectional long short-term memory network LSTM by using a meta-learning method to obtain a first neural network with a small number of sample classification capabilities;

[0082] A construction module is used to construct a second neural network based on the parameters of the first neural network, wherein the first neural network and the second neural network both adopt the same bidirectional long short-term memory network LSTM structure;

[0083] A second training module, inputting the majority class data into the first neural network and the second neural network respectively for multi-label classification training, to obtain two features output by the first neural network and one feature output by the second neural network;

[0084] A fusion module, which fuses all features output by the first neural network and the second neural network through an attention mechanism, and performs a maximum pooling operation on the fused features to obtain fused features;

[0085] a feature determination module, performing a maximum pooling operation on the features output by the second neural network and adding the features to the fusion features to obtain classification features;

[0086] A classification prediction module inputs the classification features into a classification network for classification prediction training to obtain a training value;

[0087] A loss determination module, using a loss function to calculate the labels of the majority class data and the training values ​​to obtain a loss value;

[0088] The iterative module updates the parameters of the second neural network through a back-propagation mechanism based on the loss value, and after multiple iterations, the second neural network has the long-tail distribution multi-label classification capability, thereby obtaining a trained neural network model.

[0089] According to a fifth aspect of an embodiment of the present invention, there is provided an electronic device, including:

[0090] processor;

[0091] a memory for storing instructions executable by the processor;

[0092] Wherein, the processor is configured to execute the instructions to implement the classification method of long-tail distribution data as described above or the training method of the neural network model as described above.

[0093] According to a sixth aspect of an embodiment of the present invention, a computer-readable storage medium is provided. When instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the classification method of long-tail distribution data as described above or the training method of the neural network model as described above.

[0094] According to a seventh aspect of an embodiment of the present invention, a computer program product is provided, comprising a computer program or instructions, which, when executed by a processor of an electronic device, implements the classification method for long-tail distribution data as described above or the training method for a neural network model as described above.

[0095] The technical solution provided by the embodiments of the present invention brings at least the following beneficial effects:

[0096] In an embodiment of the present invention, the text of the multi-label long-tail distribution data to be classified is input into the first neural network and the second neural network in the trained neural network model for feature extraction. The features output by the first neural network are respectively subjected to two different linear change matrices to obtain the corresponding first feature matrix and the second feature matrix; the features output by the second neural network are the third feature matrix. After that, the attention mechanism is adopted to fuse the above three feature matrices to obtain a fused feature matrix, and the fused feature matrix is ​​subjected to the maximum pooling operation to obtain the fused features; in the sequence length dimension, the third feature matrix output by the second neural network and the fused feature matrix are respectively subjected to the maximum pooling operation and then added to obtain the features input by the classification network, which are called classification features. Finally, the classification features are input into the classification network for classification prediction processing to obtain the category to which the final text belongs. The embodiment of the present invention classifies the multi-label long-tail distribution data through the trained neural network model, thereby improving the classification efficiency of the multi-label long-tail distribution data.

[0097] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention, and do not constitute an improper limitation of the present invention. In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings required for use in the embodiments or the prior art description will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0099] Figure 1It is a flow chart of a classification method for long-tail distribution data provided by an embodiment of the present invention.

[0100] Figure 2 It is a schematic diagram of the structure of a neural network provided by an embodiment of the present invention.

[0101] Figure 3 It is a schematic diagram of multi-label classification training using forward propagation provided by an embodiment of the present invention.

[0102] Figure 4 It is a flow chart of a training method of a neural network model provided by an embodiment of the present invention.

[0103] Figure 5 It is a block diagram of a classification device for long-tail distribution data provided by an embodiment of the present invention.

[0104] Figure 6 It is a block diagram of a prediction module provided by an embodiment of the present invention.

[0105] Figure 7 It is a block diagram of a training device for a neural network model provided by an embodiment of the present invention.

[0106] Figure 8 It is a block diagram of an electronic device provided by an embodiment of the present invention.

[0107] Fig. 9 It is a block diagram of a device for classifying long-tail distribution data or training a neural network model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0108] In order to enable ordinary persons in the art to better understand the technical solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings.

[0109] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Instead, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0110] In recent years, research on computer vision, deep learning, machine learning, image processing, image recognition and other technologies based on artificial intelligence has made important progress. Artificial Intelligence (AI) is an emerging science and technology that studies and develops theories, methods, technologies and application systems for simulating and extending human intelligence. Artificial Intelligence is a comprehensive discipline involving many types of technologies such as chips, big data, cloud computing, the Internet of Things, distributed storage, deep learning, machine learning, neural networks, etc. Computer vision, as an important branch of artificial intelligence, specifically allows machines to recognize the world. Computer vision technology usually includes face recognition, liveness detection, fingerprint recognition and anti-counterfeiting verification, biometric recognition, face detection, pedestrian detection, target detection, pedestrian recognition, image processing, image recognition, image semantic understanding, image retrieval, text recognition, video processing, video content recognition, behavior recognition, 3D reconstruction, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), computational photography, robot navigation and positioning and other technologies. With the research and advancement of artificial intelligence technology, this technology has been applied in many fields, such as security, urban management, traffic management, building management, park management, facial access, facial attendance, logistics management, warehouse management, robots, intelligent marketing, computational photography, mobile phone imaging, cloud services, smart homes, wearable devices, unmanned driving, automatic driving, smart medical care, facial payment, facial unlocking, fingerprint unlocking, identity verification, smart screens, smart TVs, cameras, mobile Internet, live streaming, beauty, makeup, medical beauty, and smart temperature measurement.

[0111] Technical terms:

[0112] Long-tailed distribution: A common data distribution in the real world, as shown in the figure below. Its characteristic is that the proportion of samples in most categories is at a low level, while the proportion of a few categories is very high. In this solution, the minority category refers to the tail data, and the majority category refers to the head data.

[0113] Under-sampling: Sampling from the majority class data. For example, in binary classification data, the number of positive and negative class samples is 100,000 and 1,000, and 1,000 are sampled from the 100,000 positive class samples to make the results of the two classes equivalent.

[0114] Over-sampling:

[0115] Repeated sampling is performed from the minority class data. For example, in the binary classification data, the number of positive and negative class samples is 100,000 and 1,000 respectively. 100,000 samples are repeatedly extracted from the 1,000 negative class samples to make the results of the two classes equivalent.

[0116] LSTM (Long Short-Term Memory): Long short-term memory network, a neural network for modeling sequence data.

[0117] MAML (Model-Agnostic Meta-Learning): A model-agnostic meta-learning solution, a meta-learning algorithm that optimizes network parameters through a two-layer gradient back-propagation algorithm.

[0118] BCE Loss (Binary CrossEntropy): Binary cross entropy loss, a commonly used loss function for binary classification tasks.

[0119] Focal

[0120] Loss: A classification loss commonly used in object detection tasks, which can be used in classification tasks to balance different amounts of multi-class data.

[0121] maxPooling: Maximum pooling, a nonlinear transformation operation, that is, taking the current maximum value as the transformed value.

[0122] Forward propagation: input through the input layer, all the way forward, and a result is output through the output layer.

[0123] Backpropagation: The process of updating weights by going backwards through the output.

[0124] After understanding the above technical terms, please also refer to the following embodiments.

[0125] Figure 1 is a flow chart of a classification method for long-tail distribution data provided by an embodiment of the present invention. Figure 1 As shown, the method comprises the following steps:

[0126] Step 101: Obtain text of multi-label long-tail distribution data to be classified.

[0127] Step 102: Input the text into a trained neural network model for classification prediction to obtain the category of the text, wherein the neural network model includes: a first neural network capable of classifying minority class data, a second neural network capable of classifying minority class data and majority class data, a fusion network and a classification network, the first neural network and the second neural network are respectively connected to the fusion network, and the fusion network is connected to the classification network.

[0128] In the embodiment of the present invention, the text of multi-label long-tail distribution data to be classified is input into the trained neural network model for classification prediction to obtain the category of the text. In other words, the embodiment of the present invention uses dual neural networks, fusion networks and classification networks for classification prediction, solves the classification problem of multi-label long-tail distribution data, and improves the scope of application.

[0129] The classification method for long-tail distribution data described in the present invention can be applied to terminals, servers, etc., without limitation herein. The terminal implementation device can be electronic devices such as smart phones, laptops, tablet computers, desktop computers, personal digital assistants (PDAs) and wearable devices. The server can be an independent server or a server cluster, or a server that provides cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, intermediate services, domain name services, security services, content distribution networks, or big data and artificial intelligence platforms, etc., without limitation herein.

[0130] Combine the following Figure 1 , the specific implementation steps of a classification method for long-tail distribution data provided by an embodiment of the present invention are described in detail.

[0131] In step 101, text of multi-label long-tail distribution data to be classified is obtained.

[0132] In this step, the multi-label long-tail distribution data to be classified is first obtained, and then the multi-label long-tail distribution data to be classified is converted into a data format to obtain the text of the multi-label long-tail distribution data to be classified. The specific data format conversion process is already a well-known technology for those skilled in the art and will not be repeated here.

[0133] In step 102, the text is input into a trained neural network model for classification prediction to obtain the category of the text, wherein the neural network model includes: a first neural network with minority class data classification capability, a second neural network with minority class data and majority class data classification capability, a fusion network and a classification network, the first neural network and the second neural network are respectively connected to the fusion network, and the fusion network is connected to the classification network.

[0134] In this step, after the text is input into the first neural network in the trained neural network model for processing, it is respectively input into the first linear change matrix and the second linear change matrix for feature extraction to obtain the corresponding first feature matrix and the second feature matrix; the text is input into the second neural network in the trained neural network model for feature extraction to obtain the third feature matrix; wherein the first neural network and the second neural network are both of the same bidirectional long short-term memory network LSTM structure; using the attention mechanism, on the fusion network in the trained neural network model, the first feature matrix, the second feature matrix and the third feature matrix are subjected to feature fusion and then subjected to maximum pooling operation to obtain the first fusion feature; the third feature matrix output by the second neural network is subjected to maximum pooling operation, and the feature after maximum pooling is added to the first fusion feature to obtain the classification feature; the classification feature is input into the classification network in the trained neural network model for classification prediction to obtain the category of the text.

[0135] In an embodiment of the present invention, the text of multi-label long-tail distribution data to be classified is input into the first neural network and the second neural network in the trained neural network model for feature extraction. The features output by the first neural network are respectively subjected to two different linear change matrices to obtain the corresponding first feature matrix and the second feature matrix; the features output by the second neural network are the third feature matrix. After that, the attention mechanism is used to fuse the above three feature matrices to obtain a fused feature matrix, and the fused feature matrix is ​​subjected to the maximum pooling operation to obtain the fused features; the third feature matrix output by the second neural network and the fused feature matrix are respectively subjected to the maximum pooling operation in the sequence length dimension by using maxPooling, and then added to obtain the features input by the classification network, which are called classification features. Finally, the classification features are input into the classification network for classification prediction processing to obtain the category to which the final text belongs.

[0136] In an embodiment of the present invention, the text of the multi-label long-tail distribution data to be classified is input into a trained neural network model for classification prediction to obtain the category of the text, wherein the neural network model includes: a first neural network with the ability to classify minority class data, a second neural network with the ability to classify minority class data and majority class data, a fusion network and a classification network, the first neural network and the second neural network are respectively connected to the fusion network, and the fusion network is connected to the classification network. That is to say, in an embodiment of the present invention, the multi-label long-tail distribution data is classified through the dual-stage network in the trained neural network model, which not only improves the network convergence speed, but also incorporates more classification information, thereby improving the classification efficiency of the multi-label long-tail distribution data.

[0137] It should be noted that, for the method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that this implementation disclosure is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the present invention.

[0138] Optionally, in another embodiment, based on the above embodiment, the method further includes: pre-training the neural network model in the following manner:

[0139] A training set is obtained, wherein the training set includes: minority class data and labels of minority class data; and majority class data and labels of majority class data; based on the minority class data and labels of minority class data, a bidirectional long short-term memory network LSTM is trained for multi-task binary classification using a meta-learning method to obtain a first neural network A with a small number of sample classification capabilities; based on the parameters of the first neural network A, a second neural network B is constructed, wherein the first neural network A and the second neural network B both use the same bidirectional long short-term memory network LSTM structure; the majority class data are respectively input into the first neural network and the second neural network for multi-label classification training to obtain two features output by the first neural network and one feature output by the second neural network; all features output by the first neural network and the second neural network are subjected to feature fusion through an attention mechanism, and a maximum pooling operation is performed on the fused features to obtain fused features; the features output by the second neural network B are subjected to a maximum pooling operation and then added to the fused features to obtain classification features; the classification features are input into a classification network for classification prediction training to obtain training values; and a loss function Focal Loss, calculate the labels of the majority class data and the training values ​​to obtain the loss value; based on the loss value, update the parameters of the second neural network through the back propagation mechanism, after multiple iterations, until the second neural network has the long-tail distribution multi-label classification capability, and obtain a trained neural network model.

[0140] Optionally, the method of using meta-learning to perform multi-task binary classification training on the minority class data to obtain a first neural network A with few-sample classification capabilities includes: obtaining training data, including text and labels corresponding to the text; determining the minority class data text in the training data set according to a set long-tail distribution threshold; for the minority class data text, constructing a binary classification data set one by one according to the labels corresponding to the minority class data text; for the binary classification data set, using MAML meta-learning to perform multi-task binary classification training on a bidirectional long short-term memory network LSTM to obtain a first neural network A with few-sample classification capabilities.

[0141] Optionally, the binary classification data set is trained by using a MAML meta-learning method to perform multi-task binary classification training on a bidirectional long short-term memory network LSTM to obtain a first neural network A with a small sample classification capability, including:

[0142] Select the text and corresponding labels of the first batch of samples, and the text and corresponding labels of the second batch of samples from the binary classification data set; based on the selected text and corresponding labels of the first batch of samples, calculate the first loss value L1 of the bidirectional long short-term memory network LSTM on the first batch of samples; based on the first loss value L1 on the first batch of samples, use the gradient descent method to calculate the LSTM parameters after gradient update; based on the selected text and corresponding labels of the second batch of samples, calculate the second loss value L2 of the LSTM on the second batch of samples after the parameters are updated through back propagation; repeat the above calculation steps multiple times (for example, M times, M is a natural number not equal to zero), and accumulate the second loss value L2 obtained multiple times (for example, M times) to the meta-learning loss; use the meta-learning loss to perform one-time meta-backpropagation to obtain the trained parameters of the LSTM, and the trained LSTM is called the first neural network A with few-sample classification capability.

[0143] That is to say, in the embodiment of the present invention, a two-stage training method is adopted. In the first stage, a meta-learning method is used to train on minority class data to obtain a first neural network A with a small number of sample classification capabilities. The second stage is to train on the full amount of data (i.e., majority class data). During the training process, a second neural network B is constructed based on the parameters of the first neural network A. After that, the parameters of the first neural network A are fixed as a feature extractor, and feature fusion is performed with the output of the second neural network B through the attention mechanism. The first neural network A and the second neural network B both use the same bidirectional long short-term memory network LSTM structure, as follows Figure 2 As shown, Figure 2A schematic diagram of the structure of a neural network provided for an embodiment of the present invention. The neural network can be a first neural network or a second neural network. Specifically include an embedding layer, a forward LSTM, a reverse LSTM, feature fusion (i.e., concatenation & flattening), and a classification network. Among them, the embedding layer inputs the text of the multi-label long-tail distribution data to be classified, wherein the multi-label long-tail distribution data to be classified, in this embodiment, takes "Heart is not enlarged" as an example, and is not limited to this in practical applications. And the Heart is not enlarged is converted into a data format, that is, converted into multiple text features, i.e., Heart, is, not and enlarged text features, after which each converted text is input into a forward LSTM and a reverse LSTM for feature extraction, and all the features extracted by the forward LSTM and the reverse LSTM are subjected to feature fusion, and the fused fusion features are input into multiple category classifiers in the classification network (in this embodiment, category 1 classifier, category 2 classifier, ..., category n classifier) ​​for classification to obtain the category of each text.

[0144] Based on the above neural network, the following training is performed:

[0145] Phase 1: Use minority class data for meta-learning to obtain a feature extraction network suitable for minority classes.

[0146] 1) First, obtain the training data, including text and its corresponding labels.

[0147] 2) Setting a long-tail distribution threshold τ, and determining the minority class data text in the training data set according to the set long-tail distribution threshold, that is, the samples whose category proportion is lower than the long-tail distribution threshold are recorded as the minority class, and the remaining data are recorded as the majority class.

[0148] 3) For minority class data, a 1:1 classification data set is constructed according to the corresponding labels one by one, which is called a binary classification data set, including:

[0149] After randomly selecting two categories, sampling without replacement is performed to extract a total of N samples belonging to this category, and a binary classification data set is constructed with a total of N samples from these two categories of data.

[0150] 4) For the above binary classification data set, the MAML meta-learning scheme is used to perform multi-task binary classification training on LSTM:

[0151] a) Sample a small batch of samples from the binary classification dataset, i.e., D = {x, y}, where x is the text and y is the category label, and calculate the loss value L1 of the model (i.e., LSTM) on this batch of samples.

[0152] b) Use the gradient descent method to calculate the model (i.e. LSTM) parameters after gradient update Where θ is the network parameter, α is the learning rate, L1 is the loss value obtained in step a), and θ' is the updated network parameter.

[0153] c) Sample another small batch of samples from the binary classification dataset, and calculate the loss L2 of the model (LSTM) after the LSTM parameters are updated on the batch of samples through back propagation.

[0154] d) Repeat steps a), b), and c) multiple times (for example, M times, where M is a natural number not equal to zero), and accumulate the loss values ​​obtained in step c) multiple times (M times is taken as an example in this embodiment) to obtain the meta-learning loss value, and the accumulation formula is: Wherein, L2 is the loss value obtained in step c).

[0155] e) Perform a meta-backpropagation using the meta-learning loss value, Wherein θ is the network parameter, β is the learning rate, L is the loss value obtained in step d), and θ' is the updated network parameter, wherein the network parameter is the parameter of LSTM. Afterwards, the LSTM based on the trained LSTM parameters is called the first neural network A with few-sample classification capability.

[0156] The second stage: Use the full amount of data for training to obtain a network with long-tail distribution multi-label classification capabilities.

[0157] 1) Figure 2 The classification network part is used as the first neural network A, and the parameters of the first neural network A are copied to the second neural network B, and the first neural network A is fixed without training, wherein the first neural network A and the second neural network B both use the same bidirectional long short-term memory network LSTM structure. That is, based on the parameters of the first neural network A, the second neural network B is constructed.

[0158] 2) Use Figure 3 The network structure shown uses the forward propagation process to perform multi-label classification training on the majority class data to obtain label prediction values. Figure 3 A schematic diagram of multi-label classification training using forward propagation is provided in an embodiment of the present invention, wherein the forward propagation process is as follows:

[0159] 1) Input the text into neural network B (i.e., the second neural network) for feature extraction, and use the output features as the third feature matrix Q, and input the text into neural network A (i.e., the first neural network) for feature extraction, and process the extracted features using two linear change matrices W1 and W2 respectively to obtain the first feature matrix K and the second feature matrix V. In this embodiment, the input text is represented by x, and neural network A and neural network B are represented by f respectively. A , f B It means that the first characteristic matrix K and the second characteristic matrix V are obtained by the following calculation formulas:

[0160] Q=f B (x),K=W1f A (x),V=W2f A (x)

[0161] That is, the first characteristic matrix K and the first characteristic matrix V are obtained by matrix multiplication.

[0162] 2) Use the attention mechanism to fuse features and obtain features: Among them, softmax is a normalization function used to calculate the probability value of each dimension, specifically: Among them, Q, K, V are the feature matrices obtained in step 1), and d is the dimension of latent variables.

[0163] 3) performing a maximum pooling operation on the fused feature matrix to obtain a first fused feature; performing a maximum pooling operation on the third feature matrix Q output by the neural network B to obtain a pooled feature; and adding the maximum pooled features to obtain a feature f input to the classification network, wherein the feature f is referred to as a classification feature in this article;

[0164] Among them, the feature output by the neural network B (size is batch_size*sequence_length*hidden_size) is obtained by maxpooling in the sequence length dimension to obtain f2 (size is batch_size*hidden_size), which is added to the feature f1 (size is batch_size*hidden_size obtained by maxpooling) as the final classification feature f.

[0165] Among them, batch_size: the number of samples contained in the batch; sequence_length: sequence length; Hidden_size: feature (hidden state) dimension; maxpooling is maximum pooling.

[0166] Among them, Max Pooling extracts several eigenvalues ​​from a certain Filter, and only takes the largest Pooling layer as the retained value, and discards all other eigenvalues. The largest value means that only the strongest of these features is retained, and other weak features of this type are discarded.

[0167] 4) Inputting the classification feature f into the classification network in the trained neural network model for classification prediction to obtain the category of the text.

[0168] That is to say, the feature f is input into the classification network, that is, the classification network is used for classification, and the classification result is passed through soft max to obtain the final label prediction value.

[0169] 5) Using the loss function Focal Loss, the labels of the majority class data and the training values ​​are calculated to obtain a loss value; based on the loss value, the parameters of the second neural network are updated through a back propagation mechanism, and after multiple iterations, until the second neural network has the ability to classify long-tail distribution multiple labels, a trained neural network model is obtained.

[0170] That is to say, this embodiment uses Focal Loss as the loss function and performs back propagation to further balance the sample weights of different data amounts. The loss function is:

[0171] L = -α(1-p) γ log(p)

[0172] Among them, α∈[0,1],γ∈[0,5] are hyperparameters used to balance positive and negative samples of different sizes, and p is the network output. At this time, the parameters of the neural network A are continuously updated during the back propagation process. After multiple iterations, until the neural network A has the ability of long-tail distribution multi-label classification, a trained neural network model is obtained.

[0173] In the embodiment of the present invention, a two-stage network training process is adopted. The first stage of training can enable the network to have better classification ability on long-tail data (minority classes). At the same time, the meta-learning process can give the network parameters the ability to generalize quickly and improve the network convergence speed. The second stage of training can improve the network's classification ability on the target data, while ensuring that the network does not forget the classification ability of the minority classes. The two-stage network combination process performs feature fusion between the majority class and minority class features through the attention mechanism, so that the classification features incorporate more information and are more robust.

[0174] See also Figure 4 , is a flow chart of a training method for a neural network model provided by an embodiment of the present invention, the method comprising:

[0175] Step 401: Obtain a training set, wherein the training set includes: minority class data and labels of the minority class data; and majority class data and labels of the majority class data;

[0176] Step 402: Based on the minority class data and the labels of the minority class data, a bidirectional long short-term memory network LSTM is trained for multi-task binary classification using a meta-learning method to obtain a first neural network A with a small sample classification capability;

[0177] Specifically, training data is obtained, including text and labels corresponding to the text; according to the set long-tail distribution threshold, the minority data text in the training data set is determined; for the minority data text, a binary classification data set is constructed one by one according to the labels corresponding to the minority data text; for the binary classification data set, the MAML meta-learning method is used to perform multi-task binary classification training on the bidirectional long short-term memory network LSTM, so as to obtain a first neural network A with few-sample classification capability.

[0178] The binary classification data set is trained by using MAML meta-learning to perform multi-task binary classification training on the bidirectional long short-term memory network LSTM to obtain a first neural network A with a small sample classification capability, including:

[0179] Select the text and corresponding labels of the first batch of samples, and the text and corresponding labels of the second batch of samples from the binary classification data set; based on the selected text and corresponding labels of the first batch of samples, calculate the first loss value L1 of the bidirectional long short-term memory network LSTM on the first batch of samples; based on the first loss value L1 on the first batch of samples, use the gradient descent method to calculate the LSTM parameters after the gradient update; based on the selected text and corresponding labels of the second batch of samples, calculate the second loss value L2 of the LSTM on the second batch of samples after the LSTM parameters are updated through back propagation; repeat the above calculation steps multiple times (for example, M times, M is a natural number not equal to zero), and accumulate the second loss values ​​L2 obtained multiple times to obtain a meta-learning loss value; use the meta-learning loss value to perform a meta-back propagation to obtain the trained parameters of the LSTM, and the LSTM based on the trained LSTM parameters is called the first neural network A with few-sample classification capability.

[0180] Step 403: constructing a second neural network B based on the parameters of the first neural network A, wherein the first neural network A and the second neural network B both adopt the same bidirectional long short-term memory network LSTM structure;

[0181] Step 404: inputting the majority class data into the first neural network and the second neural network respectively for multi-label classification training to obtain two features output by the first neural network and one feature output by the second neural network;

[0182] Step 405: All features output by the first neural network and the second neural network are fused through an attention mechanism, and a maximum pooling operation is performed on the fused features to obtain fused features;

[0183] Step 406: Perform a maximum pooling operation on the features output by the second neural network B, and add the features to the fusion features to obtain classification features;

[0184] Step 407: inputting the classification features into the classification network for classification prediction training to obtain training values;

[0185] Step 408: using a loss function (such as Focal Loss) to calculate the labels of the majority class data and the training values ​​to obtain a loss value;

[0186] Step 409: Based on the loss value, the parameters of the second neural network are updated through a back-propagation mechanism, and after multiple iterations, the second neural network has the long-tail distribution multi-label classification capability, thereby obtaining a trained neural network model.

[0187] In the embodiment of the present invention, for the common long-tail distribution problem in real data, relying on the text multi-label task, the LSTM network is trained by a two-stage training method combining meta-learning and focal loss, so that the classification ability of the minority class is not lost while the classification ability on the majority class data is possessed. In the minority class training stage, this embodiment uses meta-learning to perform binary classification training of multiple tasks, so that the network has the classification ability of the minority class. In the full training stage, focal loss is used to train on all data, so that the network has the classification ability on the majority class while retaining the classification ability of the minority class. Using the new network architecture, feature fusion can be flexibly performed to deal with the problem of excessive differences in long-tail distribution data, and at the same time, it has multi-label classification capabilities. It has good generalization, does not rely on specific network implementation, and can adapt to neural networks with various structures such as CNN, LSTM, and Bert. That is to say, compared with the traditional method applicable to multi-category classification tasks, this embodiment solves the classification problem of multi-label long-tail distribution data by following the paradigm of integrated learning and the comprehensive application of components such as meta-learning and focal loss, and has a wider range of application.

[0188] Figure 5 is a block diagram of a classification device for long-tail distribution data provided by an embodiment of the present invention. The device comprises: an acquisition module 51 and a prediction module 52, wherein:

[0189] An acquisition module 51 is used to acquire text of multi-label long-tail distribution data to be classified;

[0190] The prediction module 52 is used to input the text into a trained neural network model for classification prediction to obtain the category of the text, wherein the neural network model includes: a first neural network with the ability to classify minority class data, a second neural network with the ability to classify minority class data and majority class data, a fusion network and a classification network, the first neural network and the second neural network are respectively connected to the fusion network, and the fusion network is connected to the classification network.

[0191] Optionally, in another embodiment, based on the above embodiment, the prediction module 502 includes: a first feature extraction module 601, a second feature extraction module 602, a feature fusion module 603, a classification feature determination module 604 and a classification prediction module 605, and its structural block diagram is shown as follows: Figure 6 As shown,

[0192] A first feature extraction module 601 is used to input the text into the first neural network A in the trained neural network model for processing, and then input it into the first linear change matrix W1 and the second linear change matrix W2 for feature extraction to obtain the corresponding first feature matrix K and the second feature matrix V;

[0193] A second feature extraction module 602 is used to input the text into a second neural network B in a trained neural network model for feature extraction to obtain a third feature matrix Q; wherein the first neural network and the second neural network both have the same bidirectional long short-term memory network LSTM structure;

[0194] A feature fusion module 603 is used to use an attention mechanism to perform feature fusion on the first feature matrix K, the second feature matrix V and the third feature matrix Q on the fusion network in the trained neural network model, and then perform a maximum pooling operation to obtain a first fused feature f1;

[0195] A classification feature determination module 604 is used to perform a maximum pooling operation on the third feature matrix Q output by the second neural network, and add the maximum pooled feature f2 to the first fusion feature f1 to obtain a classification feature f;

[0196] The classification prediction module 605 is used to input the classification feature f into the classification network in the trained neural network model to perform classification prediction and obtain the category of the text.

[0197] Optionally, in another embodiment, based on the above embodiment, the device further includes: a training module for pre-training the neural network model.

[0198] Optionally, in another embodiment, based on the above embodiment, the training module includes:

[0199] A training set acquisition module is used to acquire a training set, wherein the training set includes: minority class data and labels of the minority class data; and majority class data and labels of the majority class data;

[0200] A first training module is used to perform multi-task binary classification training on a bidirectional long short-term memory network LSTM by using a meta-learning method based on the minority class data and the labels of the minority class data, so as to obtain a first neural network A with a small number of sample classification capabilities;

[0201] A construction module, used to construct a second neural network B based on the parameters of the first neural network A, wherein the first neural network A and the second neural network B both adopt the same bidirectional long short-term memory network LSTM structure;

[0202] A second training module is used to input the majority class data into the first neural network and the second neural network respectively for multi-label classification training to obtain two features output by the first neural network and one feature output by the second neural network;

[0203] A fusion module, used to fuse all features output by the first neural network and the second neural network through an attention mechanism, and perform a maximum pooling operation on the fused features to obtain fused features;

[0204] A feature determination module, used for performing a maximum pooling operation on the features output by the second neural network B and adding the features to the fusion features to obtain classification features;

[0205] A classification prediction module, used for inputting the classification features into a classification network for classification prediction training to obtain training values;

[0206] A loss determination module, used to calculate the labels of the majority class data and the training values ​​using a loss function Focal Loss to obtain a loss value;

[0207] The iterative module is used to update the parameters of the second neural network through a back-propagation mechanism based on the loss value, and after multiple iterations, the second neural network has the long-tail distribution multi-label classification capability to obtain a trained neural network model.

[0208] Optionally, in another embodiment, based on the above embodiment, the first training module includes:

[0209] A data acquisition module, used to acquire training data, including text and labels corresponding to the text;

[0210] A text determination module, used to determine the minority class data text in the training data set according to a set long-tail distribution threshold;

[0211] A binary classification data determination module, for constructing a binary classification data set for the minority class data texts one by one according to the labels corresponding to the minority class data texts;

[0212] The binary classification training module is used to perform multi-task binary classification training on the bidirectional long short-term memory network LSTM using the MAML meta-learning method on the binary classification data set to obtain a first neural network A with a small number of sample classification capabilities.

[0213] Optionally, in another embodiment, based on the above embodiment, the binary classification training module includes:

[0214] A selection module, used to select text and corresponding labels of the first batch of samples, and text and corresponding labels of the second batch of samples from the binary classification data set;

[0215] A first loss calculation module, used to calculate a first loss value L1 of a bidirectional long short-term memory network LSTM on the first batch of samples based on the texts and corresponding labels of the selected first batch of samples;

[0216] A parameter calculation module, used to calculate the LSTM parameters after gradient update by using a gradient descent method based on the first loss value L1 on the first batch of samples;

[0217] A second loss calculation module, used for calculating, based on the texts and corresponding labels of the selected second batch of samples, a second loss value L2 of the LSTM after parameter update on the second batch of samples through back propagation;

[0218] An accumulation module, used to repeat the above calculation steps multiple times, and accumulate the second loss value L2 obtained multiple times to the meta-learning loss;

[0219] A meta-back propagation module is used to perform a meta-back propagation using the meta-learning loss to obtain the parameters of the trained LSTM, and the trained LSTM is called the first neural network A with few-sample classification capability.

[0220] See also Figure 7A training device block diagram of a neural network model provided by an embodiment of the present invention includes: a training set acquisition module 701, a first training module 702, a construction module 703, a second training module 704, a fusion module 705, a feature determination module 706, a classification prediction module 707, a loss determination module 708 and an iteration module 709, wherein:

[0221] The training set acquisition module 701 acquires a training set, wherein the training set includes: minority class data and labels of the minority class data; and majority class data and labels of the majority class data;

[0222] A first training module 702, based on the minority class data and the labels of the minority class data, performs multi-task binary classification training on the bidirectional long short-term memory network LSTM by using a meta-learning method to obtain a first neural network A with a small number of sample classification capabilities;

[0223] A construction module 703 constructs a second neural network B based on the parameters of the first neural network A, wherein the first neural network A and the second neural network B both adopt the same bidirectional long short-term memory network LSTM structure;

[0224] A second training module 704, inputting the majority class data into the first neural network and the second neural network respectively for multi-label classification training, to obtain two features output by the first neural network and one feature output by the second neural network;

[0225] A fusion module 705 is used to fuse all features output by the first neural network and the second neural network through an attention mechanism, and perform a maximum pooling operation on the fused features to obtain fused features;

[0226] The feature determination module 706 performs a maximum pooling operation on the feature output by the second neural network B and adds the result to the fusion feature to obtain a classification feature;

[0227] The classification prediction module 707 inputs the classification features into the classification network for classification prediction training to obtain training values;

[0228] The loss determination module 708 calculates the labels of the majority class data and the training values ​​using a loss function Focal Loss to obtain a loss value;

[0229] Iteration module 709 updates the parameters of the second neural network through a back propagation mechanism based on the loss value, and after multiple iterations, the second neural network has the long-tail distribution multi-label classification capability to obtain a trained neural network model.

[0230] Optionally, in another embodiment, based on the above embodiment, the first training module includes:

[0231] A data acquisition module, used to acquire training data, including text and labels corresponding to the text;

[0232] A text determination module, used to determine the minority class data text in the training data set according to a set long-tail distribution threshold;

[0233] A binary classification data determination module, for constructing a binary classification data set for the minority class data texts one by one according to the labels corresponding to the minority class data texts;

[0234] The binary classification training module is used to perform multi-task binary classification training on the bidirectional long short-term memory network LSTM using the MAML meta-learning method on the binary classification data set to obtain a first neural network A with a small number of sample classification capabilities.

[0235] Optionally, in another embodiment, based on the above embodiment, the binary classification training module includes:

[0236] A selection module, used to select text and corresponding labels of the first batch of samples, and text and corresponding labels of the second batch of samples from the binary classification data set;

[0237] A first loss calculation module, used to calculate a first loss value L1 of a bidirectional long short-term memory network LSTM on the first batch of samples based on the texts and corresponding labels of the selected first batch of samples;

[0238] A parameter calculation module, used to calculate the LSTM parameters after gradient update by using a gradient descent method based on the first loss value L1 on the first batch of samples;

[0239] A second loss calculation module, used for calculating, based on the texts and corresponding labels of the selected second batch of samples, a second loss value L2 of the LSTM after parameter update on the second batch of samples through back propagation;

[0240] An accumulation module, used to repeat the above calculation steps multiple times (for example, M times), and accumulate the second loss value L2 obtained multiple times to the meta-learning loss;

[0241] A meta-back propagation module is used to perform a meta-back propagation using the meta-learning loss to obtain the parameters of the trained LSTM, and the trained LSTM is called the first neural network A with few-sample classification capability.

[0242] Optionally, an embodiment of the present invention further provides an electronic device, including:

[0243] processor;

[0244] a memory for storing instructions executable by the processor;

[0245] Wherein, the processor is configured to execute the instructions to implement the classification method of long-tail distribution data as described above or the training method of the neural network model as described above.

[0246] Optionally, an embodiment of the present invention also provides a computer-readable storage medium, which, when the instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to execute the classification method of long-tail distribution data as described above or the training method of the neural network model as described above.

[0247] Optionally, an embodiment of the present invention also provides a computer program product, including a computer program or instructions, which, when executed by a processor of an electronic device, implements the classification method for long-tail distribution data as described above or the training method for a neural network model as described above.

[0248] Regarding the device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the method, and will not be elaborated here.

[0249] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0250] Figure 8 800 is a block diagram of an electronic device 800 provided in an embodiment of the present invention. For example, the electronic device 800 may be a mobile terminal or a server, and the electronic device is a mobile terminal as an example for explanation in the embodiment of the present invention. For example, the electronic device 800 may be a mobile phone, a computer, a digital broadcast terminal, a message transceiver device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.

[0251] Reference Figure 8 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output (I / O) interface 812 , a sensor component 814 , and a communication component 816 .

[0252] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0253] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0254] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.

[0255] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0256] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0257] I / O interface 812 provides an interface between processing component 802 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: home button, volume button, start button, and lock button.

[0258] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the electronic device 800, and the sensor assembly 814 can also detect the position change of the electronic device 800 or a component of the electronic device 800, the presence or absence of contact between the user and the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and the temperature change of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0259] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G or 5G), or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0260] In an embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute the classification method for the long-tail distribution data shown above or the training method for the neural network model as described above.

[0261] In an embodiment, a computer-readable storage medium is also provided, and when the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device 800 can perform the classification method of the long-tail distribution data shown above or the training method of the neural network model as described above. For example, the computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0262] In an embodiment, a computer program product is also provided, including a computer program or instructions. When the computer program or instructions are executed by the processor 820 of the electronic device 800, the electronic device 800 executes the above-mentioned classification method of long-tail distribution data or the above-mentioned training method of the neural network model.

[0263] Fig. 9 is a block diagram of an apparatus 900 for classifying long-tail distribution data or training a neural network model provided by an embodiment of the present invention. For example, the apparatus 900 may be provided as a server. Fig. 9 , the apparatus 900 includes a processing component 922, which further includes one or more processors, and a memory resource represented by a memory 932 for storing instructions, such as an application, that can be executed by the processing component 922. The application stored in the memory 932 may include one or more modules, each corresponding to a set of instructions. In addition, the processing component 922 is configured to execute the instructions to perform the above method.

[0264] The device 900 may also include a power supply component 926 configured to perform power management of the device 900, a wired or wireless network interface 950 configured to connect the device 900 to a network, and an input / output (I / O) interface 958. The device 900 may operate based on an operating system stored in the memory 932, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, or the like.

[0265] Those skilled in the art will readily appreciate other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses or adaptations of the present invention that follow the general principles of the present invention and include common knowledge or customary techniques in the art that are not disclosed by the present invention. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present invention are indicated by the following claims.

[0266] It should be understood that the present invention is not limited to the exact construction that has been described above and shown in the drawings and that various modifications and changes may be made without departing from the scope thereof. The scope of the present invention is limited only by the appended claims.

Claims

1. A classification method for long-tail distribution data, characterized in that: include: Get the text of multi-label long-tail distribution data to be classified; The text is input into a trained neural network model for classification prediction to obtain the category of the text, wherein the neural network model includes: a first neural network with the ability to classify minority class data, a second neural network with the ability to classify minority class data and majority class data, a fusion network and a classification network, the first neural network and the second neural network are respectively connected to the fusion network, and the fusion network is connected to the classification network.

2. The classification method for long-tail distribution data according to claim 1, characterized in that: The step of inputting the text into a trained neural network model for classification prediction to obtain the category of the text includes: After the text is input into the first neural network in the trained neural network model for processing, it is respectively input into the first linear change matrix and the second linear change matrix for feature extraction to obtain the corresponding first feature matrix and the second feature matrix; Input the text into a second neural network in a trained neural network model for feature extraction to obtain a third feature matrix; wherein the first neural network and the second neural network both have the same bidirectional long short-term memory network LSTM structure; Using the attention mechanism on the fusion network in the trained neural network model, perform a maximum pooling operation on the first feature matrix, the second feature matrix, and the third feature matrix after feature fusion to obtain a first fusion feature; Performing a maximum pooling operation on the third feature matrix output by the second neural network, and adding the maximum pooled features to the first fusion features to obtain classification features; The classification features are input into the classification network in the trained neural network model for classification prediction to obtain the category of the text.

3. The classification method for long-tail distribution data according to claim 1 or 2, characterized in that: The method further comprises: pre-training the neural network model in the following manner: Obtain a training set, wherein the training set includes: minority class data and labels of the minority class data; and majority class data and labels of the majority class data; Based on the minority class data and the labels of the minority class data, a bidirectional long short-term memory network LSTM is trained for multi-task binary classification using a meta-learning method to obtain a first neural network with a small number of sample classification capabilities; Based on the parameters of the first neural network, a second neural network is constructed, wherein the first neural network and the second neural network both adopt the same bidirectional long short-term memory network LSTM structure; Inputting the majority class data into the first neural network and the second neural network respectively for multi-label classification training to obtain two features output by the first neural network and one feature output by the second neural network; All features output by the first neural network and the second neural network are fused through an attention mechanism, and a maximum pooling operation is performed on the fused features to obtain fused features; After performing a maximum pooling operation on the features output by the second neural network, the features are added to the fusion features to obtain classification features; Inputting the classification features into a classification network for classification prediction training to obtain training values; Using a loss function, the labels of the majority class data and the training values ​​are calculated to obtain a loss value; The parameters of the second neural network are updated through a back-propagation mechanism based on the loss value, and after multiple iterations, the second neural network has the ability to classify long-tail distribution multi-labels, thereby obtaining a trained neural network model.

4. The classification method for long-tail distribution data according to claim 3, characterized in that: The method of performing multi-task binary classification training on the minority class data by using a meta-learning method to obtain a first neural network with a small sample classification capability includes: Obtain training data, including text and labels corresponding to the text; Determining the minority class data text in the training data set according to the set long-tail distribution threshold; For the minority class data texts, construct a binary classification data set according to the labels corresponding to the minority class data texts one by one; For the binary classification data set, the MAML meta-learning method is used to perform multi-task binary classification training on the bidirectional long short-term memory network LSTM to obtain a first neural network with few-sample classification capability.

5. The classification method for long-tail distribution data according to claim 4, characterized in that: The binary classification data set is trained by using the MAML meta-learning method to perform multi-task binary classification training on the bidirectional long short-term memory network LSTM to obtain a first neural network with a small sample classification capability, including: Selecting text and corresponding labels of the first batch of samples and text and corresponding labels of the second batch of samples from the binary classification data set; Based on the texts and corresponding labels of the selected first batch of samples, calculate a first loss value of a bidirectional long short-term memory network LSTM on the first batch of samples; Based on the first loss value on the first batch of samples, the gradient descent method is used to calculate the LSTM parameters after gradient update; Based on the texts and corresponding labels of the selected second batch of samples, a second loss value of the LSTM after the LSTM parameters are updated on the second batch of samples is calculated through back propagation; Repeat the above calculation steps multiple times, and accumulate the second loss values ​​obtained multiple times to obtain the meta-learning loss value; The meta-learning loss value is used to perform a meta-back propagation to obtain the trained parameters of the LSTM, and the LSTM based on the trained LSTM parameters is called the first neural network with few-sample classification capability.

6. A training method for a neural network model, characterized in that: include: Obtain a training set, wherein the training set includes: minority class data and labels of the minority class data; and majority class data and labels of the majority class data; Based on the minority class data and the labels of the minority class data, a bidirectional long short-term memory network LSTM is trained for multi-task binary classification using a meta-learning method to obtain a first neural network with a small number of sample classification capabilities; Based on the parameters of the first neural network, a second neural network is constructed, wherein the first neural network and the second neural network both adopt the same bidirectional long short-term memory network LSTM structure; Inputting the majority class data into the first neural network and the second neural network respectively for multi-label classification training to obtain two features output by the first neural network and one feature output by the second neural network; All features output by the first neural network and the second neural network are fused through an attention mechanism, and a maximum pooling operation is performed on the fused features to obtain fused features; After performing a maximum pooling operation on the features output by the second neural network, the features are added to the fusion features to obtain classification features; Inputting the classification features into a classification network for classification prediction training to obtain training values; Using a loss function, the labels of the majority class data and the training values ​​are calculated to obtain a loss value; The parameters of the second neural network are updated through a back-propagation mechanism based on the loss value, and after multiple iterations, the second neural network has the ability to classify long-tail distribution multi-labels, thereby obtaining a trained neural network model.

7. A classification device for long-tail distribution data, characterized in that: include: An acquisition module is used to obtain the text of multi-label long-tail distribution data to be classified; The prediction module is used to input the text into a trained neural network model for classification prediction to obtain the category of the text, wherein the neural network model includes: a first neural network with minority class data classification capability, a second neural network with minority class data and majority class data classification capability, a fusion network and a classification network, the first neural network and the second neural network are respectively connected to the fusion network, and the fusion network is connected to the classification network.

8. A training device for a neural network model, characterized in that: include: A training set acquisition module is used to acquire a training set, wherein the training set includes: minority class data and labels of the minority class data; and majority class data and labels of the majority class data; A first training module, based on the minority class data and the labels of the minority class data, performs multi-task binary classification training on a bidirectional long short-term memory network LSTM by using a meta-learning method to obtain a first neural network with a small number of sample classification capabilities; A construction module is used to construct a second neural network based on the parameters of the first neural network, wherein the first neural network and the second neural network both adopt the same bidirectional long short-term memory network LSTM structure; A second training module, inputting the majority class data into the first neural network and the second neural network respectively for multi-label classification training, to obtain two features output by the first neural network and one feature output by the second neural network; A fusion module, which fuses all features output by the first neural network and the second neural network through an attention mechanism, and performs a maximum pooling operation on the fused features to obtain fused features; A feature determination module performs a maximum pooling operation on the features output by the second neural network and adds the features to the fusion features to obtain classification features; A classification prediction module inputs the classification features into a classification network for classification prediction training to obtain a training value; A loss determination module, using a loss function to calculate the labels of the majority class data and the training values ​​to obtain a loss value; The iterative module updates the parameters of the second neural network through a back-propagation mechanism based on the loss value, and after multiple iterations, the second neural network has the long-tail distribution multi-label classification capability, thereby obtaining a trained neural network model.

9. An electronic device, characterized in that: include: processor; a memory for storing instructions executable by the processor; Wherein, the processor is configured to execute the instructions to implement the classification method of long-tail distribution data as described in any one of claims 1 to 5 or the training method of the neural network model as described in claim 6.

10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to execute the classification method for long-tail distribution data as described in any one of claims 1 to 5 or the training method for a neural network model as described in claim 6.