Multi-granularity financial text noise open classification method based on large model
Through the multi-grained financial text noise open classification method based on large models, semantic features, clustering processing and particle-spheric attribute analysis are extracted, and the problems of insufficient classification and recognition capabilities and poor adaptability of financial text noise in the existing technology are solved, achieving more accurate noise recognition and classification effects.
Patent Information
- Application Number
- CN202510608156.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-13
AI Technical Summary
The prior art has problems such as insufficient recognition capability, poor adaptability and defects in noise processing methods in financial text noise classification, especially when processing data with complex noise and non-spherical distributions.
A multi-grained financial text noise open classification method based on large models is adopted to extract semantic features of financial text and cluster processing to obtain multiple spheres of different labels, and classify them according to the position of the sample in the sphere and the properties of the sphere, including the identification of clean samples, internal distribution noise, and external distribution noise.
It improves the accuracy of noise recognition and classification accuracy, enhances the ability to identify unknown categories, ensures efficient and stable performance in complex financial environments, and is also suitable for data environments of various distribution types.
Smart Images

Figure CN120123518A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text noise classification, and in particular to an open classification method for multi-granularity financial text noise based on a large model. Background Art
[0002] With the rapid development of fintech, financial dialogue systems have become a key technology for banks and financial service institutions to provide customer service and intelligent consultation. These dialogue systems need to be able to accurately identify and classify the query intentions of users, including regular financial transaction requests and non-predefined new intentions. Traditional intention classification systems usually can only identify the intention categories defined during the training process and have insufficient processing ability for newly emerged or unknown intention categories, which limits the adaptability and utility of noise classification methods.
[0003] In addition, existing open intention classification methods usually rely on large-scale and high-quality labeled datasets, which are often difficult to obtain in practical applications. In practical applications, the datasets often contain noise, and label noise can usually be divided into two categories: in-distribution (IND) noise and out-of-distribution (OOD) noise. IND noise occurs when an intention belonging to a predefined known category is mislabeled as another known category, while OOD noise occurs when an unknown intention is mislabeled as a predefined known category. In most real-world scenarios, IND noise and OOD noise exist simultaneously, which can be referred to as complex noise.
[0004] There are few existing noise processing methods for open intention classification, and there are a series of defects in existing noise processing research, which limit its direct application in the representation learning of open intention classification. The specific defects include: (1) Limitations of prototype representation: Current distance-based noise processing methods usually use category prototypes to represent category distributions and judge whether a sample is noise based on the distance between the sample and the prototype. On the one hand, these prototypes can only represent the central position of the category and cannot comprehensively reflect the overall distribution of the category, resulting in inaccurate distance-based noise detection. On the other hand, traditional clustering methods such as K-means clustering preset the number of clusters and cannot truly reflect the actual distribution of subclasses within each intention, thus affecting the effect of representation learning.
[0005] (2) Insufficiency of noise sample recognition methods: Current sample selection methods mainly rely on fixed distance thresholds or probability criteria to identify clean samples and IND noise. This coarse-grained global selection criterion fails to consider the distribution differences within the category and cannot be effectively applied to data environments of various distribution types. Therefore, the adaptability and scalability of these methods in variable practical application scenarios are limited. Summary of the Invention
[0006] The object of the present invention is to overcome the problems of the prior art, and provide a multi-granularity financial text noise open classification method based on a large model.
[0007] The object of the present invention is achieved by the following technical solutions: A multi-granularity financial text noise open classification method based on a large model, the method includes a training step: Extract the semantic features of financial texts; Perform clustering processing on the semantic features to obtain multiple granular balls with different labels, and calculate the attributes of the multiple granular balls with different labels, including capacity, purity, centroid and radius; According to the position of the sample in the granular ball and the attributes of the granular ball, classify the samples in the granular ball as clean samples, in-distribution noise, and out-of-distribution noise, including the following sub-steps: Define the granular balls with granular ball purity greater than the high purity threshold and granular ball capacity greater than the high capacity threshold as high-quality granular balls. The noise classification of the samples in the high-quality granular balls includes: if the distance from the sample to the centroid is less than the radius of the granular ball and the sample and the granular ball have the same label, classify the sample as a clean sample; if the distance from the sample to the centroid is less than the radius of the granular ball and the sample and the granular ball label are different, classify the sample as in-distribution noise; Define the granular balls with granular ball purity less than the low purity threshold and granular ball capacity less than the low capacity threshold as low-quality granular balls, and classify the samples in the low-quality granular balls as out-of-distribution noise.
[0008] In one example, the extraction of the semantic features of financial texts includes: Freeze the underlying parameters of the large model, and use the training data set of financial texts to train the top-level parameters of the large model. The top level of the large model includes an embedding layer and an encoding layer; Use the trained large model to extract the semantic features of financial texts.
[0009] In one example, the clustering processing of the semantic features to obtain multiple granular balls with different labels includes the following sub-steps: Initialize all samples corresponding to the semantic features as a granular ball; Calculate the attributes of the granular ball, and judge whether to perform granular ball splitting processing according to the granular ball attributes. If so, randomly select samples with each different label in the granular ball as the initial centroid of the new granular ball, calculate the distance from all samples to the initial centroid, and assign the samples to the new granular ball represented by the nearest centroid, so as to obtain multiple granular balls with different labels.
[0010] In one example, the judgment of whether to perform granular ball splitting processing according to the granular ball attributes includes: If the granular ball purity is less than the first purity threshold and if the granular ball capacity is greater than the first capacity threshold, then perform granular ball splitting.
[0011] In one example, the method further includes classifying the samples within the granules as uncertain samples: When classifying the samples in high-quality granules for noise, if the distance from the sample to the centroid is greater than the radius of the granule, classify the sample as an uncertain sample; and / or, Classify the samples in the granules with granule purity greater than the low purity threshold and less than the high purity threshold as uncertain samples; Discard the uncertain samples and do not include them in the calculation of the loss function.
[0012] In one example, after classifying the samples within the granules as clean samples and in-distribution noise, it further includes: Calculate the first cross-entropy loss function for the clean samples; After modifying the in-distribution noise labels to the labels of the granules where they are located, calculate the second cross-entropy loss function; Construct a first total loss function using the first cross-entropy loss function and the second cross-entropy loss function, and then perform optimization processing on the classification model.
[0013] In one example, after classifying the samples within the granules as out-of-distribution noise, it further includes: Calculate the intra-class loss function and the inter-class loss function based on the centroid of the high-quality granules and the label of the high-quality granules; calculate the open space loss function based on the centroid of the high-quality granules and the centroid of the low-quality granules; Construct a second total loss function using the first cross-entropy loss function, the second cross-entropy loss function, the intra-class loss function, the inter-class loss function, and the open space loss function, and then perform optimization processing on the classification model.
[0014] It should be further noted that the technical features corresponding to the above examples can be combined or replaced with each other to form a new technical solution.
[0015] In one example, the method further includes an application step: Extract the semantic features of the financial text to be classified; For each sample to be classified corresponding to the semantic features, determine the centroid of the first known category with the closest distance; the known categories use the centroids and radii of the various known categories represented by the high-quality granules as the final category boundaries; If the distance from the sample to be classified to the centroid of the first known category is less than or equal to the radius corresponding to the centroid of the first known category, classify the sample to be classified as the known category corresponding to the centroid of the first known category; if the distance from the sample to be classified to the centroid of the first known category is greater than the radius corresponding to the centroid of the first known category, classify the sample to be classified as an unknown category.
[0016] Compared with the prior art, the beneficial effects of the present invention are: 1. In one example, multiple granulocytes with different labels are obtained through clustering processing, and the attributes of each granulocyte are calculated. At this time, each category is finally represented by multiple granulocytes with the same label but different attributes. The centroid of multiple granulocytes with the same label represents multiple centers of the category, and the radius reflects the distribution range of the category. That is, through the multi-granularity characterization method, the multi-granularity prototype and distribution range of the category are effectively characterized, which can truly reflect the actual distribution of sub-categories within each category, consider the distribution differences within the category, and thus improve the representation learning effect, can effectively distinguish the feature representations of different categories and sub-categories, and provide a more accurate reference for subsequent noise recognition and classification tasks. At the same time, each category is finally represented by multiple granulocytes with the same label but different attributes, so that each category is not limited to a single boundary, but multiple boundaries are set according to the actual distribution of the data, ensuring high adaptability to various data distributions. This flexible boundary setting significantly improves the accuracy of classification. Especially when dealing with non-spherical or heterogeneous distributed data, it can effectively distinguish known categories and unknown categories, and enhance the robustness of the classification method.
[0017] Furthermore, according to the position of the sample in the granulocyte and the attributes of the granulocyte, the sample is classified into clean samples, in-distribution noise, and out-of-distribution noise. Through fine-grained feature analysis, the accuracy of noise recognition is improved, so as to maintain high efficiency and stable performance in a complex financial environment; at the same time, by identifying and classifying out-of-distribution noise, the out-of-distribution noise can be used as a placeholder for unknown categories for representation learning. In the overall representation space, the representation space of known categories is compressed, and the representation space of unknown categories is expanded, avoiding the representation space being completely occupied by known categories, leaving room for the development of potential unknown categories, effectively expanding the recognition ability of unknown categories, ensuring that the representation of known categories is more compact, so that it can be quickly and accurately classified when unknown intentions appear; in addition, classifying according to the position of the sample in the granulocyte and the attributes of the granulocyte can be applied to data environments of various distribution types.
[0018] 2. In one example, freezing the underlying parameters of the pre-trained large model can utilize the general language features already mastered by the pre-trained large model; training the top-layer parameters of the large model enables the large model to capture the nuances in financial texts, enhances the large model's understanding ability of the unique expressions in financial texts, and improves the adaptability of the large model to text speech in the financial field.
[0019] 3. In one example, splitting the granulocyte according to the granulocyte purity threshold can further subdivide the granulocyte including samples of multiple categories, thereby reducing the possibility of misclassification; splitting the granulocyte according to the capacity threshold of the granulocyte can split the granulocyte with too large a capacity, which can avoid excessive aggregation of samples in the category and further improve the classification accuracy.
[0020] 4. In one example, the samples located outside the radius of the grain sphere are classified as uncertain noise, and the loss function is not calculated for them during subsequent model training to avoid misjudging their noise categories and affecting subsequent representation learning.
[0021] 5. In one example, the classification model is optimized through the first cross-entropy loss function, which can effectively measure the difference between the predicted value and the true value, thereby guiding the optimization of the classification model to consolidate the learning performance of the classification model for correct standard samples; the second cross-entropy loss function is calculated after modifying the in-distribution noise label to the label of the grain sphere where it is located to optimize the large model, which can correct the deviation caused by incorrect labels; the large model is optimized through the intra-class loss function, which can ensure that the samples within the same category are more compactly distributed, making the centroids of the grain spheres with the same label closer to each other, thereby improving the consistency of the features of the intra-class samples; the large model is optimized through the inter-class loss function, which can increase the discrimination between different categories, making the centroids of the grain spheres with different labels farther from each other, thereby maintaining a clear boundary in the feature space; the large model is optimized through the open space loss function, which can reserve more space for unknown categories in the representation space, encouraging the centroids of known categories to be far from the centroids of unknown categories, facilitating the identification of unknown intentions during the inference stage. In summary, the present invention designs the total loss function using the above loss functions, which can enable the large model to better learn the features of known categories while remaining sensitive to unknown categories. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The following further describes in detail the specific embodiments of the present invention with reference to the accompanying drawings. The accompanying drawings provided herein are used to provide a further understanding of the present application and constitute a part of the present application. The same reference numerals are used to represent the same or similar parts in these drawings. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation of the present application.
[0023] Figure 1 It is a flowchart of the method provided by an example of the present invention; Figure 2 It is a flowchart of the method provided by a preferred example of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] The technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0025] In the description of the present invention, it should be noted that the directions or positional relationships indicated by "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the use of ordinal numbers (e.g., "first and second", "first to fourth", etc.) is for differentiating objects and is not limited to this order, and should not be construed as indicating or implying relative importance.
[0026] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "mounted", "connected", and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.
[0027] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0028] In one example, as Figure 1 shown, a multi-granularity financial text noise open classification method based on a large model, the method includes a training step:
[0029] S1: Extract semantic features of financial texts.
[0030] In step S1, a pre-trained large model (large language model) is used to extract semantic features of financial texts. The large language model has powerful natural language processing capabilities and high efficiency in capturing complex text relationships, and can be any one of BERT, RoBERTa, ALBERT, ELECTRA, etc.
[0031] Specifically, the extraction of semantic features of financial texts is to extract semantic features from the input sentences and labels through a large model, and then output feature vectors and labels as the input for the intention classification task in the financial field, which is used for subsequent classifiers (classification models) to classify. Optionally, the extraction of semantic features of financial texts includes the following sub-steps:
[0032] S11: Input data processing, including text preprocessing sub-steps and word segmentation and encoding processing sub-steps; among them, the text preprocessing sub-steps include: preprocessing the input financial statements, including removing stop words, punctuation, and text standardization, etc. The word segmentation and encoding processing sub-steps include: using the tokenizer of a pre-trained large model (such as BERT) to perform sub-word segmentation on the text, converting the text units (words or sub-words or punctuation marks, etc.) after word segmentation into a unique integer identifier, and adding special tokens [CLS] and [SEP], where [CLS] is the classification token and [SEP] is the separator token. At the same time, generate corresponding position encodings and sentence encodings to prepare for the next feature extraction.
[0033] S12: Calculate the features of the financial text. The pre-trained model (such as BERT) performs word vector calculation to generate the context representation of each text unit. Preferably, a multi-layer self-attention mechanism is used for feature extraction, and transformation and normalization processing are performed through a feed-forward neural network.
[0034] S13: Obtain the final semantic features of the financial text, including a feature aggregation sub-step and a feature vectorization sub-step. Among them, the feature aggregation sub-step includes: using a pooling method (such as average pooling or max pooling) to aggregate the features of all text units, or selecting the hidden state of the [CLS] token as the global feature. The feature vectorization sub-step includes: performing feature normalization processing to obtain the final feature vector, that is, the semantic feature of the financial text, for subsequent classification tasks.
[0035] S2: Perform clustering processing on the semantic features to obtain multiple granular balls with different labels, and calculate the attributes of the multiple granular balls with different labels, including capacity (size), purity, centroid, and radius.
[0036] Among them, the capacity of the granular ball represents the number of samples contained in the granular ball; the centroid of the granular ball represents the mean of all sample features within the granular ball; the radius of the granular ball represents the average Euclidean distance representation of all samples within the granular ball to the centroid; the label of the granular ball represents the label of the category with the highest proportion in the granular ball; the purity of the granular ball represents the proportion of samples within the granular ball that belong to the label
[0037] In this step, all samples corresponding to semantic features (data points represented by semantic features) are initialized as a large granule. It is determined whether to split based on granule purity and capacity. After confirming that splitting is to continue, for the granule to be split, it is classified into multiple granules according to the labels and distances of the samples contained in the granule. Then, splitting judgment and splitting processing are carried out again until no splitting is required, and thus multiple granules with different labels are obtained. Further, by calculating the attributes of multiple granules with different labels, multiple granules with the same label but different attributes are obtained. At this time, each category is finally represented by multiple granules with the same label but different attributes. The centroid of multiple granules with the same label represents multiple centers of the category, and the radius reflects the distribution range of the category, that is: through the multi-granularity representation method, this application effectively depicts the multi-granularity prototype and distribution range of the category, can truly reflect the actual distribution of subclasses within each category, takes into account the distribution differences within the category, and thus improves the representation learning effect, can effectively distinguish the feature representations of different categories and subclasses, and provides a more accurate reference for subsequent noise recognition and classification tasks. At the same time, each category is finally represented by multiple granules with the same label but different attributes, so that each category is not limited to a single boundary, but multiple boundaries are set according to the actual distribution of the data, ensuring high adaptability to various data distributions. This flexible boundary setting significantly improves the accuracy of classification. Especially when dealing with data with non-spherical or heterogeneous distributions, it can effectively distinguish known categories and unknown categories and enhance the robustness of the classification method.
[0038] S3: Classify the samples in the granule into clean samples, in-distribution noise, and out-of-distribution noise according to the position of the sample in the granule (distance from the centroid) and the attributes of the granule, including the following sub-steps: S31: Define granules with granule purity greater than the high purity threshold and granule capacity greater than the high capacity threshold as high-quality granules. The noise classification of the samples in the high-quality granules includes: if the distance from the sample to the centroid is less than the radius of the granule and the sample has the same label as the granule, classify the sample as a clean sample; if the distance from the sample to the centroid is less than the radius of the granule and the sample has a different label from the granule, classify the sample as in-distribution noise.
[0039] Among them, the granule purity represents the proportion of samples belonging to the label in the granule. The high purity threshold is a preset parameter used to determine whether the granule purity is high enough. The user can customize it or set the high purity threshold according to historical experience. Optionally, in this example, the high purity threshold is 0.9. The granule capacity represents the number of samples included in the granule. The high capacity threshold is a preset parameter used to determine whether the granule capacity is large enough. The user can set it according to the data distribution, specifically given according to the number of training samples in each category of the dataset. For example, if there are 100 samples in each class of the training data, the high-capacity threshold of the granule is set to 25 at this time. However, if there are 20 samples in each class, the high-capacity threshold is set to 8. Generally speaking, the more samples each class contains, the larger the value of the high-capacity threshold of the granule. When a granule has a high purity (granule purity > greater than the high-purity threshold ), and the number of samples contained in the granule is large (granule capacity > high-capacity threshold ), it indicates that the granule can represent a subclass of the class corresponding to its label. Further, a high-quality granule means that the granule has a strong representativeness of the class subclass, that is, it can represent the distribution information of the class.
[0040] In step S31, if the distance from the sample to the centroid is less than the radius of the granule and the sample and the granule have the same label, it means that the sample is inside the granule and is consistent with the label of the granule. These samples are not interfered by noise, and this part of the samples is classified as clean samples and added to the clean sample set ; if the distance from the sample to the centroid is less than the radius of the granule and the sample and the granule labels are different, it means that although these samples are inside the granule, their labels are inconsistent with the label of the granule. Then this part of the samples may be noise samples, but still belong to the distribution range of the known class. Therefore, this part of the samples is classified as in-distribution noise and added to the in-distribution noise set .
[0041] S32: Define the granules with a granule purity less than the low-purity threshold and a granule capacity less than the low-capacity threshold as low-quality granules, and classify the samples inside the low-quality granules as out-of-distribution noise.
[0042] Among them, the low-purity threshold is a preset parameter used to determine whether the granule purity is low enough. The user can customize or set the low-purity threshold according to historical experience. Optionally, in this example, the high-purity threshold is 0.5. The low-capacity threshold is a preset parameter used to determine whether the granule capacity is small enough. The user sets it according to the number of training samples in each category of the dataset. If there are 100 samples in each class of the training data, the low-capacity threshold can be set to 5. However, if there are 20 samples in each class, the low-capacity threshold can be set to 3. Generally speaking, the fewer samples each class contains, the smaller the low-capacity threshold. It should be noted that in the present invention, the high-purity threshold > low-purity threshold , and the high-capacity threshold > Low volume threshold . When the purity of a granule is low (granule purity < low purity threshold ), and the number of samples contained in the granule is small (granule volume < low volume threshold ), it indicates that the granule cannot effectively represent the category corresponding to its label. Further, a low-quality granule means that the granule has poor representativeness of the subclass of the category it represents, and the samples it represents are out-of-distribution samples.
[0043] In step S32, when the granule purity is less than the low purity threshold, it indicates that the sample labels in the granule are inconsistent and may contain samples of multiple categories, which means that these samples may not come from the same known category; when the granule volume is less than the low volume threshold, it means that the number of samples in the granule is small and not enough to represent a valid category, which further increases the possibility that these samples belong to an unknown category. Therefore, the samples in the low-quality granule are classified as out-of-distribution noise and added to the out-of-distribution noise set , to avoid misclassifying this part of the samples as known categories.
[0044] In this example, according to the position of the sample in the granule and the attributes of the granule, the samples are classified into clean samples, in-distribution noise, and out-of-distribution noise. Through fine-grained feature analysis, the accuracy of noise recognition is improved, thus maintaining efficient and stable performance in a complex financial environment; at the same time, by identifying and classifying out-of-distribution noise, the out-of-distribution noise can be used as a placeholder for unknown categories for representation learning. In the overall representation space, the representation space of known categories is compressed and the representation space of unknown categories is expanded, avoiding the representation space being completely occupied by known categories and leaving room for the development of potential unknown categories, effectively expanding the recognition ability of unknown categories and ensuring that the representation of known categories is more compact, so as to quickly and accurately classify when unknown intentions appear; in addition, classifying and processing according to the position of the sample in the granule and the attributes of the granule can be applied to data of various distribution types.
[0045] Preferably, before the semantic feature step of financial texts, it further includes:
[0046] S01: Select and construct a pre-trained large language model.
[0047] Select an advanced pre-trained large language model as the basic framework for processing the intention classification of the financial dialogue system. Large models usually include an input layer, a preprocessing module, an embedding layer, an encoding layer, and an output layer, etc.
[0048] S02: Collect and divide financial dialogue intention data, including financial data collection and preprocessing sub-steps, and dataset division sub-steps.
[0049] Among them, the sub-steps of financial data collection and preprocessing include:
[0050] (1) Data collection: Collect user consultation text data from the financial dialogue system to ensure that the data covers a wide range of financial-related topics, such as account inquiries, transaction operations, etc.
[0051] (2) Data annotation: Invite financial experts to carefully annotate the collected text data to clarify the intention category of each consultation text. The annotated intention categories should cover account management, transaction processing, product consultation, complaint handling, etc.
[0052] (3) Data cleaning: Thoroughly clean the collected text data, remove irrelevant information, such as HTML tags, special characters, etc., and correct spelling mistakes, unify terms and expressions to improve data quality.
[0053] Further, in the design of the open intention classification task of the present invention, dataset division plays a key role, aiming to create a test environment that simulates the real scenario, where the test set will contain categories not seen during training. In addition, in-distribution noise and out-of-distribution noise will be deliberately introduced into the training data. Among them, in-distribution noise (IND noise) involves mislabeled known category samples. Out-of-distribution noise (OOD noise) involves unknown samples mislabeled as known categories. At this time, the sub-steps of dataset division include:
[0054] (1) Divide known intention categories and unknown intention categories:
[0055] To effectively train and test the performance of the classification model in dealing with open categories, in the category selection of the present invention, a certain proportion of intentions are selected as known categories and only these categories are used during the training phase. In the category definition, in the experiment, the known categories range from the 1st category to the Nth category, and all unseen or newly introduced intention categories are defined as the N + 1th category.
[0056] (2) Construct the training set:
[0057] From the determined known category data, 80% of the data is randomly selected as the training set. For the IND noise design, by randomly swapping the labels of some known samples with other known class labels, the phenomenon of mislabeling is simulated. For the OOD noise design, samples that are similar in appearance to the known categories but belong to unknown categories are introduced and these samples are mislabeled as known category labels to enhance the model's adaptability to new categories.
[0058] (3) Construct the validation set:
[0059] In terms of data distribution: 10% of the known-class data is set aside as the validation set for model tuning and performance evaluation. In terms of noise management, ensure that the validation set does not contain any form of noise to guarantee the accuracy and reliability of the evaluation results.
[0060] (4)Construct the test set:
[0061] In terms of data composition, the test set consists of the remaining 10% of the known-class data and the unknown-class data, and is used to finally evaluate the model's ability to recognize new intents. In terms of noise control, the test set also does not contain noise to ensure a fair and unbiased evaluation of the model's ability to recognize unknown intents.
[0062] In one example, extracting the semantic features of financial texts includes:
[0063] Freeze the underlying parameters of the large model, and use the training dataset of financial texts to train the top-layer parameters of the large model. The top layer of the large model includes an embedding layer and an encoding layer;
[0064] Use the trained large model to extract the semantic features of financial texts.
[0065] To adapt to the intent classification task in the financial field, the present invention adopts a specific strategy in the representation learning stage, that is, freezing the underlying parameters of the pre-trained large model and only training the top-layer parameters. This strategy allows the use of the general language features already mastered by the pre-trained large model, while adapting to the specific semantic requirements of the financial field through fine-tuning of the top layer. In this example, it involves the initialization settings of the large model, including the initialization of the learning rate and optimizer, parameter initialization, and the setting of batch size and early stopping strategy. Specifically, for the initialization of the learning rate and optimizer, in the model initialization stage, set the initial learning rate to r1, and use the Adam optimizer for parameter optimization. For parameter initialization, in this example, keep the underlying parameters unchanged as the values in the pre-trained model, while the top-layer weight parameters are randomly initialized using the Xavier initialization method to promote the learning efficiency of the large model in the open classification task of financial text noise. For the batch size and early stopping strategy, set an appropriate batch size and configure the early stopping strategy to prevent overfitting and ensure timely stopping of training when the performance on the validation set no longer improves.
[0066] In this example, fine-tuning the embedding layer and encoding layer of the large model can adapt to the specific needs of the financial field and enhance the model's understanding of financial proprietary terms. Initialize the optimal validation score to 0, and the optimal model is the initial model.
[0067] In one example, clustering the semantic features to obtain multiple granule balls with different labels includes the following sub-steps:
[0068] S21: Initialize all samples corresponding to the semantic features as one granule ball;
[0069] S22: Calculate the attributes of the granules (capacity, purity, centroid, and radius), determine whether to perform granule splitting based on the granule attributes. If so, randomly select samples with each different label from the granules as the initial centroids of the new granules, calculate the distances from all samples to the initial centroids, and assign the samples to the new granules represented by the nearest centroids, thereby obtaining multiple granules with different labels.
[0070] Among them, a granule is an adaptive clustering unit that can characterize the data distribution at multiple granularity levels. Let the sample set , where is the feature vector of the sample, is the class label of the sample, represents the sample index. By clustering this sample set , a set consisting of a series of granules can be obtained, where each granule is composed of samples, is the number of granules, is the granule index.
[0071] Optionally, in step S22, determining whether to perform granule splitting based on the granule attributes includes:
[0072] If the granule purity is less than the first purity threshold and if the granule capacity is greater than the first capacity threshold, then granule splitting is performed.
[0073] Among them, the first purity threshold is a preset parameter, which can be customized by the user or set according to historical experience for the low purity threshold. Exemplarily, in this example, the first purity threshold is greater than the low purity threshold. The first capacity threshold is a preset parameter, which is set by the user according to the number of training samples in each category in the dataset. In this example, the first capacity threshold is between the low capacity threshold and the high capacity threshold. If the granule purity is less than the first purity threshold and the granule capacity is greater than the first capacity threshold, the data points inside the granule are not consistent enough and may contain multiple different categories or features. By splitting this granule, the data points can be redistributed into more pure sub - granules, thereby improving the overall quality of clustering. If the granule capacity is greater than the first capacity threshold, it means that the granule contains too many samples and may contain multiple different sub - patterns or categories. In this case, the sample distribution inside the granule may be relatively complex, resulting in a low purity of the granule. Therefore, at this time, the granule containing multiple sub - patterns can be decomposed into multiple smaller granules, and each small granule can better represent a specific sub - pattern or category, thereby improving the purity and classification effect of the granule. Further, if the granule purity is greater than or equal to the first purity threshold, or the granule capacity is less than or equal to the first capacity threshold, then the granule is added to the final granule set .
[0074] Further, when it is determined according to the particle ball attribute that particle ball splitting processing is required, the following sub-steps are included:
[0075] (1) Centroid selection; specifically, calculate the number of different labels included in the particle ball , , represents the unique value function; among the samples included in the particle ball , randomly select a representative point for each sample corresponding to a different label as the initial centroid of the new particle ball.
[0076] (2) Distance calculation; specifically, calculate the distances from all samples in the particle ball to the initial centroids of the new particle balls, such as the Euclidean distance.
[0077] (3) Particle ball division. Specifically, for each sample, according to the distance between the sample and each initial centroid, allocate the sample to the new particle ball where the nearest centroid is located, and finally form new particle balls, and continue to judge whether the new particle balls need to continue splitting until all particle balls meet the condition that the purity is greater than or equal to the first purity threshold or the capacity is less than or equal to the first capacity threshold, then stop splitting, and finally obtain particle balls to form a particle ball set , and this particle ball set can represent the multi-granularity characteristics of the data.
[0078] In one example, the method further includes classifying the samples in the particle ball as uncertain samples:
[0079] When classifying the samples in the high-quality particle ball as noise, if the distance from the sample to the centroid is greater than the radius of the particle ball, classify the sample as an uncertain sample; and / or,
[0080] Classify the samples in the particle balls with particle ball purity greater than the low purity threshold and less than the high purity threshold as uncertain samples.
[0081] Among them, the uncertain samples are samples with uncertain noise types. When classifying the samples in the high-quality particle ball as noise, if the distance from the sample to the centroid is greater than the radius of the particle ball, it means that the sample is far from the center of the particle ball and exceeds the radius range of the particle ball, indicating that the noise type of the sample is uncertain. If the purity of the particle ball is greater than the low purity threshold and less than the high purity threshold, it means that there may be multiple types of samples mixed inside the particle ball, and the noise category attribution of the sample is not clear enough. Therefore, if the low purity threshold < particle ball purity < high purity threshold , the sample in the granule is classified as an uncertain sample and added to the uncertain sample set. Preferably, in this example, the above two uncertain sample classification methods are simultaneously used to classify the samples in the granule.
[0082] Preferably, the above examples are combined. After classifying the samples in the granule into clean samples, in-distribution noise, out-of-distribution noise, and uncertain samples according to the position of the sample in the granule and the attributes of the granule, centroid collection processing is performed. Specifically, the centroids of high-quality granules and low-quality granules are collected as follows: ; ;
[0083] Among them, is the centroid of all high-quality granules, representing the center of the known class; is the centroid of all low-quality granules, representing the center of the unknown class granules.
[0084] In one example, after classifying the samples in the granule into clean samples and in-distribution noise, it further includes:
[0085] S41: Calculate the first cross-entropy loss function of the clean samples : ; Among them, represents the logarithmic function.
[0086] S42: Modify the in-distribution noise label to the label of the granule where it is located to correct the bias caused by the wrong label, and then calculate the second cross-entropy loss function : ; Among them, represents the th granule category.
[0087] S43: Use the first cross-entropy loss function and the second cross-entropy loss function to construct the first total loss function, and then optimize the classification model. Among them, the classification model can be a neural network based on deep learning, such as a convolutional neural network, a recurrent neural network, a BERT model, etc.
[0088] Preferably, for uncertain samples, since their true labels cannot be accurately determined, they are directly discarded and not included in the calculation of the loss function.
[0089] In one example, after classifying the samples in the granule into out-of-distribution noise, it further includes:
[0090] S41': Calculate the intra-class loss function and the inter-class loss function based on the centroid of high-quality granules and the labels of high-quality granules; calculate the open space loss function based on the centroid of high-quality granules and the centroid of low-quality granules.
[0091] Specifically, the intra-class loss function has the following calculation expression: ; where represents the indicator function; represents the class of the th granule; represents the centroid of the th granule; represents the centroid of the th granule.
[0092] Furthermore, the inter-class loss function has the following calculation expression: ; where represents a small constant.
[0093] Furthermore, the open space loss function has the following calculation expression: .
[0094] S42': Use the first cross-entropy loss function, the second cross-entropy loss function, the intra-class loss function, the inter-class loss function, and the open space loss function to construct the second total loss function : ;
[0095] Use the second total loss function to perform backpropagation on the classification model, calculate the gradient, and update the trainable parameters of the model to gradually improve the performance of the model in open intent classification.
[0096] Preferably, after optimizing the model using the loss function, it further includes model update processing, including:
[0097] (1) Validation set evaluation: After each round of training, calculate the evaluation index score of the model on the validation set to evaluate the current performance of the model.
[0098] (2) Update the optimal validation score and the optimal model: If the current evaluation score of the classification model is higher than the historical optimal score, set the score and the corresponding model parameters as the new optimal results.
[0099] (3) Continuing training: Subsequent training iterations are based on the parameters of the current optimal model, continuously improving the adaptability of the model for classification in complex noise environments.
[0100] Further, after the model update process, it also includes training stop judgment and boundary preservation. Specifically, if the validation scores do not improve for a consecutive number of times, such as 10 times, the early stopping strategy is triggered to stop the representation learning and save the optimal model parameters. At the same time, the centroids and their radii of each known category represented by high-quality granules are recorded as the final category boundaries.
[0101] Combining the above examples, the preferred open classification method of the present invention is obtained. As Figure 2 shown, the method includes the following steps at this time: S100: Select and construct a pre-trained large language model and initialize the large model; S200: Extract the semantic features of financial texts; S300: Perform clustering on the semantic features to obtain granules with multiple different labels, and calculate the attributes of the granules with multiple different labels; S400: Classify the samples in the granules as clean samples, in-distribution noise, out-of-distribution noise, and uncertainty samples according to the position of the samples in the granules and the attributes of the granules; S500: Calculate the loss functions of different types of samples and noise, and then construct a total loss function, and use the total loss function to optimize the model parameters; specifically, calculate the first cross-entropy loss function of clean samples; after modifying the in-distribution noise labels to the labels of the granules where they are located, calculate the second cross-entropy loss function; calculate the intra-class loss function and inter-class loss function according to the centroids and their corresponding labels of high-quality granules, and calculate the open space loss function according to the centroids of high-quality granules and the centroids of low-quality granules; then use the first cross-entropy loss function, the second cross-entropy loss function, the intra-class loss function, the inter-class loss function, and the open space loss function to construct a second total loss function, and use the second total loss function to optimize the classification model.
[0102] S600: Model update process. If the iteration stop strategy is triggered, stop learning and save the optimal model parameters. At the same time, record the centroids and their radii of each known category represented by high-quality granules as the final category boundaries.
[0103] In an example, after the model training and validation are completed, the test set data is input into the optimal model. By extracting the semantic features of the financial text to be classified, and then using the centroids and radii of each known category saved to perform open intent classification, that is, to test or apply the open intent classification method. The method includes the following steps at this time: S100’: Extract the semantic features of the financial text to be classified; S200': For each sample to be classified corresponding to the semantic features, calculate the Euclidean distances between the sample to be classified and all the centroids of the known categories, then find the first centroid of the known category that is closest to each sample to be classified, and compare this distance with the first centroid of the known category; at this time, the centroids and radii of the known categories represented by high-quality granulocytes are used as the final category boundaries for the known categories. S300': If the distance from the sample to be classified to the first centroid of the known category is less than or equal to the radius corresponding to the first centroid of the known category, classify the sample to be classified as the known category corresponding to the first centroid of the known category; if the distance from the sample to be classified to the first centroid of the known category is greater than the radius corresponding to the first centroid of the known category, classify the sample to be classified as an unknown category.
[0104] Through the above method, the present invention can make an efficient and accurate distinction between known categories and unknown categories during the inference stage, greatly improving the practicality and robustness of open intent classification.
[0105] The above specific implementation manners are detailed descriptions of the present invention. It cannot be determined that the specific implementation manners of the present invention are only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions and substitutions can still be made, which should all be regarded as belonging to the protection scope of the present invention.
Claims
1. A multi-granularity financial text noise open classification method based on a large model, characterized by: The method comprises the steps of training: Extracting semantic features of financial texts; Clustering is performed on the semantic features to obtain multiple spheres with different labels, and the properties of the spheres with different labels are calculated, including capacity, purity, centroid and radius; According to the position of the sample in the sphere and the properties of the sphere, the samples in the sphere are classified into clean samples, noise within the distribution, and noise outside the distribution, including the following sub-steps: The spheres whose purity is greater than the high purity threshold and whose capacity is greater than the high capacity threshold are defined as high-quality spheres. The noise classification of samples in the high-quality spheres includes: if the distance from the sample to the centroid is less than the radius of the sphere, and the sample and the sphere have the same label, the sample is classified as a clean sample; If the distance from the sample to the centroid is less than the radius of the sphere, and the sample and sphere labels are different, the sample is classified as in-distribution noise; The spheres whose purity is less than the low purity threshold and whose capacity is less than the low capacity threshold are defined as low-quality spheres, and the samples in the low-quality spheres are classified as out-of-distribution noise.
2. The multi-granularity financial text noise open classification method based on a large model according to claim 1 is characterized in that: The extracting of semantic features of financial texts includes: Freeze the bottom-level parameters of the big model and use the training data set of financial text to train the top-level parameters of the big model. The top level of the big model includes the embedding layer and the encoding layer. The trained large model is used to extract semantic features of financial texts.
3. The multi-granularity financial text noise open classification method based on a large model according to claim 1 is characterized in that: The method of clustering the semantic features to obtain multiple spheres with different labels includes the following sub-steps: Initialize the samples corresponding to all semantic features into a sphere; Calculate the properties of the sphere, and determine whether to split the sphere according to the properties of the sphere. If so, randomly select each sample with a different label from the sphere as the initial centroid of the new sphere, calculate the distance from all samples to the initial centroid, and assign the samples to the new sphere represented by the nearest centroid, thereby obtaining multiple spheres with different labels.
4. The multi-granularity financial text noise open classification method based on a large model according to claim 3 is characterized in that: The step of judging whether to perform a pellet splitting process according to pellet properties includes: If the purity of the ball is less than the first purity threshold, and if the capacity of the ball is greater than the first capacity threshold, the ball is split.
5. The multi-granularity financial text noise open classification method based on a large model according to claim 1 is characterized in that: The method further includes classifying the sample in the sphere as an uncertain sample: When performing noise classification on samples in a high-quality sphere, if the distance from the sample to the centroid is greater than the radius of the sphere, the sample is classified as an uncertain sample; and / or, Classify the samples in the pellets whose purity is greater than the low purity threshold and less than the high purity threshold as uncertain samples; Uncertain samples are discarded and not included in the calculation of the loss function.
6. The multi-granularity financial text noise open classification method based on a large model according to claim 1 is characterized in that: After classifying the samples in the sphere into clean samples and noise in the distribution, it also includes: Calculate the first cross entropy loss function of clean samples; After modifying the noise labels in the distribution to the labels of the spheres they belong to, calculate the second cross entropy loss function; The first cross entropy loss function and the second cross entropy loss function are used to construct the first total loss function, and then the classification model is optimized.
7. The multi-granularity financial text noise open classification method based on a large model according to claim 6 is characterized in that: After classifying the samples in the sphere as out-of-distribution noise, it also includes: According to the centroid of high-quality spheres and the labels of high-quality spheres, the intra-class loss function and the inter-class loss function are calculated; according to the centroid of high-quality spheres and the centroid of low-quality spheres, the open space loss function is calculated; The second total loss function is constructed using the first cross entropy loss function, the second cross entropy loss function, the intra-class loss function, the inter-class loss function, and the open space loss function, and then the classification model is optimized.
8. The multi-granularity financial text noise open classification method based on a large model according to claim 1 is characterized in that: The method further comprises the step of applying: Extract semantic features of financial text to be classified; For the samples to be classified corresponding to each semantic feature, determine the centroid of the first known category that is closest to it; The known categories use the centroid and radius of each known category represented by a high-quality particle sphere as the final category boundary; If the distance between the sample to be classified and the centroid of the first known category is less than or equal to the radius corresponding to the centroid of the first known category, the sample to be classified is classified into the known category corresponding to the centroid of the first known category; If the distance between the sample to be classified and the centroid of the first known category is greater than the radius corresponding to the centroid of the first known category, the sample to be classified is classified as an unknown category.
Citation Information
Patent Citations
Image label noise learning method based on granular ball calculation and contrast learning
CN119478545A
Long-tail learning image classification method with in-distribution and out-distribution noise labels
CN119580004A
Device and method for multiparameter measurements of microparticles in a fluid
US20120296570A1
Cited By
Large model-based noise text open intention classification method and system
CN120578766A
Image noise mark feature selection method and system, storage medium and computer
CN120894642A