Open Classification Method for Multi-Granularity Financial Text Noise Based on Large Models

Through the multi-grained financial text noise open classification method based on large models, using particle-spheric clustering and loss function optimization, the adaptability and accuracy of financial text noise classification method in complex environments is solved, and efficient identification and classification of unknown intentions is achieved.

CN120123518BActive Publication Date: 2025-07-08SOUTHWESTERN UNIV OF FINANCE & ECONOMICS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510608156.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-13
Publication Date
2025-07-08
Estimated Expiration
2045-05-13

AI Technical Summary

Technical Problem

Existing financial text noise classification methods are insufficient in handling complex noise environments, especially the limited ability to identify unknown intention categories and are difficult to obtain by relying on large-scale high-quality annotation datasets.

Method used

The multi-grained financial text noise open classification method based on large models is adopted. By extracting semantic features and performing clustering processing, multiple particle balls with different labels are formed, the particle ball properties are calculated, and the samples are classified as clean samples, internal distribution noise and external distribution noise are classified according to the sample position and particle ball properties, and the classification model is optimized using cross-entropy loss functions.

Benefits of technology

It improves the accuracy of noise recognition and the robustness of classification, can effectively distinguish known and unknown categories, enhances the adaptability to financial text and classification accuracy, and is suitable for data environments of various distribution types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120123518B_ABST
    Figure CN120123518B_ABST
Patent Text Reader

Abstract

The present invention discloses an open classification method for multi-granularity financial text noise based on large models, belonging to the technical field of text noise classification, including: extracting semantic features of financial texts; performing clustering processing on the semantic features to obtain multiple granule balls with different labels, and calculating the attributes of the multiple granule balls with different labels; classifying the samples within the granule balls as clean samples, in-distribution noise, and out-of-distribution noise according to the positions of the samples in the granule balls and the attributes of the granule balls. By performing clustering processing to obtain multiple granule balls with different labels and calculating the attributes of each granule ball, each category is finally represented by multiple granule balls with the same label but different attributes at this time, which can effectively reflect the multi-granularity prototypes and distribution ranges of the categories and improve the representation learning effect. At the same time, classifying according to the positions of the samples in the granule balls and the attributes of the granule balls can be applicable to data environments of various distribution types.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of text noise classification, and particularly to an open classification method for multi-granularity financial text noise based on a large model. Background Art

[0002] With the rapid development of fintech, financial dialogue systems have become a key technology for banks and financial service institutions to provide customer service and intelligent consultation. These dialogue systems need to be able to accurately identify and classify the query intentions of users, including regular financial transaction requests and non-predefined new intentions. Traditional intention classification systems can usually only identify the intention categories defined during the training process, and have insufficient processing capabilities for newly emerging or unknown intention categories, which limits the adaptability and utility of noise classification methods.

[0003] In addition, existing open intention classification methods usually rely on large-scale and high-quality labeled data sets, which are often difficult to obtain in practical applications. In practical applications, the data sets often contain noise, and label noise can usually be divided into two categories: in-distribution (IND) noise and out-of-distribution (OOD) noise. IND noise occurs when an intention belonging to a predefined known category is mislabeled as another known category, while OOD noise occurs when an unknown intention is mislabeled as a predefined known category. In most real-world scenarios, IND noise and OOD noise exist simultaneously, which can be referred to as complex noise.

[0004] There are few existing noise processing methods for open intention classification, and there are a series of defects in existing noise processing research, which limit its direct application in the representation learning of open intention classification. The specific defects include:

[0005] (1) Limitations of prototype representation: Current distance-based noise processing methods usually use category prototypes to represent category distributions and determine whether a sample is noise based on the distance between the sample and the prototype. On the one hand, these prototypes can only represent the central position of the category and cannot comprehensively reflect the overall distribution of the category, resulting in inaccurate distance-based noise detection. On the other hand, traditional clustering methods such as K-means clustering preset the number of clusters and cannot truly reflect the actual distribution of subclasses within each intention, thus affecting the effect of representation learning.

[0006] (2) Insufficiency of noise sample identification methods: Current sample selection methods mainly rely on fixed distance thresholds or probability criteria to identify clean samples and IND noise. This coarse-grained global selection criterion fails to consider the distribution differences within the category and cannot be effectively applied to data environments of various distribution types. Therefore, the adaptability and scalability of these methods in variable practical application scenarios are limited. Summary of the Invention

[0007] The object of the present invention is to overcome the problems of the prior art, and provides an open classification method for multi-granularity financial text noise based on a large model.

[0008] The object of the present invention is achieved by the following technical solutions: An open classification method for multi-granularity financial text noise based on a large model, the method includes a training step:

[0009] Extract the semantic features of financial texts;

[0010] Perform clustering processing on the semantic features to obtain granule balls with multiple different labels, and calculate the attributes of the granule balls with multiple different labels, including capacity, purity, centroid and radius;

[0011] According to the position of the sample in the granule ball and the attributes of the granule ball, classify the samples in the granule ball into clean samples, in-distribution noise, and out-of-distribution noise, including the following sub-steps:

[0012] Define the granule balls with granule ball purity greater than the high purity threshold and granule ball capacity greater than the high capacity threshold as high-quality granule balls. The noise classification of the samples in the high-quality granule balls includes: if the distance from the sample to the centroid is less than the radius of the granule ball, and the sample and the granule ball have the same label, classify the sample as a clean sample; if the distance from the sample to the centroid is less than the radius of the granule ball, and the sample and the granule ball label are different, classify the sample as in-distribution noise;

[0013] Define the granule balls with granule ball purity less than the low purity threshold and granule ball capacity less than the low capacity threshold as low-quality granule balls, and classify the samples in the low-quality granule balls as out-of-distribution noise.

[0014] In one example, the extraction of the semantic features of financial texts includes:

[0015] Freeze the underlying parameters of the large model, and use the training data set of financial texts to train the top-layer parameters of the large model. The top layer of the large model includes an embedding layer and an encoding layer;

[0016] Use the trained large model to extract the semantic features of financial texts.

[0017] In one example, the clustering processing of the semantic features to obtain granule balls with multiple different labels includes the following sub-steps:

[0018] Initialize all samples corresponding to the semantic features as one granule ball;

[0019] Calculate the attributes of the grain balls, and determine whether to perform grain ball splitting processing according to the grain ball attributes. If so, randomly select samples with each different label from the grain balls as the initial centroids of the new grain balls, calculate the distances from all samples to the initial centroids, and assign the samples to the new grain balls represented by the nearest centroids, thereby obtaining multiple grain balls with different labels.

[0020] In one example, the determining whether to perform grain ball splitting processing according to the grain ball attributes includes:

[0021] If the purity of the grain ball is less than the first purity threshold and the capacity of the grain ball is greater than the first capacity threshold, then perform grain ball splitting.

[0022] In one example, the method further includes classifying the samples in the grain balls as uncertain samples:

[0023] When classifying the samples in the high-quality grain balls as noise, if the distance from the sample to the centroid is greater than the radius of the grain ball, classify the sample as an uncertain sample; and / or,

[0024] Classify the samples in the grain balls with a grain ball purity greater than the low purity threshold and less than the high purity threshold as uncertain samples;

[0025] Discard the uncertain samples and do not include them in the calculation of the loss function.

[0026] In one example, after classifying the samples in the grain balls as clean samples and in-distribution noise, it further includes:

[0027] Calculate the first cross-entropy loss function for the clean samples;

[0028] After modifying the in-distribution noise labels to the labels of the grain balls where they are located, calculate the second cross-entropy loss function;

[0029] Use the first cross-entropy loss function and the second cross-entropy loss function to construct the first total loss function, and then perform optimization processing on the classification model.

[0030] In one example, after classifying the samples in the grain balls as out-of-distribution noise, it further includes:

[0031] According to the centroids of the high-quality grain balls and the labels of the high-quality grain balls, calculate the within-class loss function and the between-class loss function; according to the centroids of the high-quality grain balls and the centroids of the low-quality grain balls, calculate the open space loss function;

[0032] Use the first cross-entropy loss function, the second cross-entropy loss function, the within-class loss function, the between-class loss function, and the open space loss function to construct the second total loss function, and then perform optimization processing on the classification model.

[0033] It should be further noted that the technical features corresponding to the above examples can be combined or replaced with each other to form a new technical solution.

[0034] In one example, the method further includes an application step:

[0035] Extract the semantic features of the financial text to be classified;

[0036] For each sample to be classified corresponding to the semantic features, determine the first known class centroid with the closest distance; the known classes use the centroids and radii of each known class represented by high-quality granular balls as the final class boundaries;

[0037] If the distance from the sample to be classified to the first known class centroid is less than or equal to the radius corresponding to the first known class centroid, classify the sample to be classified as the known class corresponding to the first known class centroid; if the distance from the sample to be classified to the first known class centroid is greater than the radius corresponding to the first known class centroid, classify the sample to be classified as an unknown class.

[0038] Compared with the prior art, the beneficial effects of the present invention are:

[0039] 1. In one example, multiple granular balls with different labels are obtained through clustering processing, and the attributes of each granular ball are calculated. At this time, each class is finally represented by multiple granular balls with the same label but different attributes. The centroid of multiple granular balls with the same label represents multiple centers of the class, and the radius reflects the distribution range of the class, that is: the present application effectively characterizes the multi-granularity prototype and distribution range of the class through the multi-granularity representation method, can truly reflect the actual distribution of subclasses within each class, considers the distribution differences within the class, and thus improves the representation learning effect, can effectively distinguish the feature representations of different classes and subclasses, and provides a more accurate reference for subsequent noise recognition and classification tasks. At the same time, each class is finally represented by multiple granular balls with the same label but different attributes, so that each class is not limited to a single boundary, but multiple boundaries are set according to the actual distribution of the data, ensuring high adaptability to various data distributions. This flexible boundary setting significantly improves the accuracy of classification. Especially when dealing with data with non-spherical or heterogeneous distributions, it can effectively distinguish known classes and unknown classes and enhance the robustness of the classification method.

[0040] Furthermore, the samples are classified into clean samples, in-distribution noise, and out-of-distribution noise according to the position of the samples in the granules and the attributes of the granules. Through fine-grained feature analysis, the accuracy of noise recognition is improved, thus maintaining efficient and stable performance in a complex financial environment. At the same time, by identifying and classifying out-of-distribution noise, the out-of-distribution noise can be used as a placeholder for unknown classes for representation learning. In the overall representation space, the representation space of known classes is compressed and the representation space of unknown classes is expanded, avoiding the representation space being completely occupied by known classes and leaving room for the development of potential unknown classes, effectively expanding the recognition ability of unknown classes and ensuring that the representation of known classes is more compact, so that accurate classification can be quickly achieved when unknown intentions appear. In addition, classifying according to the position of the samples in the granules and the attributes of the granules can be applied to data environments of various distribution types.

[0041] 2. In one example, freezing the underlying parameters of the pre-trained large model can utilize the general language features already mastered by the pre-trained large model; training the top-level parameters of the large model enables the large model to capture the nuances in financial texts, enhances the large model's understanding ability of the unique expressions in financial texts, and improves the adaptability of the large model to text and speech in the financial field.

[0042] 3. In one example, granule splitting according to the granule purity threshold can further subdivide the granules containing samples of multiple categories, thus reducing the possibility of misclassification; granule splitting according to the capacity threshold of the granules can split the granules with too large a capacity, avoiding over-aggregation of samples in the category and further improving the classification accuracy.

[0043] 4. In one example, the samples located outside the granule radius are classified as uncertain noise, and the loss function is not calculated for them in subsequent model training to avoid misjudging their noise categories and affecting subsequent representation learning.

[0044] 5. In one example, optimizing the classification model through the first cross-entropy loss function can effectively measure the difference between the predicted value and the true value, thereby guiding the optimization of the classification model to consolidate the learning performance of the classification model for correct standard samples; modifying the in-distribution noise labels to the labels of the corresponding granular balls and then calculating the second cross-entropy loss function to optimize the large model can correct the bias caused by incorrect labels; optimizing the large model through the intra-class loss function can ensure that the sample distribution within the same category is more compact, making the centroids of the granular balls with the same label closer to each other, thereby improving the consistency of the features of the intra-class samples; optimizing the large model through the inter-class loss function can increase the distinguishability between different categories, making the centroids of the granular balls with different labels farther from each other, thereby maintaining a clear boundary in the feature space; optimizing the large model through the open-space loss function can reserve more space for unknown categories in the representation space, encouraging the centroids of known categories to be far from the centroids of unknown categories, which is convenient for identifying unknown intentions during the inference stage. In summary, by using the above loss functions to design the total loss function, the present invention can enable the large model to better learn the features of known categories while remaining sensitive to unknown categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The following further elaborates on the specific embodiments of the present invention in conjunction with the accompanying drawings. The accompanying drawings provided herein are used to provide a further understanding of the present application and form a part of the present application. The same reference numerals are used to represent the same or similar parts in these drawings. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application.

[0046] Figure 1 It is a flowchart of the method provided for an example of the present invention;

[0047] Figure 2 It is a flowchart of the method provided for a preferred example of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] The technical solutions of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0049] In the description of the present invention, it should be noted that the directions or positional relationships indicated by "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. are based on the directions or positional relationships shown in the drawings. This is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it should not be construed as a limitation to the present invention. In addition, the use of ordinal numbers (e.g., "first and second", "first to fourth", etc.) is for distinguishing objects and is not limited to this order, and should not be construed as indicating or implying relative importance.

[0050] In the description of the present invention, it should be noted that unless otherwise clearly specified and limited, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those of ordinary skill in the art, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0051] In addition, the technical features involved in different embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.

[0052] In one example, as Figure 1 shown, a multi-granularity financial text noise open classification method based on a large model, the method includes a training step:

[0053] S1: Extract the semantic features of the financial text.

[0054] In step S1, a pre-trained large model (large language model) is used to extract the semantic features of the financial text. The large language model has powerful natural language processing capabilities and high efficiency in capturing complex text relationships, and can be any one of BERT, RoBERTa, ALBERT, ELECTRA, etc.

[0055] Specifically, the extraction of the semantic features of the financial text is to extract the semantic features from the input sentence and label through the large model, and then output the feature vector and label as the input for the intention classification task in the financial field, which is used for subsequent classification by a classifier (classification model). Optionally, the extraction of the semantic features of the financial text includes the following sub-steps:

[0056] S11: Input data processing, including text preprocessing sub-steps and tokenization and encoding processing sub-steps; among them, the text preprocessing sub-steps include: preprocessing the input financial statements, including removing stop words, punctuation, and performing text normalization, etc. The tokenization and encoding processing sub-steps include: using the tokenizer of a pre-trained large model (such as BERT) to perform sub-word segmentation on the text, converting the tokenized text units (words or sub-words or punctuation marks, etc.) into a unique integer identifier, and adding special tokens [CLS] and [SEP], where [CLS] is the classification token and [SEP] is the separator token. At the same time, generate corresponding position encodings and sentence encodings to prepare for the next feature extraction.

[0057] S12: Calculate the features of the financial text. A pre-trained model (such as BERT) is used to calculate word vectors to generate the context representation of each text unit. Preferably, a multi-layer self-attention mechanism is used for feature extraction, and transformation and normalization processing are performed through a feed-forward neural network.

[0058] S13: Obtain the final semantic features of the financial text, including a feature aggregation sub-step and a feature vectorization sub-step. Among them, the feature aggregation sub-step includes: using a pooling method (such as average pooling or max pooling) to aggregate the features of all text units, or selecting the hidden state of the [CLS] token as the global feature. The feature vectorization sub-step includes: performing feature normalization processing to obtain the final feature vector, that is, the semantic feature of the financial text, for subsequent classification tasks.

[0059] S2: Perform clustering processing on the semantic features to obtain multiple granule balls with different labels, and calculate the attributes of the multiple granule balls with different labels, including capacity (size), purity, centroid, and radius.

[0060] Among them, the granule ball capacity represents the number of samples contained in the granule ball; the granule ball centroid represents the mean of all sample features within the granule ball; the granule ball radius represents the average Euclidean distance representation of all samples within the granule ball to the centroid; the granule ball label represents the label of the category with the highest proportion in the granule ball; the granule ball purity represents the proportion of samples within the granule ball that belong to the label

[0061] ​In this step, all samples corresponding to semantic features (data points represented by semantic features) are initialized as a large granule. It is determined whether to split according to the granule purity and capacity. After confirming that splitting is to continue, for the granule to be split, it is classified into multiple granules according to the labels and distances of the samples contained in it. Then, splitting judgment and splitting processing are carried out again until no splitting is required, and multiple granules with different labels are obtained. Further, by calculating the attributes of multiple granules with different labels, multiple granules with the same label but different attributes are obtained. At this time, each category is finally represented by multiple granules with the same label but different attributes. The centroid of multiple granules with the same label represents multiple centers of the category, and the radius reflects the distribution range of the category, that is: This application effectively characterizes the multi-granularity prototype and distribution range of categories through a multi-granularity representation method, can truly reflect the actual distribution of sub-categories within each category, takes into account the distribution differences within the category, and thus improves the representation learning effect, can effectively distinguish the feature representations of different categories and sub-categories, and provides a more accurate reference for subsequent noise recognition and classification tasks. At the same time, each category is finally represented by multiple granules with the same label but different attributes, so that each category is not limited to a single boundary, but multiple boundaries are set according to the actual distribution of the data, ensuring high adaptability to various data distributions. This flexible boundary setting significantly improves the accuracy of classification. Especially when dealing with data with non-spherical or heterogeneous distributions, it can effectively distinguish known categories and unknown categories and enhance the robustness of the classification method.

[0062] S3: Classify the samples in the granule as clean samples, in-distribution noise, and out-of-distribution noise according to the position of the sample in the granule (the distance from the centroid) and the attributes of the granule, including the following sub-steps:

[0063] S31: Define granules with granule purity greater than the high purity threshold and granule capacity greater than the high capacity threshold as high-quality granules. The noise classification of the samples in the high-quality granules includes: If the distance from the sample to the centroid is less than the radius of the granule and the sample has the same label as the granule, classify the sample as a clean sample; if the distance from the sample to the centroid is less than the radius of the granule and the sample has a different label from the granule, classify the sample as in-distribution noise.

[0064] Among them, the granule purity represents the proportion of samples belonging to the label in the granule, and the high purity threshold is a preset parameter used to judge whether the granule purity is high enough. The user can customize it or set the high purity threshold according to historical experience. Optionally, in this example, the high purity threshold is 0.9. The granule capacity represents the number of samples included in the granule, and the high capacity threshold is a preset parameter used to determine whether the granule capacity is large enough. Users can set it according to the data distribution, specifically given based on the number of training samples in each category of the dataset. For example, if there are 100 samples in each category of the training data, the high-capacity threshold of the granule is set to 25 at this time. However, if there are 20 samples in each category, the high-capacity threshold is set to 8. Generally speaking, the more samples each category contains, the larger the value of the high-capacity threshold of the granule. When a granule has a high purity (granule purity > greater than the high-purity threshold ), and the number of samples contained in the granule is large (granule capacity > high-capacity threshold ), it indicates that this granule can represent a subclass of the class corresponding to its label. Further, a high-quality granule means that the granule has a strong representativeness of the class subclass, that is, it can represent the distribution information of the class.

[0065] In step S31, if the distance from the sample to the centroid is less than the radius of the granule and the sample and the granule have the same label, it indicates that the sample is inside the granule and has the same label as the granule. These samples are not interfered by noise, and this part of the samples is classified as clean samples and added to the clean sample set ; if the distance from the sample to the centroid is less than the radius of the granule and the sample and the granule label are different, it indicates that although these samples are inside the granule, their labels are different from the granule label. Then this part of the samples may be noise samples, but they still belong to the distribution range of the known class, so this part of the samples is classified as in-distribution noise and added to the in-distribution noise set .

[0066] S32: Define the granules with granule purity less than the low-purity threshold and granule capacity less than the low-capacity threshold as low-quality granules, and classify the samples inside the low-quality granules as out-of-distribution noise.

[0067] Among them, the low-purity threshold is a preset parameter used to determine whether the granule purity is low enough. Users can customize or set the low-purity threshold according to historical experience. Optionally, in this example, the high-purity threshold is 0.5. The low-capacity threshold is a preset parameter used to determine whether the granule capacity is small enough. Users set it according to the number of training samples in each category of the dataset. If there are 100 samples in each category of the training data, the low-capacity threshold can be set to 5. However, if there are 20 samples in each category, the low-capacity threshold can be set to 3. Generally speaking, the fewer samples each category contains, the smaller the low-capacity threshold. It should be noted that in the present invention, the high-purity threshold > low-purity threshold , and the high-capacity threshold > Low volume threshold When the purity of a granule is low (granule purity < low purity threshold ), and the number of samples contained in the granule is small (granule volume < low volume threshold ), it indicates that the granule cannot effectively represent the category corresponding to its label. Further, a low-quality granule means that the granule has poor representativeness of the subclass of the category it represents, and the samples it represents are out-of-distribution samples.

[0068] In step S32, when the granule purity is less than the low purity threshold, it means that the sample labels in the granule are inconsistent and may contain samples of multiple categories, indicating that these samples may not come from the same known category; when the granule volume is less than the low volume threshold, it means that the number of samples in the granule is small and not enough to represent an effective category, which further increases the possibility that these samples belong to an unknown category. Therefore, the samples in the low-quality granule are classified as out-of-distribution noise and added to the out-of-distribution noise set , avoiding misclassifying this part of the samples as known categories.

[0069] In this example, according to the position of the sample in the granule and the attributes of the granule, the samples are classified into clean samples, in-distribution noise, and out-of-distribution noise. Through fine-grained feature analysis, the accuracy of noise recognition is improved, thus maintaining efficient and stable performance in a complex financial environment; at the same time, by identifying and classifying out-of-distribution noise, the out-of-distribution noise can be used as a placeholder for unknown categories for representation learning. In the overall representation space, the representation space of known categories is compressed and the representation space of unknown categories is expanded, avoiding the representation space being completely occupied by known categories and leaving room for the development of potential unknown categories, effectively expanding the recognition ability of unknown categories, ensuring that the representation of known categories is more compact, so as to quickly and accurately classify when unknown intentions appear; in addition, classifying and processing according to the position of the sample in the granule and the attributes of the granule can be applied to data of various distribution types.

[0070] Preferably, before the semantic feature step of financial texts, it further includes:

[0071] S01: Select and construct a pre-trained large language model.

[0072] Select an advanced pre-trained large language model as the basic framework for processing the intention classification of the financial dialogue system. Large models usually include an input layer, a preprocessing module, an embedding layer, an encoding layer, and an output layer, etc.

[0073] S02: Collect and partition financial dialogue intention data, including financial data collection and preprocessing sub-steps, and dataset partitioning sub-steps.

[0074] Among them, the sub-steps of financial data collection and preprocessing include:

[0075] (1) Data collection: Collect user consultation text data from the financial dialogue system to ensure that the data covers a wide range of financial-related topics, such as account inquiries, transaction operations, etc.

[0076] (2) Data annotation: Invite financial experts to carefully annotate the collected text data to clarify the intent category of each consultation text. The annotated intent categories should cover account management, transaction processing, product consultation, complaint handling, etc.

[0077] (3) Data cleaning: Thoroughly clean the collected text data, remove irrelevant information such as HTML tags, special characters, etc., correct spelling mistakes, and unify terms and expressions to improve data quality.

[0078] Further, in the design of the open intent classification task of the present invention, dataset partitioning plays a key role, aiming to create a test environment that simulates the real scenario, where the test set will contain categories not seen during training. In addition, in-distribution noise and out-of-distribution noise will be deliberately introduced into the training data. Among them, in-distribution noise (IND noise) involves mislabeled samples of known categories. Out-of-distribution noise (OOD noise) involves unknown samples mislabeled as known categories. At this time, the sub-steps of dataset partitioning include:

[0079] (1) Divide known intent categories and unknown intent categories:

[0080] To effectively train and test the performance of the classification model in dealing with open categories, in the category selection of the present invention, a certain proportion of intents are selected as known categories and only these categories are used during the training phase. In the category definition, in the experiment, the known categories range from the 1st category to the Nth category, and all unseen or newly introduced intent categories are defined as the N + 1st category.

[0081] (2) Construct the training set:

[0082] From the determined known category data, 80% of the data is randomly selected as the training set. For the IND noise design, by randomly swapping the labels of some known samples with other known category labels, the phenomenon of mislabeling is simulated. For the OOD noise design, samples that are similar in appearance to the known categories but belong to unknown categories are introduced and these samples are mislabeled as the labels of known categories to enhance the model's adaptability to new categories.

[0083] (3) Construct the validation set:

[0084] In terms of data allocation: 10% of the known-class data is set aside as the validation set for model tuning and performance evaluation. In terms of noise management, ensure that the validation set does not contain any form of noise to guarantee the accuracy and reliability of the evaluation results.

[0085] (4)Construct the test set:

[0086] In terms of data composition, the test set consists of the remaining 10% of the known-class data and the unknown-class data, and is used to finally evaluate the model's ability to recognize new intents. In terms of noise control, the test set also does not contain noise to ensure a fair and unbiased evaluation of the model's ability to recognize unknown intents.

[0087] In one example, the semantic features of financial texts are extracted, including:

[0088] Freeze the underlying parameters of the large model, and use the training dataset of financial texts to train the top-layer parameters of the large model. The top layer of the large model includes an embedding layer and an encoding layer;

[0089] Use the trained large model to extract the semantic features of financial texts.

[0090] To adapt to the intent classification task in the financial field, the present invention adopts a specific strategy in the representation learning stage, that is, freezing the underlying parameters of the pre-trained large model and only training the top-layer parameters. This strategy allows the use of the general language features already mastered by the pre-trained large model, while adapting to the specific semantic requirements of the financial field through fine-tuning of the top layer. In this example, it involves the initialization settings of the large model, including the initialization of the learning rate and optimizer, parameter initialization, and the setting of batch size and early stopping strategy. Specifically, for the initialization of the learning rate and optimizer, at the model initialization stage, set the initial learning rate to r1, and use the Adam optimizer for parameter optimization. For parameter initialization, in this example, keep the underlying parameters unchanged as the values in the pre-trained model, while the top-layer weight parameters are randomly initialized using the Xavier initialization method to promote the learning efficiency of the large model in the open classification task of financial text noise. For the batch size and early stopping strategy, set an appropriate batch size and configure the early stopping strategy to prevent overfitting and ensure timely stopping of training when the performance on the validation set no longer improves.

[0091] In this example, fine-tuning the embedding layer and encoding layer of the large model can adapt to the specific requirements of the financial field and enhance the model's understanding of financial proprietary terms. Initialize the optimal validation score to 0, and the optimal model to the initial model.

[0092] In one example, clustering processing is performed on the semantic features to obtain multiple granule balls with different labels, including the following sub-steps:

[0093] S21: Initialize all samples corresponding to the semantic features as one granule ball;

[0094] S22: Calculate the attributes of the granular balls (capacity, purity, centroid, and radius), determine whether to perform granular ball splitting processing according to the granular ball attributes. If so, randomly select samples with each different label from the granular balls as the initial centroids of the new granular balls, calculate the distances from all samples to the initial centroids, and assign the samples to the new granular balls represented by the nearest centroids, thereby obtaining multiple granular balls with different labels.

[0095] Among them, the granular ball is an adaptive clustering unit that can characterize the data distribution at multiple granularity levels. Let the sample set , where is the feature vector of the sample, is the class label of the sample, represents the sample index. By clustering this sample set , a set consisting of a series of granular balls can be obtained, where each granular ball is composed of samples, is the number of granular balls, is the granular ball index.

[0096] Optionally, in step S22, determining whether to perform granular ball splitting processing according to the granular ball attributes includes:

[0097] If the purity of the granular ball is less than the first purity threshold and if the capacity of the granular ball is greater than the first capacity threshold, then perform granular ball splitting.

[0098] Among them, the first purity threshold is a preset parameter, and the user can customize it or set the low purity threshold according to historical experience. Exemplarily, in this example, the first purity threshold is greater than the low purity threshold. The first capacity threshold is a preset parameter, and the user sets it according to the number of training samples in each category in the dataset. In this example, the first capacity threshold is between the low capacity threshold and the high capacity threshold. If the purity of the granular ball is less than the first purity threshold and the capacity of the granular ball is greater than the first capacity threshold, the data points inside the granular ball are not consistent enough, and there may be multiple different categories or features. By splitting this granular ball, the data points can be redistributed into purer sub-granular balls, thereby improving the overall quality of clustering. If the capacity of the granular ball is greater than the first capacity threshold, it means that the granular ball contains too many samples and may contain multiple different sub-patterns or categories. In this case, the sample distribution inside the granular ball may be relatively complex, resulting in a low purity of the granular ball. Therefore, at this time, the granular ball containing multiple sub-patterns can be decomposed into multiple smaller granular balls, and each small granular ball can better represent a specific sub-pattern or category, thereby improving the purity and classification effect of the granular ball. Further, if the purity of the granular ball is greater than or equal to the first purity threshold, or the capacity of the granular ball is less than or equal to the first capacity threshold, then add this granular ball to the final granular ball set .

[0099] Further, when it is determined according to the granule ball attribute that granule ball splitting processing is required, the following sub-steps are included:

[0100] (1) Centroid selection; specifically, calculate the number of different labels included in the granule ball , represents the unique value function; among the samples included in the granule ball , randomly select a representative point for each sample corresponding to a different label as the initial centroid of the new granule ball.

[0101] (2) Distance calculation; specifically, calculate the distance from all samples in the granule ball to the initial centroids of the new granule balls, such as the Euclidean distance.

[0102] (3) Granule ball division. Specifically, for each sample, according to the distance between the sample and each initial centroid, allocate the sample to the new granule ball where the nearest centroid is located, and finally form new granule balls, and continue to judge whether the new granule balls need to continue splitting until all granule balls meet the condition that the purity is greater than or equal to the first purity threshold or the capacity is less than or equal to the first capacity threshold, then stop splitting, and finally obtain granule balls to form a granule ball set , and this granule ball set can represent the multi-granularity features of the data.

[0103] In one example, the method further includes classifying the samples in the granule ball as uncertain samples:

[0104] When classifying the samples in the high-quality granule ball as noise, if the distance from the sample to the centroid is greater than the radius of the granule ball, classify the sample as an uncertain sample; and / or,

[0105] Classify the samples in the granule ball with a purity greater than the low purity threshold and less than the high purity threshold as uncertain samples.

[0106] Among them, the uncertain samples are samples with uncertain noise types. When classifying the samples in the high-quality granule ball as noise, if the distance from the sample to the centroid is greater than the radius of the granule ball, it means that the sample is far from the center of the granule ball and exceeds the radius range of the granule ball, indicating that the noise type of the sample is uncertain. If the purity of the granule ball is greater than the low purity threshold and less than the high purity threshold, it means that there may be multiple types of samples mixed inside the granule ball, and the noise category attribution of the sample is not clear enough. Therefore, if the low purity threshold < granule ball purity < high purity threshold ​If the sample in the granule is classified as an uncertain sample, it is added to the set of uncertain samples. Preferably, in this example, the above two methods for classifying uncertain samples are simultaneously used to classify the samples within the granule.

[0107] Preferably, the above examples are combined. After classifying the samples within the granule into clean samples, in-distribution noise, out-of-distribution noise, and uncertain samples according to the position of the sample in the granule and the attributes of the granule, centroid collection processing is performed. Specifically, the centroids of high-quality granules and low-quality granules are collected as follows:

[0108] ;

[0109] ;

[0110] Among them, is the centroid of all high-quality granules, representing the center of the known class; is the centroid of all low-quality granules, representing the center of the granules of the unknown class.

[0111] In one example, after classifying the samples within the granule into clean samples and in-distribution noise, it further includes:

[0112] S41: Calculate the first cross-entropy loss function of the clean samples :

[0113] ;

[0114] Among them, represents the logarithmic function.

[0115] S42: Modify the in-distribution noise label to the label of the granule where it is located to correct the bias caused by the wrong label, and then calculate the second cross-entropy loss function :

[0116] ;

[0117] Among them, represents the th granule category.

[0118] S43: Use the first cross-entropy loss function and the second cross-entropy loss function to construct the first total loss function, and then optimize the classification model. Among them, the classification model can be a neural network based on deep learning, such as a convolutional neural network, a recurrent neural network, a BERT model, etc.

[0119] Preferably, for uncertain samples, since their true labels cannot be accurately determined, they are directly discarded and not included in the calculation of the loss function.

[0120] In one example, after classifying the samples within the grain balls as out-of-distribution noise, it further includes:

[0121] S41’: Calculate the within-class loss function and the between-class loss function based on the centroids of high-quality grain balls and the labels of high-quality grain balls; calculate the open space loss function based on the centroids of high-quality grain balls and the centroids of low-quality grain balls.

[0122] Specifically, the within-class loss function The calculation expression is:

[0123] ;

[0124] Wherein, represents the indicator function; represents the category of the th grain ball; represents the centroid of the th grain ball; represents the centroid of the

[0125] Further, the between-class loss function The calculation expression is:

[0126] ;

[0127] Wherein, represents a small constant.

[0128] Further, the open space loss function The calculation expression is:

[0129] .

[0130] S42’: Construct the second total loss function using the first cross-entropy loss function, the second cross-entropy loss function, the within-class loss function, the between-class loss function, and the open space loss function:

[0131] ;

[0132] Use the second total loss function to perform backpropagation on the classification model, calculate the gradients, and update the trainable parameters of the model to gradually improve the performance of the model in open intent classification.

[0133] Preferably, after optimizing the model using the loss function, it further includes model update processing, including:

[0134] (1) Validation set evaluation: After each round of training is completed, calculate the evaluation index scores of the model on the validation set to evaluate the current performance of the model.

[0135] (2) Update the optimal verification score and the optimal model: If the current evaluation score of the classification model is higher than the historical optimal score, set this score and the corresponding model parameters as the new optimal result.

[0136] (3) Continue training: Subsequent training iterations are based on the parameters of the current optimal model, continuously improving the adaptability of the model in classifying in a complex noise environment.

[0137] Further, after the model update process, it also includes training stop judgment and boundary preservation. Specifically, if the verification score does not improve for several consecutive times, such as 10 times, the early stopping strategy is triggered, the representation learning is stopped, and the optimal model parameters are saved. At the same time, record the centroids and their radii of each known category represented by the high-quality granular balls as the final category boundaries.

[0138] Combine the above examples to obtain the preferred open classification method of the present invention, as Figure 2 shown, the method includes the following steps at this time:

[0139] S100: Select and construct a pre-trained large language model and initialize the large model;

[0140] S200: Extract the semantic features of financial texts;

[0141] S300: Perform clustering processing on the semantic features to obtain granular balls with multiple different labels, and calculate the attributes of the granular balls with multiple different labels;

[0142] S400: Classify the samples in the granular balls as clean samples, in-distribution noise, out-of-distribution noise, and uncertainty samples according to the position of the samples in the granular balls and the attributes of the granular balls;

[0143] S500: Calculate the loss functions of different types of samples and noise, and then construct a total loss function, and use the total loss function to optimize the model parameters; specifically, calculate the first cross-entropy loss function of clean samples; after modifying the in-distribution noise labels to the labels of the granular balls where they are located, calculate the second cross-entropy loss function; calculate the within-class loss function and between-class loss function according to the centroids and their corresponding labels of the high-quality granular balls, and calculate the open space loss function according to the centroids of the high-quality granular balls and the centroids of the low-quality granular balls; then use the first cross-entropy loss function, the second cross-entropy loss function, the within-class loss function, the between-class loss function, and the open space loss function to construct a second total loss function, and use the second total loss function to optimize the classification model.

[0144] S600: Model update process, if the iteration stop strategy is triggered, stop learning and save the optimal model parameters, and at the same time record the centroids and their radii of each known category represented by the high-quality granular balls as the final category boundaries.

[0145] In one example, after completing model training and validation, the test set data is input into the optimal model. By extracting the semantic features of the financial text to be classified, and then using the centroids and radii of each known category saved, open intent classification is performed, that is, the test or application of the open intent classification method. At this time, the method includes the following steps:

[0146] S100’: Extract the semantic features of the financial text to be classified;

[0147] S200’: For each sample to be classified corresponding to the semantic features, calculate the Euclidean distance between the sample to be classified and the centroids of all known categories, then find the centroid of the first known category closest to each sample to be classified, and compare this distance with the centroid of the first known category; At this time, the known categories use the centroids and radii of each known category represented by high-quality granular balls as the final category boundaries;

[0148] S300’: If the distance from the sample to be classified to the centroid of the first known category is less than or equal to the radius corresponding to the centroid of the first known category, classify the sample to be classified as the known category corresponding to the centroid of the first known category; If the distance from the sample to be classified to the centroid of the first known category is greater than the radius corresponding to the centroid of the first known category, classify the sample to be classified as an unknown category.

[0149] Through the above method, the present invention can make an efficient and accurate distinction between known categories and unknown categories during the inference stage, greatly improving the practicability and robustness of open intent classification.

[0150] The above specific implementation manners are detailed descriptions of the present invention. It cannot be determined that the specific implementation manners of the present invention are only limited to these descriptions. For those of ordinary skill in the technical field to which the present invention belongs, without departing from the concept of the present invention, several simple deductions and substitutions can still be made, and all should be regarded as belonging to the protection scope of the present invention.

Claims

1. An open classification method for multi-granularity financial text noise based on large models, characterized in that, The method includes a training step: Extract the semantic features of financial texts; Cluster the semantic features to obtain multiple granules with different labels, and calculate the attributes of the granules with different labels, including capacity, purity, centroid, and radius; According to the position of the samples in the granules and the attributes of the granules, classify the samples in the granules into clean samples, in-distribution noise, and out-of-distribution noise, including the following sub-steps: Define the granules with granule purity greater than the high purity threshold and granule capacity greater than the high capacity threshold as high-quality granules. The noise classification of the samples in the high-quality granules includes: if the distance from the sample to the centroid is less than the radius of the granule and the sample and the granule have the same label, classify the sample as a clean sample; If the distance from the sample to the centroid is less than the radius of the granule and the sample and the granule have different labels, classify the sample as in-distribution noise; Define the granules with granule purity less than the low purity threshold and granule capacity less than the low capacity threshold as low-quality granules, and classify the samples in the low-quality granules as out-of-distribution noise.

2. The multi-granularity financial text noise open classification method based on a large model according to claim 1, characterized in that The extraction of the semantic features of financial texts includes: Freeze the underlying parameters of the large model, and use the training dataset of financial texts to train the top-layer parameters of the large model. The top layer of the large model includes an embedding layer and an encoding layer; Use the trained large model to extract the semantic features of financial texts.

3. The multi-granularity financial text noise open classification method based on a large model according to claim 1, wherein The clustering process of the semantic features to obtain multiple granules with different labels includes the following sub-steps: Initialize the samples corresponding to all semantic features as a single granule; Calculate the attributes of the granule. According to the granule attributes, judge whether to perform granule splitting. If so, randomly select samples with each different label from the granule as the initial centroids of the new granules, calculate the distances from all samples to the initial centroids, and assign the samples to the new granules represented by the nearest centroid, thereby obtaining multiple granules with different labels.

4. The method for open classification of multi-granularity financial text noise based on a large model according to claim 3, wherein The judgment of whether to perform granule splitting according to the granule attributes includes: If the granule purity is less than the first purity threshold and the granule capacity is greater than the first capacity threshold, then perform granule splitting.

5. The multi-granularity financial text noise open classification method based on a large model according to claim 1, wherein, The method further includes classifying the samples in the granules as uncertain samples: When performing noise classification on the samples in the high-quality granules, if the distance from the sample to the centroid is greater than the radius of the granule, classify the sample as an uncertain sample; and / or, Classify the samples in the granules with granule purity greater than the low purity threshold and less than the high purity threshold as uncertain samples; Discard the uncertain samples and do not include them in the calculation of the loss function.

6. The open classification method for multi-granularity financial text noise based on a large model according to claim 1, characterized in that After classifying the samples in the granules as clean samples and in-distribution noise, it further includes: Calculate the first cross-entropy loss function of the clean samples; After modifying the in-distribution noise labels to the labels of the granules where they are located, calculate the second cross-entropy loss function; Use the first cross-entropy loss function and the second cross-entropy loss function to construct the first total loss function, and then optimize the classification model.

7. The multi-granularity financial text noise open classification method based on a large model according to claim 6, wherein After classifying the samples in the granules as out-of-distribution noise, it further includes: According to the centroid of the high-quality granules and the labels of the high-quality granules, calculate the intra-class loss function and the inter-class loss function; according to the centroid of the high-quality granules and the centroid of the low-quality granules, calculate the open space loss function; Construct a second total loss function using the first cross-entropy loss function, the second cross-entropy loss function, the intra-class loss function, the inter-class loss function, and the open space loss function, and then optimize the classification model.

8. The multi-granularity financial text noise open classification method based on a large model according to claim 1, wherein The method further includes an application step: Extract the semantic features of the financial text to be classified; For each sample to be classified corresponding to the semantic features, determine the first known class centroid with the closest distance; The known classes use the centroids and radii of each known class represented by high-quality granular balls as the final class boundaries; If the distance from the sample to be classified to the first known class centroid is less than or equal to the radius corresponding to the first known class centroid, classify the sample to be classified as the known class corresponding to the first known class centroid; If the distance from the sample to be classified to the first known class centroid is greater than the radius corresponding to the first known class centroid, classify the sample to be classified as an unknown class.