A large model-based noise text open intent classification method and system
Through granular structure modeling and joint loss function, the difficulties of traditional intent classification systems in identifying unknown intents and processing noise in open environments are solved, and the model's high robustness and generalization ability in complex environments are achieved.
Patent Information
- Application Number
- CN202511071956.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-01
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-08-01
AI Technical Summary
Traditional intent classification systems have difficulty identifying known and unknown intents in open environments. Existing methods are not very accurate in dealing with label noise and have difficulty distinguishing between in-distribution noise and out-of-distribution noise, resulting in insufficient model generalization capabilities.
The granular structure modeling is adopted, combined with the global granular-level OOD tendency index and the sample-level consistency index. Through unsupervised granular clustering and joint loss function, the noise samples in the training data are identified and processed, and a multi-granularity decision boundary is constructed to improve the model robustness.
It effectively improves the recognition robustness and generalization ability of the model in complex open environments, can adaptively capture multi-prototype structures, and enhance the understanding of known classes and the generalization ability of unknown classes.
Smart Images

Figure CN120578766B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of natural language processing and artificial intelligence technology, and in particular to a method and system for classifying open intent in noisy text based on a large model, which is suitable for intent understanding tasks in text classification, dialogue systems, and other open-world scenarios. Background Art
[0002] With the rapid advancement of technology, conversational systems have become a key technology for providing customer service and intelligent consulting. For example, in scenarios such as intelligent customer service, financial Q&A, and human-computer interaction, it is essential to accurately identify and classify user query intent. Traditional intent classification systems typically only recognize intent categories defined during training and lack the ability to handle newly emerging or unknown intent categories. This limits the system's adaptability and utility in open environments. As the task of open intent classification grows in importance, how to simultaneously identify known intents and detect unknown intents in real-world scenarios has become a research hotspot. Traditional methods often rely on clean training data and assume that training and test categories are identical, making them difficult to address the common label noise problem in the real world. Furthermore, existing methods often discard noisy samples, particularly out-of-distribution (OOD) noise, as interference during training, failing to fully tap its potential value.
[0003] To address these issues, some studies have attempted to identify noise through clustering or to enhance model robustness using multi-prototype structures to assist in identifying label-noisy samples. However, these methods typically require pre-specifying the number of clusters for each category, and most rely on supervised clustering with noisy labels. This makes them susceptible to label errors, resulting in poor clustering quality and limiting the method's generalization capabilities. Furthermore, existing classification methods rely on methods such as sample loss or distance from the class center to identify noise types, which are not sufficiently accurate and make it difficult to distinguish between noise from known class labels (in-distribution noise samples, IND) mixed in the training data and noise introduced by pseudo-labeling of unknown class samples (out-of-distribution noise, OOD).
[0004] Therefore, there is an urgent need for a robust method that can adaptively characterize the structure of training data, effectively identify various types of noise, and utilize noise to promote open intent classification. Summary of the Invention
[0005] The purpose of the present invention is to overcome the problems existing in the prior art and provide a noisy text open intent classification method and system based on a large model. It takes the granular structure as the core modeling unit, integrates multi-granularity structural information, noise pattern recognition mechanism and joint representation learning strategy, and can simultaneously complete the classification of known intents and the rejection of unknown intents in the presence of label noise in the training data, effectively improving the recognition robustness and generalization ability of the model in complex open environments.
[0006] The object of the present invention is achieved through the following technical solutions:
[0007] In a first aspect, a method for classifying open intent in noisy text based on a large model is provided, comprising the following steps:
[0008] S1. Granular Spherical Structure Modeling: We use a pre-trained large language model to extract text features, cluster these features through unsupervised granular sphere clustering, and construct intra-class structural representations in the feature space.
[0009] S2. Multi-level noise identification: Based on the global granular-level OOD tendency index and the sample-level consistency index, the training samples are divided into clean samples, in-distribution noise samples, out-of-distribution noise samples, and uncertain samples;
[0010] S3. Model Training: Construct a beneficial noise-guided joint loss function and use the sample partitioning results to perform differential training on the model. The joint loss function introduces the OOD dispersion loss, prototype softmax loss, and prototype margin loss in a weighted manner.
[0011] S4: Intention Reasoning: Construct a multi-granularity decision boundary based on a granular structure to determine the category of the input sample.
[0012] In some embodiments, extracting text features using a pre-trained large language model includes:
[0013] Freeze the bottom encoding layer of the pre-trained large language model and fine-tune the top layer structure;
[0014] Extract text features using the adjusted pre-trained large language model.
[0015] In some embodiments, the global sphere-level OOD tendency index integrates the relative positions between spheres and the neighborhood label distribution information to determine the extent to which a sphere contains noise samples outside the distribution; the sample-level consistency index is used to evaluate the rationality of the distribution of each sample within its sphere and the reliability of its label. This index combines the local neighborhood density and label consistency information to determine the extent to which a sample is a clean sample or a noise sample within the distribution.
[0016] In some embodiments, the global spherical-level OOD tendency index is calculated as follows:
[0017] , j represents the ball number, Indicates the OOD tendency index value of the global particle level of the target particle, It represents the average distance between the centers of mass of the nearest spheres with the same pseudo-label as the target sphere. Indicates the proportion of inconsistent pseudo labels among the several spheres closest to the target sphere.
[0018] In some embodiments, the sample-level consistency index is calculated as follows:
[0019] , where i represents the sample number, represents a sample in the sphere, Indicates the consistency index value of a sample within the sphere, Indicates the tag purity of the pellet, represents the center of mass of the sphere, represents the maximum radius of the sphere, and exp() represents the exponential function with e as the base.
[0020] In some embodiments, the method of dividing training samples into clean samples, in-distribution noise samples, out-of-distribution noise samples, and uncertain samples based on the global granular-level OOD tendency index and the sample-level consistency index includes:
[0021] Based on the combined results of the global granular-level OOD tendency index and the sample-level consistency index, the following four categories of discrimination criteria are set:
[0022] like and , it is judged as a clean sample;
[0023] like and , then it is determined to be a noise sample within the distribution, represents the true label of the sample, Indicates the pellet label;
[0024] like , it is determined to be a noise sample outside the distribution;
[0025] In other cases, the samples are determined to be uncertain;
[0026] in, and is the empirical threshold, obtained based on training set statistics.
[0027] In some embodiments, the joint loss function is defined as:
[0028] ,in, represents the joint loss function, represents the OOD dispersion loss, represents the prototype Softmax loss, represents the prototype interval loss, 、 、 is the weight coefficient of the loss term.
[0029] Secondly, a large-model-based noisy text open intent classification system is provided, including:
[0030] The granular structure modeling module is configured to extract text features using a pre-trained large language model, cluster the text features through unsupervised granular clustering, and construct intra-class structure representation in the feature space;
[0031] The multi-level noise recognition module is configured to classify training samples into clean samples, in-distribution noise samples, out-of-distribution noise samples, and uncertain samples based on the global granular-level OOD tendency index and the sample-level consistency index.
[0032] A model training module is configured to construct a beneficial noise-guided joint loss function and perform differential training on the model using the sample partitioning results. The joint loss function introduces OOD dispersion loss, prototype softmax loss, and prototype margin loss in a weighted manner.
[0033] The intention reasoning module is configured to construct a multi-granularity decision boundary based on a granular sphere structure to determine the category of the input sample.
[0034] In a third aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program. When the computer program is executed by a processor, the method for classifying open intent of noisy text based on a large model described in the first aspect is implemented.
[0035] In a fourth aspect, an electronic device is provided, comprising a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, and when the processor executes the computer instructions, the large model-based noisy text open intention classification method described in the first aspect is executed.
[0036] It should be further explained that the technical features corresponding to the above embodiments can be combined or replaced with each other to form a new technical solution if there is no conflict.
[0037] Compared with the prior art, the present invention has the following beneficial effects:
[0038] 1. This invention uses spherical structures as the core modeling unit, integrating multi-granularity structural information, noise pattern recognition mechanisms, and joint representation learning strategies. It can simultaneously classify known intents and reject unknown intents in the presence of label noise in the training data, effectively improving the model's recognition robustness and generalization capabilities in complex open environments.
[0039] 2. This paper employs advanced large-model fine-tuning techniques. By freezing the underlying encoding layer and fine-tuning the top-level structure, the model is better adapted to the target task's data distribution, thereby extracting discriminative semantic features, enhancing understanding of known classes and generalization to unknown classes. Furthermore, based on an unsupervised granular sphere clustering algorithm, a representation of intra-class structure is constructed in the feature space, adaptively capturing the multi-prototype structure and local density variations in the sample distribution, thereby enhancing the ability to model intra-class distribution heterogeneity.
[0040] 3. The present invention provides a multi-level noise identification method based on granular structure. Starting from the granular structure, the distribution consistency of samples in local and global structures is comprehensively considered, and two structured evaluation indicators are designed, namely the OOD tendency indicator at the global granular level and the consistency indicator at the sample level, which are used to jointly judge whether the sample is a clean sample, IND noise, OOD noise or uncertain sample, thereby improving the accuracy of sample noise type identification.
[0041] 4. This paper proposes a representation learning method guided by beneficial noise and designs three jointly optimized customized loss functions to guide feature space structure learning from multiple aspects, thereby improving the model's generalization ability and unknown class discrimination ability in open environments. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] Figure 1 This is a flow chart of a method for classifying open intent in noisy text based on a large model according to an embodiment of the present invention;
[0043] Figure 2 A schematic diagram of a noisy text open intent classification system based on a large model is shown in an embodiment of the present invention. DETAILED DESCRIPTION
[0044] The technical solutions of the present invention are described clearly and completely below in conjunction with the accompanying drawings. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. The components of the embodiments of the present application generally described and shown in the drawings herein can be arranged and designed in various different configurations. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of the present invention.
[0045] It should be noted that the defects existing in the solutions in the above-mentioned prior art are the results obtained by the inventor after practice and careful research. Therefore, the discovery process of the above-mentioned problems and the solutions proposed in the embodiments of this application below for the above-mentioned problems should be the contributions made by the inventor to this application in the process of invention and creation, and should not be understood as technical contents known to technical personnel in this field.
[0046] In response to the technical problems pointed out in the background technology, the embodiments provided by the present invention are as follows:
[0047] Reference Figure 1 In an exemplary embodiment, a method for classifying open intent of noisy text based on a large model is provided, comprising the following steps:
[0048] S1. Granular Spherical Structure Modeling: We use a pre-trained large language model to extract text features, cluster these features through unsupervised granular sphere clustering, and construct intra-class structural representations in the feature space.
[0049] S2. Multi-level noise identification: Based on the global granular-level OOD tendency index and the sample-level consistency index, the training samples are divided into clean samples, in-distribution noise samples, out-of-distribution noise samples, and uncertain samples;
[0050] S3. Model Training: Construct a beneficial noise-guided joint loss function and use the sample partitioning results to perform differential training on the model. The joint loss function introduces the OOD dispersion loss, prototype softmax loss, and prototype margin loss in a weighted manner.
[0051] S4: Intention Reasoning: Construct a multi-granularity decision boundary based on a granular structure to determine the category of the input sample.
[0052] In step S1, advanced large-model fine-tuning techniques, such as BERT or RoBERTa, are used to tailor the upper-layer parameters of the model to the complex semantics and specific terminology of the open intent classification task. By freezing the underlying encoding layers and fine-tuning the top-level structure, the model is better adapted to the target task's data distribution, extracting discriminative semantic features and enhancing understanding of known classes and generalization to unknown classes. An unsupervised granular ball clustering algorithm is used to construct a representation of intra-class structure in the feature space. Granular ball computing, a clustering strategy based on regional density, provides a new approach for constructing structure-aware representation spaces. Through unsupervised partitioning and fusion of the feature space, data distribution characteristics can be more accurately characterized, laying the foundation for open intent recognition in complex noisy environments. Granular balls are microscopic clustering units that simultaneously express distribution structure and semantic consistency. They can adaptively capture the multi-prototype structure and local density variations in the sample distribution, thereby enhancing the ability to model intra-class distribution heterogeneity.
[0053] In step S2, to identify the complex noise types in the training set, the present invention designs two structured indicators:
[0054] The global sphere-level OOD tendency indicator (G-ODT) is used to assess whether a sphere deviates from the main structure of a cluster of similar spheres. This indicator combines the relative position of spheres and the distribution of neighborhood labels to effectively determine the extent to which a sphere may contain OOD noise samples.
[0055] The sample-level consistency indicator (S-ICD) is used to evaluate the rationality of the distribution of each sample within its sphere and the reliability of its label. This indicator combines local neighborhood density and label consistency information to accurately determine the degree to which a sample is a clean sample or an IND noise sample.
[0056] Based on the above two indicators, the present invention divides training samples into four categories: clean samples, IND noise samples (in-distribution noise samples), OOD noise samples (out-of-distribution noise samples) and uncertain samples, providing reliable labels and structural information support for subsequent representation learning.
[0057] In step S3, a beneficial noise-guided representation learning mechanism is constructed. Based on multi-level noise recognition, the present invention further proposes a jointly optimized representation learning framework, wherein:
[0058] Introducing OOD Dispersion Loss to penalize OOD noise samples that are close to any known class prototype, pushing them away from known class areas to assist in open space modeling;
[0059] Introducing the Prototype Softmax Loss to enhance the intra-class aggregation of clean samples and label-corrected IND noise samples, and constructing a compact class representation;
[0060] Prototype Margin Loss is introduced to expand the interval between prototypes of different categories and enhance the model's ability to discriminate between known categories.
[0061] The three loss functions are jointly optimized in a weighted manner to jointly construct a structure-aware and robust discriminative feature space.
[0062] In step S4, during the inference phase, the present invention constructs a multi-granularity discrimination boundary based on the granular sphere representation, and uses the granular sphere center of mass and radius to define the spatial boundary of each class, achieving flexible intra-class attraction and out-of-class rejection capabilities, thereby improving the model's recognition performance for unknown intents.
[0063] In summary, the present invention combines the granular structure with the multi-level structure indicators to achieve comprehensive modeling of the open intent classification task in a complex noise environment. It has the advantages of strong recognition robustness, good open generalization ability, and strong interpretability. It is suitable for practical scenarios with high requirements for intent recognition, such as intelligent customer service, financial Q&A, and human-computer interaction.
[0064] In another exemplary embodiment, based on the inventive concept of the above embodiment, the implementation details of each step are described in detail by taking financial text as an example.
[0065] 1. Spatial approximation method based on sphere clustering
[0066] An unsupervised granular sphere clustering method is proposed to automatically mine latent class structures from noisy data to achieve robust structure modeling and subsequent noise identification.
[0067] Current mainstream research typically uses clustering algorithms to estimate prototypes for each category to aid in identifying label-noisy samples. However, these methods typically require pre-specifying the number of clusters for each category and rely on supervised clustering with noisy labels, making them susceptible to label errors and resulting in poor clustering quality. To overcome these issues, the present invention employs a fully unsupervised granular sphere clustering method that adaptively models the structure of the latent representation space without relying on any label information. This method more accurately perceives the local density, structural continuity, and coverage of samples in space.
[0068] 1.1 Selection and Construction of Pre-trained Large Language Model
[0069] Financial Text Collection: Collect user inquiry text data from financial dialogue systems, ensuring that the data covers a wide range of financial topics, such as account inquiries and trading operations. Financial experts are invited to carefully annotate the collected text data to clearly define the intent category of each inquiry text. Annotated intent categories should cover account management, transaction processing, product inquiries, and complaint handling. The collected text data is then thoroughly cleaned to remove irrelevant information, such as HTML tags and special characters. Spelling errors are corrected, and terminology and expressions are standardized to improve data quality.
[0070] Base model selection: Select a currently popular pre-trained large language model as the foundation for domain-specific intent classification tasks. Recommended models include BERT, RoBERTa, ALBERT, and ELECTRA. These models are widely adopted for their powerful semantic modeling capabilities and ability to efficiently capture complex language structures. They can provide stable and generalizable feature representations for subsequent tasks.
[0071] Model Architecture: The selected model typically consists of an input layer, a preprocessing module, an embedding layer, an encoding layer, and an output layer. We fine-tune the embedding and encoding layers to address the linguistic characteristics of specific domains, improving the model's understanding of domain terminology and semantic structure. Furthermore, during training, the optimal validation score is initialized to 0; this initial model serves as the baseline for comparison.
[0072] Data preprocessing: includes text preprocessing sub-steps and word segmentation and encoding processing sub-steps; among them, the text preprocessing sub-step includes: preprocessing the input financial statements, including removing stop words and punctuation, and performing text standardization, etc., in preparation for the next step of feature extraction.
[0073] Compute features for financial text. Pre-trained models (such as BERT) compute word embeddings and generate contextual representations for each text unit. A multi-layer self-attention mechanism is preferably used for feature extraction, followed by transformation and normalization using a feedforward neural network.
[0074] 1.2 Definition and properties of spheres
[0075] Granular-Ball is an adaptive clustering unit that can describe the true distribution of data. ,in is the characteristic vector of the sample, is the category label of the sample. By clustering the sample set, a series of spheres can be obtained, which are represented as a set , where each ball Depend on samples and has the following properties:
[0076] Ball size : the number of samples contained in the sphere;
[0077] Center of mass of sphere : the mean of all sample features within the sphere;
[0078] Mean radius of spheres : the average Euclidean distance from all samples in the sphere to the centroid;
[0079] Maximum radius of the sphere : the maximum Euclidean distance from all samples in the sphere to the centroid;
[0080] Ball Label : The label of the category with the highest proportion in the sphere;
[0081] Pellet purity :The ball belongs to the label The sample proportion.
[0082] 1.3 Unsupervised granular sphere clustering
[0083] In the specific implementation, first a sample set is given , we get the sample feature representation extracted by the pre-trained large language model and get the feature set ,in is the representation vector of the sample, Then, in this feature space, all Perform sphere clustering operations to construct sphere sets to fit the true distribution structure of the data.
[0084] The sphere clustering process consists of two stages: sphere generation and sphere merging. The specific operations are as follows:
[0085] 1.3.1 Ball generation stage:
[0086] Initially, the entire feature set is considered as a whole sphere , and the whole is represented as multiple spheres by recursive splitting. For each sphere, its distribution metric is defined ( ):
[0087]
[0088] Only when the current ball is divided can The partitioning operation is performed only when the value decreases. The partitioning strategy is as follows:
[0089] First find the center of mass of the sphere The farthest sample ; Find the distance among the remaining samples The farthest sample ;by and 、 The midpoint of the original ball is used as the initial center of mass of the two sub-balls; the samples in the original ball are divided into the ball closest to the new initial center of mass according to the Euclidean distance, and divided into two sub-balls. , , and calculate the weighted distribution measure:
[0090]
[0091] If this value is less than the original ball If the value is less than the threshold, the division result is retained; otherwise, the division is terminated. To avoid over-division, if the number of samples in the sphere is less than the threshold , then no further division is done.
[0092] 1.3.2 Ball merging stage:
[0093] After the division is completed, if the two balls and The sum of the distance from the centroid minus its radius is less than the set threshold , that is, satisfying:
[0094]
[0095] It is considered that the two structures overlap and a merge operation should be performed, which is to merge the two spheres into one sphere, and the sphere properties are calculated according to the definition in 1.1.
[0096] Finally, we get a collection of spheres ,This set has a compact structure and complete distribution, and can serve as the basic structure for subsequent noise recognition and learning.
[0097] 2. Multi-level noise sample recognition method
[0098] This paper proposes a multi-level noise recognition module based on a granular sphere structure to improve the accuracy of sample noise type identification. Unlike existing methods that rely on sample loss or distance from the class center, this method takes the granular sphere structure into consideration and comprehensively considers the consistency of sample distribution in both local and global structures. It then designs two structured evaluation metrics: the global granular sphere-level Out-of-Detection (OOD) tendency metric (G-ODT) and the sample-level consistency metric (S-ICD). These metrics are used to jointly determine whether a sample is clean, contains IND noise, OOD noise, or is uncertain.
[0099] 2.1 Global granular-level OOD tendency indicator (G-ODT)
[0100] This metric measures the degree to which a sphere deviates from the known distribution of its labels in feature space. Based on our observations, if a sphere is primarily composed of OOD samples, its centroid will be far from the centroid of similar spheres, and there will be strong label inconsistency in its local neighborhood. Therefore, the G-ODT metric combines two factors:
[0101] Centroid distance: Calculate the average distance between the centroids of the nearest M spheres among the spheres with the same pseudo-label as the target sphere. :
[0102]
[0103] in Representation and Labeling Same distance The center of mass of the M nearest spheres.
[0104] Label consistency: Evaluating pellets There is label inconsistency between and its local neighborhood. Specifically, we select The nearest T balls, and calculate the pseudo labels and Inconsistent proportions :
[0105]
[0106] in Is an indicator function that returns 1 if the condition is true.
[0107] Finally, G-ODT is defined as the product of these two values:
[0108]
[0109] The larger the G-ODT value, the more likely it is that the samples inside the sphere are dominated by OOD noise samples.
[0110] 2.2 Sample-level consistency index (S-ICD)
[0111] This indicator is used to measure the representativeness of the sample within the pellet and is defined by combining the purity of the pellet and the relative position of the sample.
[0112] Set ball The tag purity is , the center of mass is , the maximum radius is , then a sample in the sphere The S-ICD value is calculated as follows:
[0113]
[0114] The larger the value, the closer the sample is to the centroid and the higher the label consistency of the sphere it is in, and the greater the possibility that it is a clean or IND noise sample. exp() represents an exponential function with base e.
[0115] 2.3 Sample Classification Rules and Category Prototype Selection
[0116] According to the combined results of G-ODT and S-ICD and two thresholds and The following four categories of discrimination criteria are proposed:
[0117] like and , it is judged as a clean sample;
[0118] like and , then it is determined to be a noise sample within the distribution, represents the true label of the sample, Indicates the pellet label;
[0119] like , it is determined to be a noise sample outside the distribution;
[0120] In other cases, the samples are determined to be uncertain;
[0121] in, and is the empirical threshold, obtained based on training set statistics.
[0122] This method classifies all training samples into four categories based on the aforementioned rules: clean samples, IND noise, OOD noise, and uncertain samples. After classification, IND noise samples are reassigned with pseudo-labels corresponding to their spheres for subsequent training. Uncertain samples are excluded from model training due to their identification uncertainty.
[0123] In addition, the present invention further selects the tag purity greater than the threshold And the number of samples included exceeds The centroid of the sphere is used as the class prototype, which is convenient for representing the structure of each category in subsequent representation learning and improving the ability of category discrimination.
[0124] 3. Beneficial Noise-Guided Representation Learning Method
[0125] This paper proposes a representation learning method guided by beneficial noise, aiming to improve the model's generalization and unknown class discrimination capabilities in open environments. Unlike existing methods that consider out-of-distribution (OOD) noise harmful and directly eliminate it, this method considers OOD noise as an approximate representation of unknown classes, assisting in open space modeling. Based on the previously identified sample categories (clean samples, IND noise, OOD noise) and the multi-prototype structure of each category, three jointly optimized customized loss functions are designed to guide feature space structure learning from multiple perspectives.
[0126] Assume that the OOD noise sample set is , the clean and IND sample sets are , the prototype set of known categories is , where each prototype Corresponding category label Based on this structure, this method proposes the following three optimization objectives:
[0127] 3.1 OOD Dispersion Loss
[0128] This loss aims to exclude OOD noise samples from all known category prototype areas and guide them away from known category prototypes. , select the nearest Known class prototypes , and define its dispersion loss as:
[0129]
[0130] in, Controls the penalty strength. The total loss for all OOD samples is:
[0131] .
[0132] 3.2 Prototype Softmax Loss
[0133] This loss is used to guide clean samples and corrected IND noise samples to gather in the prototype area of their corresponding categories, forming a compact category distribution. and its label , select the closest prototype with the same label , define its prototype Softmax loss as:
[0134] ,
[0135] The total loss is:
[0136] .
[0137] 3.3 Prototype Margin Loss
[0138] This loss is used to improve the clarity of the boundary between classes and enhance the rejection ability of the model. , with its closest prototype For reference, the loss is:
[0139] ,
[0140] in, is the class interval hyperparameter. The total loss is:
[0141] .
[0142] 3.4 Overall optimization goal
[0143] The three loss functions are jointly optimized in a linear combination, and the total loss is defined as:
[0144] ,
[0145] in, 、 、 is the weight coefficient of the loss term.
[0146] Through a triple optimization mechanism guided by beneficial noise, the present invention effectively achieves modeling of unknown space, aggregation of known classes, and enhancement of inter-class boundaries in the representation learning process, thereby significantly improving the recognition robustness and generalization ability of the model in open intent classification tasks.
[0147] 3.5 Gradient backpropagation and parameter modification
[0148] The above loss function is used for backpropagation to calculate the gradient and update the trainable parameters of the model to gradually improve the performance of the model in open intent classification.
[0149] 3.6 Model Update
[0150] 3.6.1 Validation Set Evaluation
[0151] After completing each round of training, calculate the evaluation index score of the model on the validation set to evaluate the current performance of the model
[0152] 3.6.2 Update the best validation score and the best model
[0153] If the current evaluation score is higher than the historical optimal score, the score and the corresponding model parameters are set as the new optimal result.
[0154] If the current evaluation score is lower than or equal to the historical best score, the optimal model will not be updated.
[0155] 3.6.3 Continue training
[0156] Subsequent training iterations are performed based on the parameters of the current optimal model, which continuously improves the adaptability of the model in complex noisy environments.
[0157] 3.7 Training Stop and Boundary Preservation
[0158] If the validation score does not improve after 10 consecutive times, the early stopping strategy is triggered, the representation learning is stopped, and the optimal model parameters are saved. At the same time, the centroid and radius of each known category represented by the high-quality particle sphere are recorded as the final category boundary.
[0159] 4. Inference Method Based on Granular Sphere Structure: This paper utilizes an open-intent classification inference method based on a granular sphere structure. This method relies on the granular sphere set constructed during the representation learning phase and establishes multi-granular, locally adaptive decision boundaries for each category during the inference process, thereby achieving accurate recognition of known class samples and effective rejection of unknown class samples.
[0160] Specifically, for each known category, the spheres obtained in the training phase are screened to find the spheres whose purity is greater than a preset threshold. And the number of samples is greater than the threshold The spheres are used as effective structural units for this category. Each sphere defines its coverage by its center of mass and average radius, forming an approximate description of the embedding space distribution of this category.
[0161] During the inference process, for any query sample to be classified, the following discriminant steps are performed in sequence:
[0162] 1. Calculate the Euclidean distance between the sample and the centroid of all known categories of particles;
[0163] 2. Determine whether the sample falls within the coverage radius of any ball:
[0164] If it falls within the radius of at least one sphere, it will be assigned to the category corresponding to the centroid closest to it in the sphere covering it;
[0165] If it does not fall within the coverage of any sphere, it is considered an unknown class sample.
[0166] In another exemplary embodiment, based on the same inventive concept as the method embodiment, a noise text open intention classification system based on a large model is provided, such as Figure 2 Shown, including:
[0167] The granular structure modeling module is configured to extract text features using a pre-trained large language model, cluster the text features through unsupervised granular clustering, and construct intra-class structure representation in the feature space;
[0168] The multi-level noise recognition module is configured to classify training samples into clean samples, in-distribution noise samples, out-of-distribution noise samples, and uncertain samples based on the global granular-level OOD tendency index and the sample-level consistency index.
[0169] A model training module is configured to construct a beneficial noise-guided joint loss function and perform differential training on the model using the sample partitioning results. The joint loss function introduces OOD dispersion loss, prototype softmax loss, and prototype margin loss in a weighted manner.
[0170] The intention reasoning module is configured to construct a multi-granularity decision boundary based on a granular sphere structure to determine the category of the input sample.
[0171] First, in the spherical structure modeling module, after fine-tuning the large model for feature extraction, an unsupervised spherical clustering mechanism is used to cluster the feature representations of the input samples, forming adaptive and robust local spherical structures. This clustering process effectively approximates the original data distribution, thereby alleviating the impact of label noise and providing structured priors for subsequent noise recognition and representation learning.
[0172] Secondly, a multi-level noise recognition strategy was constructed in the multi-level noise recognition module, comprehensively considering the structural consistency at the sphere level and the sample level. Specifically, it includes: (1) a global sphere-level OOD tendency index, which is used to assess whether the sphere as a whole has an abnormal distribution or a tendency to deviate from the center of the main category; (2) a sample-level consistency index, which is used to measure the representativeness and consistency of the sample within the sphere to which it belongs. Based on the combined judgment of these two indicators, the training samples are automatically divided into four categories: clean samples, in-distribution noise samples (IND noise), out-of-distribution noise samples (OOD noise), and uncertain samples.
[0173] Next, in the model training module, a representation learning mechanism based on beneficial noise guidance is proposed. The above sample partitioning results are used for differential training to improve the model's discriminability and generalization capabilities. This module designs three jointly optimized loss functions to: (1) push OOD noise samples away from known class prototypes to construct open space boundaries; (2) guide clean samples and corrected IND noise samples to cluster towards their corresponding class prototypes to enhance intra-class compactness; and (3) introduce distance intervals between different class prototypes to expand the inter-class discrimination boundary.
[0174] Finally, the intention reasoning module uses the trained model for reasoning and constructs a multi-granularity decision boundary based on the granular structure to determine the category of the input sample.
[0175] In another exemplary embodiment, based on the same inventive concept as the method embodiment, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, a method for classifying open intent of noisy text based on a large model provided by an embodiment of the present invention is implemented. Based on this understanding, the technical solution of this embodiment, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0176] In another exemplary embodiment, based on the same inventive concept as the method embodiment, an electronic device is provided, including a memory and a processor, wherein the memory stores computer instructions that can be executed on the processor, and when the processor runs the computer instructions, it executes a large model-based noisy text open intention classification method provided in an embodiment of the present invention.
[0177] The processor may be a single-core or multi-core central processing unit or a specific integrated circuit, or one or more integrated circuits configured to implement the present invention.
[0178] Embodiments of the subject matter and functional operations described in this specification may be implemented in: tangibly embodied computer software or firmware, computer hardware including the structures disclosed in this specification and their structural equivalents, or a combination of one or more thereof. Embodiments of the subject matter described in this specification may be implemented as one or more computer programs, i.e., one or more modules of computer program instructions encoded on a tangible, non-transitory program carrier for execution by a data processing apparatus or to control the operation of the data processing apparatus. Alternatively or in addition, the program instructions may be encoded on an artificially generated propagated signal, such as a machine-generated electrical, optical, or electromagnetic signal, that is generated to encode and transmit information to a suitable receiver apparatus for execution by the data processing apparatus.
[0179] The processes and logic flows described in this specification can be performed by one or more programmable computers executing one or more computer programs to perform the corresponding functions by operating on input data and generating output. The processes and logic flows can also be performed by, and apparatus can be implemented as, special purpose logic circuitry, such as an FPGA (field programmable gate array) or an ASIC (application-specific integrated circuit).
[0180] Processors suitable for executing computer programs include, for example, general-purpose and / or special-purpose microprocessors, or any other type of central processing unit. Typically, a central processing unit will receive instructions and data from a read-only memory and / or random access memory. The basic components of a computer include a central processing unit for implementing or executing instructions and one or more memory devices for storing instructions and data. Typically, a computer will also include one or more mass storage devices for storing data, such as magnetic disks, magneto-optical disks, or optical disks, or the computer will be operably coupled to such a mass storage device to receive data from it or to transmit data to it, or both. However, a computer does not necessarily have such a device. In addition, a computer can be embedded in another device, such as a mobile phone, a personal digital assistant (PDA), a mobile audio or video player, a game console, a global positioning system (GPS) receiver, or a portable storage device such as a universal serial bus (USB) flash drive, to name a few.
[0181] It should be understood that each box in the flow chart or block diagram can represent a module, program segment or a part of code, and the part of the module, program segment or code comprises one or more executable instructions for realizing the logical function of the provision. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs the function or action of the provision, or can be implemented with a combination of dedicated hardware and computer instructions.
[0182] The above specific implementation methods are detailed descriptions of the present invention. It cannot be considered that the specific implementation methods of the present invention are limited to these descriptions. For ordinary technicians in the technical field to which the present invention belongs, they can make several simple deductions and substitutions without departing from the concept of the present invention, which should be regarded as falling within the scope of protection of the present invention.
Claims
1. A method for classifying open intent of noisy text based on a large model, characterized by: The following steps are involved: S1. Granular Spherical Structure Modeling: We use a pre-trained large language model to extract text features, cluster these features through unsupervised granular sphere clustering, and construct intra-class structural representations in the feature space. S2. Multi-level noise identification: Based on the global sphere-level OOD tendency index and the sample-level consistency index, the training samples are divided into clean samples, in-distribution noise samples, out-of-distribution noise samples, and uncertain samples. The global sphere-level OOD tendency index integrates the relative position between spheres and the neighborhood label distribution information to determine the extent to which a sphere contains out-of-distribution noise samples. The sample-level consistency index is used to evaluate the rationality of the distribution of each sample within its sphere and the reliability of its label. This index combines the local neighborhood density and label consistency information to determine the extent to which a sample is a clean sample or an in-distribution noise sample. S3. Model Training: Construct a beneficial noise-guided joint loss function and use the sample partitioning results to perform differential training on the model. The joint loss function introduces the OOD dispersion loss, prototype softmax loss, and prototype margin loss in a weighted manner. S4: Intention Reasoning: Construct a multi-granularity decision boundary based on a granular structure to determine the category of the input sample.
2. A method for classifying open intent of noisy text based on a large model according to claim 1, characterized in that: The method of extracting text features using a pre-trained large language model includes: Freeze the bottom encoding layer of the pre-trained large language model and fine-tune the top layer structure; Extract text features using the adjusted pre-trained large language model.
3. The method for classifying open intent of noisy text based on a large model according to claim 1 is characterized in that: The calculation of the global granular-level OOD tendency index is as follows: , where j represents the ball number, Indicates the OOD tendency index value of the global particle level of the target particle, It represents the average distance between the centers of mass of the nearest spheres with the same pseudo-label as the target sphere. Indicates the proportion of inconsistent pseudo labels among the several spheres closest to the target sphere.
4. A method for classifying open intent of noisy text based on a large model according to claim 3, characterized in that: The sample-level consistency index is calculated as follows: , where i represents the sample number, represents a sample in the sphere, Indicates the consistency index value of a sample within the sphere, Indicates the tag purity of the pellet, represents the center of mass of the sphere, represents the maximum radius of the sphere, and exp() represents the exponential function with e as the base.
5. A method for classifying open intent of noisy text based on a large model according to claim 4, characterized in that: The OOD tendency index based on the global granular level and the sample-level consistency index divides the training samples into clean samples, in-distribution noise samples, out-of-distribution noise samples and uncertain samples, including: Based on the combined results of the global granular-level OOD tendency index and the sample-level consistency index, the following four categories of discrimination criteria are set: like and , it is judged as a clean sample; like and , then it is determined to be a noise sample within the distribution, represents the true label of the sample, Indicates the pellet label; like , it is determined to be a noise sample outside the distribution; In other cases, the samples are determined to be uncertain; in, and is the empirical threshold, obtained based on training set statistics.
6. A method for classifying open intent of noisy text based on a large model according to claim 1, characterized in that: The joint loss function is defined as: ,in, represents the joint loss function, represents the OOD dispersion loss, represents the prototype Softmax loss, represents the prototype interval loss, 、 、 is the weight coefficient of the loss term.
7. A noisy text open intent classification system based on a large model, characterized by: include: The granular structure modeling module is configured to extract text features using a pre-trained large language model, cluster the text features through unsupervised granular clustering, and construct intra-class structure representation in the feature space; The multi-level noise recognition module is configured to divide the training samples into clean samples, in-distribution noise samples, out-of-distribution noise samples, and uncertain samples based on the global sphere-level OOD tendency index and the sample-level consistency index. The global sphere-level OOD tendency index integrates the relative position between spheres and the neighborhood label distribution information to determine the extent to which a sphere contains out-of-distribution noise samples. The sample-level consistency index is used to evaluate the rationality of the distribution of each sample within its sphere and the reliability of its label. This index combines the local neighborhood density and label consistency information to determine the extent to which a sample is a clean sample or an in-distribution noise sample. A model training module is configured to construct a beneficial noise-guided joint loss function and perform differential training on the model using the sample partitioning results. The joint loss function introduces the OOD dispersion loss, prototype softmax loss, and prototype margin loss in a weighted manner. The intention reasoning module is configured to construct a multi-granularity decision boundary based on a granular sphere structure to determine the category of the input sample.
Citation Information
Patent Citations
Database simplification method and system based on granular ball face clustering image quality evaluation
CN114003752A
Weak supervision continuous text classification method and device for open environment
CN116401363A