Small sample language consciousness classification method and system

By adopting a small sample language awareness classification method with multi-head attention mechanism and dynamic mask training strategy in language intention classification, the problem of low accuracy of language intention classification in the prior art is solved, and more efficient language intention recognition and interactive guidance are achieved.

CN120067777AInactive Publication Date: 2025-05-30杭州威灿科技有限公司

Patent Information

Application Number
CN202510561643.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to adapt to the diversity and complexity of natural language in the classification of language intentions, resulting in low accuracy.

Method used

A small sample language awareness classification method based on multi-head attention mechanism is adopted, and the model's understanding of the intrinsic structure and semantic relationships of the language is improved through dynamic mask training strategies and preprocessing techniques.

Benefits of technology

It significantly improves the accuracy and adaptability of language intention classification, can more accurately identify unseen or non-standard expression intentions, and improves the effectiveness of small sample learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067777A_ABST
    Figure CN120067777A_ABST
Patent Text Reader

Abstract

The invention discloses a small sample language consciousness classification method and system, and relates to the field of model training, and the method comprises the steps: receiving a small sample supervised data set uploaded by a user; preprocessing the small sample supervised data set to obtain preprocessed training data; training a language awareness classification model by adopting a dynamic mask training strategy based on the preprocessed training data; in the process of training the language awareness classification model, the accuracy rate of the language awareness classification model is evaluated in real time, and when the accuracy rate of the language awareness classification model reaches a preset threshold value, the language awareness classification model is deployed to the electronic equipment; the electronic equipment calls a language awareness classification model to conduct intention classification on the obtained real-time dialogue text, and an intention classification result is obtained. The accuracy of language intention classification can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of model training, and in particular, to a small-sample language awareness classification method and system. Background Art

[0002] In modern information processing systems, language intent classification is a key task in the field of natural language processing, which is of great significance for improving the efficiency of human-computer interaction and realizing intelligent services. Especially in the legal and administrative fields, accurately identifying the intent in conversations is crucial for the efficiency and fairness of case handling.

[0003] Currently, text intent is mainly identified by manually defining keywords, phrases, and grammar rules, but this method relies on a rule base. For example, by setting keywords such as "I want" and "how" to judge the query intent of users.

[0004] However, due to the diversity, complexity, and real-time update nature of natural language, the above method has poor adaptability to newly emerging expressions and domain-specific terms, and it is difficult to adapt to the diversity and changes of language expressions, resulting in a low accuracy rate of language intent classification. Summary of the Invention

[0005] The embodiments of the present application provide a small-sample language awareness classification method and system for effectively improving the accuracy of language intent classification.

[0006] To achieve the above object, the embodiments of the present application adopt the following technical solutions:

[0007] In a first aspect, a small-sample language awareness classification method is provided, which is applied to an electronic device. The electronic device stores a language awareness classification model constructed based on a multi-head attention mechanism. The method includes:

[0008] Receiving a small-sample supervised data set uploaded by a user, where the small-sample supervised data set includes labeled conversation texts and corresponding intent categories;

[0009] Preprocessing the small-sample supervised data set to obtain preprocessed training data;

[0010] Based on the preprocessed training data, training the language awareness classification model using a dynamic masking training strategy, where the dynamic masking training strategy is used to randomly mask some tokens in the input text and control the model to learn semantic relationships from the context;

[0011] During the training process of the language awareness classification model, the accuracy of the language awareness classification model is evaluated in real time, and when the accuracy of the language awareness classification model reaches a preset threshold, the language awareness classification model is deployed to the electronic device, so that the electronic device can call the language awareness classification model to classify the intent of the obtained real-time conversation text and obtain an intent classification result.

[0012] In a possible implementation manner of the first aspect, the electronic device deploys an intent classification platform, the intent classification platform includes a data processing layer, a model training layer, and an application layer, the model training layer includes the language awareness classification model, and the application layer is connected to at least one database.

[0013] In another possible implementation manner of the first aspect, the deploying the language awareness classification model to the electronic device includes:

[0014] Deploy the language awareness classification model to the application layer of the electronic device, where the application layer is used to receive real-time conversation text, call the language awareness classification model to classify the intent of the real-time conversation text, obtain an intent classification result, determine a corresponding reply template from a preset reply template library according to the intent classification result, generate guiding reply content, and display the guiding reply content through the front-end interface of the electronic device to achieve conversation guidance.

[0015] In another possible implementation manner of the first aspect, the preprocessing of the small-sample supervised data set includes:

[0016] Perform text cleaning on the conversation text;

[0017] Use a pre-trained tokenizer to perform tokenization on the conversation text after text cleaning;

[0018] Convert the tokenized conversation text into a word vector representation to obtain preprocessed training data.

[0019] In another possible implementation manner of the first aspect, the application layer calls the language awareness classification model to classify the intent of the real-time conversation text and obtain an intent classification result, including:

[0020] The application layer preprocesses the real-time conversation text to obtain preprocessed conversation text;

[0021] Input the preprocessed conversation text into the language awareness classification model to obtain a probability distribution of intent categories;

[0022] Take the intent category with the highest probability as the intent classification result. Among them, when the intent classification result is lower than a preset confidence threshold, mark the intent classification result as an uncertain category.

[0023] In another possible implementation of the first aspect, the language awareness classification model constructed based on the multi-head attention mechanism is the language awareness classification model based on the Transformer architecture, and the language awareness classification model includes an encoder and a classification head;

[0024] Among them, the encoder includes multiple layers of multi-head self-attention layers and a feed-forward neural network layer for extracting the context representation of the real-time conversation text; the classification head includes a fully connected layer and a softmax layer for mapping the context representation to the probability distribution of intent categories.

[0025] In another possible implementation of the first aspect, training the language awareness classification model using a dynamic masking training strategy based on the preprocessed training data includes:

[0026] During the training process, randomly mask a preset number of tokens in the preprocessed training data so that the language awareness classification model predicts the masked tokens and the correct intent categories to obtain a prediction result;

[0027] Use the cross-entropy loss function to calculate the loss between the prediction result and the true label, and update the model parameters of the language awareness classification model through backpropagation according to the loss.

[0028] In another possible implementation of the first aspect, after deploying the language awareness classification model to the electronic device, it further includes:

[0029] Divide the small-sample supervised data set into K subsets of equal size to train the language awareness classification model K times, where each training selects a different subset as the validation set and the remaining subsets as the training set, and K is an integer greater than or equal to 2;

[0030] Calculate the average accuracy of K trainings;

[0031] When the average accuracy reaches a preset accuracy threshold, stop training and save the model parameters of the language awareness classification model.

[0032] In a second aspect, the present application provides an electronic device, including:

[0033] A memory configured to store instructions; and

[0034] A processor, configured to call the instructions from the memory and capable of implementing the above-mentioned few-shot language awareness classification method when executing the instructions.

[0035] In a third aspect, the present application provides a few-shot language awareness classification system, including:

[0036] The above-mentioned electronic device.

[0037] Through the above technical solution, a language awareness classification model is adopted. Its architecture can deeply capture the context dependencies in the input text through the multi-head self-attention mechanism, enabling the language awareness classification model to understand the true meaning of words in a specific context, rather than simply relying on surface keyword matching, effectively solving the problem that traditional methods are difficult to handle language complexity. In addition, by receiving the few-shot supervised data set uploaded by the user and performing model training after preprocessing, the demand for data is greatly reduced. Since it adopts a dynamic masking training strategy, during the training process, some tokens in the input text are randomly masked, forcing the model to not only predict the intention category but also use the context information to predict the masked tokens. This training method greatly enhances the model's understanding depth of the internal structure and semantic relationships of language, enabling it to learn more robust and generalizable language representations from limited data, improving the model's adaptability when facing unseen or non-standard expressions, and effectively enhancing the effect of few-shot learning. Additionally, after the model is deployed at the application layer, it can not only perform intention classification but also select or generate guiding response content from a preset response template library according to the classification result and display it through the front-end interface, realizing a closed-loop from intention understanding to actual dialogue guidance. In the legal and administrative fields, this technical solution can more accurately understand the user's consultation, request, or statement, provide more accurate guidance or information, thereby significantly improving service efficiency and processing accuracy. In summary, this solution effectively solves the problems of poor adaptability and low accuracy rate of traditional rule methods in language intention classification tasks, significantly enhancing the processing ability for the diversity and complexity of natural language, especially in the application scenarios of few-shot and legal administration, achieving more accurate language intention recognition and interaction guidance.

[0038] Other features and advantages of the embodiments of the present application will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic flowchart of a few-shot language awareness classification method provided by an embodiment of the present application;

[0040] Figure 2 It is a schematic structural diagram of an intention classification platform provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0041] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. It should be understood that the specific implementation manners described herein are only for explaining and interpreting the embodiments of this application, and are not used to limit the embodiments of this application. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0042] It should be noted that if there are directional indications (such as up, down, left, right, front, back,...) involved in the embodiments of this application, the directional indications are only used to explain the relative positional relationships and movement conditions among components in a specific posture (as shown in the accompanying drawings). If the specific posture changes, the directional indications will also change accordingly.

[0043] In addition, if there are descriptions involving "first", "second", etc. in the embodiments of this application, the descriptions of "first", "second", etc. are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions among various embodiments may be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by this application.

[0044] Figure 1 A flowchart showing the process of a few-shot language awareness classification method according to an embodiment of this application is schematically shown. As Figure 1 shown, an embodiment of this application provides a few-shot language awareness classification method, which is applied to an electronic device. The electronic device stores a language awareness classification model constructed based on the multi-head attention mechanism. The method may include the following steps.

[0045] S110. Receive a few-shot supervised data set uploaded by a user. The few-shot supervised data set includes labeled dialogue texts and corresponding intent categories;

[0046] S120. Preprocess the few-shot supervised data set to obtain preprocessed training data;

[0047] S130. Based on the preprocessed training data, train the language awareness classification model using a dynamic masking training strategy. The dynamic masking training strategy is used to randomly mask some tokens in the input text and control the model to learn semantic relationships from the context;

[0048] S140. During the training process of the language awareness classification model, the accuracy rate of the language awareness classification model is evaluated in real time. When the accuracy rate of the language awareness classification model reaches the preset threshold, the language awareness classification model is deployed to the electronic device so that the electronic device can call the language awareness classification model to classify the intent of the obtained real-time conversation text and obtain the intent classification result.

[0049] First, receive the small supervised data set uploaded by the user. The small supervised data set includes the labeled conversation text and the corresponding intent categories. In this embodiment, the electronic device provides a data upload function through the user interface, allowing the user to upload the pre-prepared small supervised data set to the system. The small supervised data set refers to a data set with a relatively small amount of data (usually 10 - 100 samples for each intent category) compared to the large-scale training data required by traditional deep learning models, but each sample is manually labeled to ensure data quality.

[0050] This data set mainly includes two parts: conversation text and the corresponding intent category labels. The conversation text refers to the written record of human-machine or human-human conversations in a specific scenario, which can include various language forms such as questions, statements, requests, etc. The intent category refers to the classification of the user intent expressed by the conversation text. For example, in the legal and administrative field, it can include categories such as "consult legal provisions", "apply to consult legal professionals", "submit evidence materials", "voluntarily accept punishment", etc.

[0051] The format of the data set is usually a structured JSON or CSV file, and each record contains the original conversation text and the corresponding intent category label. It should be noted that during data reception, preliminary verification will be performed to check whether the data format meets the requirements, whether the labels are complete, and count the sample distribution of each intent category to ensure the balance of the data set. If data imbalance is found, such as too few or too many samples in some categories, the user will be prompted to adjust the data set structure to improve the effect of subsequent training.

[0052] The data preprocessing process mainly includes three sub-steps: text cleaning, word segmentation processing, and vectorization conversion. First, in the text cleaning link, the original conversation text is normalized, including removing redundant white spaces, standardizing punctuation marks, unifying case, processing special characters, converting emoji, etc. Especially for the text in the legal and administrative field, specific abbreviations, terms, and citation formats also need to be processed to ensure text consistency.

[0053] Secondly, use a pre-trained tokenizer to tokenize the cleaned text. In a Chinese environment, character-level tokenization or tokenization methods combined with a vocabulary, such as BERT's WordPiece tokenizer, are usually adopted. The tokenization process splits continuous text into discrete tokens and assigns a unique ID to each token. For out-of-vocabulary words (i.e., words not present in the tokenizer's vocabulary), the tokenizer breaks them down into smaller sub-word units to ensure that any input text can be effectively processed. Additionally, the tokenizer adds special tokens, such as the sentence start token and the sentence end token, for subsequent model processing. Finally, convert the tokenized token sequence into a word vector representation. A pre-trained language model, such as BERT's embedding layer, can be used to map each token ID to a dense vector of a fixed dimension (usually 768 dimensions). These vectors capture the semantic information and context relationships of the tokens, providing rich feature representations for subsequent model training. At the same time, an attention mask can also be generated to identify valid tokens and padding tokens, ensuring that the model only focuses on valid information.

[0054] After preprocessing, the original text data is converted into a structured numerical representation, including an input ID sequence, an attention mask, a token type ID, etc., for subsequent model training.

[0055] The dynamic masking training strategy is used to randomly mask some tokens in the input text and control the model to learn semantic relationships from the context. The dynamic masking training strategy of this embodiment is specifically optimized for small-sample intent classification tasks.

[0056] In traditional supervised learning, the model directly learns the mapping relationship from the input text to the intent label, and overfitting is likely to occur, especially in small-sample scenarios. The dynamic masking training strategy forces the model to learn more robust feature representations and improve generalization ability by randomly masking some tokens in the input text. Specifically, in each training iteration, the system randomly selects 15% of the tokens in the input text for masking and replaces these tokens with special tokens. The selection of the mask follows a specific strategy: 80% of the selected tokens are replaced with special tokens, 10% are replaced with random tokens, and 10% remain unchanged. The above hybrid strategy can reduce the inconsistency between pre-training and fine-tuning and improve the robustness of the model.

[0057] It is worth noting that the masking operation is dynamic, that is, new masking positions are generated in each training epoch to ensure that the model can fully learn different parts of the text. During the training process, on the one hand, it is necessary to predict the masked tokens, and on the other hand, it is necessary to correctly classify the intent category of the text. The above multi-task learning paradigm prompts the model to establish a deeper semantic understanding ability. The loss function of the model consists of two parts: the masked language model loss and the intent classification loss, and the total loss is obtained by weighted summation. Among them, the masked language model loss uses cross-entropy to calculate the difference between the predicted tokens and the true tokens, and the intent classification loss also uses cross-entropy to calculate the difference between the predicted intent and the true intent label. The Adam optimizer is used in the training process, the learning rate is set to 2e-5, and a linear learning rate decay strategy is used.

[0058] To prevent overfitting, weight decay and gradient clipping techniques can also be introduced. In addition, considering the characteristics of the small-sample scenario, an early stopping mechanism is also adopted in the training process, that is, when the performance on the validation set does not improve for several consecutive rounds, the training is automatically stopped to avoid overfitting the training data.

[0059] During the training process of the language awareness classification model, the accuracy of the language awareness classification model is evaluated in real time, and when the accuracy of the language awareness classification model reaches a preset threshold, the language awareness classification model is deployed to the electronic device so that the electronic device can call the language awareness classification model to classify the intent of the acquired real-time conversation text to obtain an intent classification result. The cross-validation method can be used to divide the small-sample data set into a training set and a validation set, usually in a ratio of 8:2. After each training round ends, the model performance is automatically evaluated on the validation set, and the evaluation metrics are calculated.

[0060] The evaluation metric can be accuracy. When the accuracy of the model reaches a preset threshold (usually set to 95% or higher), the model saving and deployment process is automatically triggered. Model saving includes complete information such as model parameters, configuration information, and vocabulary to ensure that the model can be fully restored and used. The model deployment process refers to deploying the optimized model to the application layer of the electronic device and integrating it into the existing business process. After deployment, the electronic device can receive conversation text in real time and call the deployed language awareness classification model for processing. The processing process includes: performing the same preprocessing as in the training stage on the real-time conversation text, inputting the preprocessed text into the model, obtaining the probability distribution of the intent categories output by the model, and selecting the category with the highest probability as the classification result. At the same time, the confidence of the classification is recorded. When the confidence is lower than the preset threshold (usually 0.7), the result is marked as an uncertain category, indicating that manual intervention or further confirmation is required.

[0061] Based on the classification result, the corresponding reply template can be selected from the preset reply template library to generate the guiding reply content, which is then displayed to the user through the front-end interface of the electronic device to achieve intelligent dialogue guidance.

[0062] In summary, this solution significantly improves the effect of language intention classification by introducing a few-shot language awareness classification method based on the multi-head attention mechanism, effectively overcoming many drawbacks of the traditional rule-based method mentioned in the background art. The traditional method relies heavily on manually defined keywords and rule libraries, which is not only time-consuming and laborious, but also has a low accuracy rate when faced with the inherent diversity, complexity, and continuously evolving characteristics of natural language. Newly emerging expressions, domain-specific terms, or slang are often not covered by the preset rules, resulting in low accuracy of intention recognition in practical applications, especially in professional fields such as law and administration that require high precision and strong adaptability, directly affecting the efficiency of human-computer interaction, the quality of intelligent services, and even the efficiency and fairness of case handling. In the application in the field of law and administration, this embodiment can accurately identify the dialogue intention, realize the human-machine collaborative guidance for the investigated object to admit guilt and accept punishment or voluntarily accept the punishment, and strongly promote the rapid handling of administrative cases and the quick adjudication of minor cases in judicial procedures.

[0063] This embodiment uses a language awareness classification model. Its architecture, through the multi-head self-attention mechanism, can deeply capture the context dependencies in the input text, enabling the language awareness classification model to understand the true meaning of words in a specific context rather than relying solely on surface keyword matching. This effectively solves the problem that traditional methods are difficult to handle language complexity. In addition, by receiving the small supervised dataset uploaded by the user and performing model training after preprocessing, the data requirements are greatly reduced. Due to its adoption of the dynamic masking training strategy, during the training process, some tokens in the input text are randomly masked, forcing the model to not only predict the intent category but also use the context information to predict the masked tokens. This training method greatly enhances the model's understanding depth of the internal structure and semantic relationships of language, enabling it to learn more robust and generalized language representations from limited data, improving the model's adaptability when facing unseen or non-standard expressions, and effectively enhancing the effect of few-shot learning. Additionally, after the model is deployed at the application layer, it can not only perform intent classification but also select or generate guiding response content from the preset response template library according to the classification results and display it through the front-end interface, realizing a closed-loop from intent understanding to actual dialogue guidance. In the legal and administrative fields, this technical solution can more accurately understand the user's inquiries, requests, or statements, provide more accurate guidance or information, thereby significantly improving service efficiency and processing accuracy. In summary, this solution effectively solves the problems of poor adaptability and low accuracy of traditional rule methods in language intent classification tasks, significantly enhancing the processing ability for the diversity and complexity of natural language, especially in the few-shot and legal administrative application scenarios, achieving more accurate language intent recognition and interaction guidance.

[0064] In one implementation manner of this embodiment, an intent classification platform is deployed on the electronic device. The intent classification platform includes a data processing layer, a model training layer, and an application layer. The model training layer includes a language awareness classification model, and the application layer is connected to at least one database.

[0065] Figure 2 The following shows a schematic structural diagram of an intent classification platform provided by an embodiment of the present application, as Figure 2 As shown, in this embodiment, the intent classification platform deployed on the electronic device adopts a hierarchical architecture design. The overall architecture of the intent classification platform includes a data processing layer, a model training layer, and an application layer, and each layer interacts through standardized interfaces. In actual deployment, the intent classification platform can be deployed on a single server or adopt a distributed deployment method, deploying different layers on different server clusters to meet the needs of different scale application scenarios.

[0066] The data processing layer is used for tasks such as the collection, cleaning, annotation, and feature extraction of raw data, providing high-quality data support for model training and applications. The data processing layer includes a data collection module, a data cleaning module, a data annotation module, and a feature extraction module.

[0067] Specifically, the data collection module supports the access of multiple data sources, including but not limited to structured databases, semi-structured log files, unstructured text documents, and real-time data streams. In the legal and administrative fields, data sources can include historical case records, legal provisions, administrative regulations, user consultation records, etc. Data collection adopts a combination of batch processing and stream processing. Batch processing is used to process historical data, and stream processing is used to process real-time generated data. The data cleaning module is responsible for cleaning and normalizing the collected raw data, including operations such as removing noise data, handling missing values, correcting format errors, and deduplication. The data annotation module is used to add labels to the cleaned data, supporting both manual annotation and automatic annotation. Among them, automatic annotation can use a rule engine or a pre-trained model for preliminary annotation, and then be reviewed and corrected by humans. The feature extraction module is used to extract effective features from the annotated data.

[0068] The model training layer is used to train a language awareness classification model based on the processed data. The model training layer includes a model selection module, a parameter configuration module, a training execution module, and an evaluation and optimization module. The model selection module provides a variety of preset model architectures, including traditional machine learning models (such as support vector machines, random forests, etc.) and deep learning models. In the legal and administrative fields, considering the complexity and professionalism of text understanding, a pre-trained language model is usually selected as the basis, such as a legal domain-specific BERT model. The parameter configuration module is responsible for setting the hyperparameters of model training, including the learning rate, batch size, number of training epochs, optimizer type, etc. The selection of hyperparameters can be automatically determined by methods such as grid search, random search, or Bayesian optimization, or can be manually set by experts based on experience. The training execution module is responsible for the actual model training process, supporting both single-machine training and distributed training modes. Single-machine training is suitable for small models or datasets, while distributed training is suitable for large models or large-scale datasets. The evaluation and optimization module is used to evaluate and optimize the trained model, and the evaluation metrics include accuracy, precision, recall, F1 score, etc.

[0069] The language awareness classification model adopts a deep learning model based on the Transformer architecture, which can effectively capture semantic information and context relationships in the text. The input of the model is a text sequence that has been tokenized and encoded, and the output is the probability distribution of each intention category. The model architecture consists of three parts: an embedding layer, an encoding layer, and a classification layer. The embedding layer converts the input data into a dense vector representation. The encoding layer is composed of multiple Transformer blocks, and each Transformer block contains a multi-head self-attention mechanism and a feed-forward neural network, which can capture long-range dependencies in the sequence.

[0070] The classification layer maps the output of the encoding layer to the intention category space, usually implemented using a fully connected layer plus a softmax activation function. The model is trained using a cross-entropy loss function, and the model parameters are updated through the backpropagation algorithm.

[0071] The application layer is the front end of the intention classification platform, which is used to receive user requests, call the model for inference, generate responses, and interact with users. The application layer adopts a microservices architecture, including components such as an API gateway, service registration and discovery, load balancing, service orchestration, monitoring and alerting.

[0072] Specifically, the application layer is connected to at least one database, which is used to store and manage various types of data required for the operation of the system. The database can be a relational database, an in-memory database, etc. Among them, the relational database can be MySQL, which is used to store structured data such as user information, case records, operation logs, etc. The graph database is used to store and query complex relational data, such as a legal knowledge graph, and supports efficient graph traversal and path query. The database connection adopts connection pool technology.

[0073] The main data stored in the database includes four categories: user data, conversation data, model data, and system data. User data includes user basic information, permission settings, usage history, etc., which is used for user management and personalized services. Conversation data includes conversation history, intention classification results, user feedback, etc., which is used for conversation management and system optimization. Model data includes model parameters, feature dictionaries, evaluation results, etc., which is used for model deployment and update. System data includes configuration information, operation logs, performance metrics, etc., which is used for system maintenance and monitoring.

[0074] The interaction between the application layer and the database can adopt the Data Access Object (DAO) pattern, which encapsulates the data access logic in a dedicated class and separates it from the business logic. The DAO layer provides a unified data access interface, shielding the implementation details of the underlying database, making it possible to easily switch different database implementations without affecting the upper-layer business logic.

[0075] The intent classification platform of this embodiment is applicable to complex application scenarios in the legal and administrative fields, capable of effectively handling professional terms, complex semantics, and multi-turn conversations, providing users with accurate and timely intent recognition and response services, and significantly improving the intelligent level of administrative processing and judicial services.

[0076] In one implementation of this embodiment, deploying the language awareness classification model to an electronic device includes the following steps:

[0077] S210. Deploy the language awareness classification model to the application layer of the electronic device. The application layer is used to receive real-time conversation text, call the language awareness classification model to perform intent classification on the real-time conversation text to obtain an intent classification result, determine a corresponding reply template from a preset reply template library according to the intent classification result, generate guiding reply content, and display the guiding reply content through the front-end interface of the electronic device to achieve conversation guidance.

[0078] In this embodiment, the model deployment adopts a hierarchical architecture design. Deploying the language awareness classification model to the application layer of the electronic device realizes the effective integration of the model and the business logic. Specifically, the application layer is the core functional layer in the software architecture of the electronic device, located between the data layer and the presentation layer, and is responsible for implementing the main business logic and function processing of the system. To deploy the language awareness classification model to the application layer, it is first necessary to optimize the trained model, including model compression, quantization, and format conversion. Model compression uses knowledge distillation technology to transfer the knowledge of the original large model to a smaller model with fewer parameters, significantly reducing the model size while maintaining the classification performance. For example, the original BERT-based classification model may have 110M parameters, which can be compressed to a lightweight model of about 30M parameters through knowledge distillation. Model quantization is to convert the floating-point parameters of the model (usually in FP32 format) to a low-precision representation (such as INT8 or INT4 format), reducing each parameter from occupying 32 bits to 8 bits or 4 bits, further reducing the model size and improving the inference speed. Format conversion is to convert the model from the training framework format to the format supported by the deployment framework to adapt to different hardware environments and inference engines.

[0079] The optimized model is integrated into the application layer through the model service framework. The application layer adopts a microservices architecture, taking the model service as an independent microservices component and communicating with other service components through RESTful API interfaces. The above architecture design has high scalability and fault tolerance, allowing the model service to be independently extended and updated without affecting other parts of the system.

[0080] On the front-end interface, users can provide conversation content by means of a text input box, voice input (converted into text by combining speech recognition technology), or uploading a text file. In addition, conversation text from other systems or applications can be received through an API interface to achieve seamless integration between systems. For the received real-time conversation text, the application layer first performs preprocessing, including text normalization, word segmentation, and feature extraction. The preprocessing adopts the same process and parameters as in the model training stage to ensure the consistency of the input data and avoid performance degradation caused by differences in data distribution.

[0081] After the preprocessing is completed, the application layer calls the deployed language awareness classification model to perform intent classification on the processed text. Specifically, the application layer sends the text features after preprocessing to the model service through a predefined interface. After receiving the request, the model service inputs the features into the language awareness classification model. Through the calculation of the multi-head attention mechanism, the probability distribution of each intent category is obtained. The output of the model is a probability vector. The application layer determines the final intent classification result based on the probability vector, usually selecting the category with the highest probability as the classification result.

[0082] At the same time, the confidence level of the classification can also be calculated, that is, the highest probability value. When the confidence level is lower than the preset threshold, the classification result is marked as uncertain and may require manual intervention or further confirmation, effectively avoiding the model making wrong decisions in uncertain situations.

[0083] After obtaining the intent classification result, the application layer needs to determine the corresponding reply template from the preset reply template library. The reply template library is a structured data set that contains reply templates designed in advance for different intent categories. Each intent category may correspond to multiple reply templates to increase the diversity and naturalness of the replies. The design of the reply template library adopts a hierarchical structure. First, it is grouped by intent category. Each category contains multiple templates, and each template contains a fixed part and a variable part. The fixed part is the framework of the template and remains unchanged; the variable part is the content that needs to be filled dynamically according to the context, such as user name, time, location, etc. For example, for the intent of consulting legal provisions, the reply template may be "According to Article {clause number} of the '{law name}', {clause content}. Do you have any other questions about the {legal field}?" The content within the curly brackets is the variable part that needs to be filled dynamically.

[0084] The application layer queries the reply template library according to the intent classification result to obtain the template set corresponding to the intent category. To increase the naturalness of the reply and avoid repetition, a template can be randomly selected from the template set.

[0085] After selecting a reply template, the application layer needs to generate guiding reply content, including two steps: variable filling and content optimization. Variable filling replaces the variable parts in the template with actual content, and the content source may be information input by the user, relevant data in the system database, or information obtained by calling external services through APIs. For example, when a user consults a specific legal clause, relevant clause content can be retrieved from a legal database and filled into the reply template. Content optimization is to polish and adjust the generated reply to ensure that the language is fluent and natural, conforms to grammar norms, and adjusts the language style and professionalism according to user characteristics (such as professional background, age, etc.). In the legal and administrative field, the guiding direction of the reply can be adjusted according to the case type and processing stage. For example, for cases of voluntarily accepting penalties, the reply content will guide the user to understand the legal consequences and procedures of voluntarily accepting penalties.

[0086] After generating the guiding reply content, the application layer displays the reply content to the user through the front-end interface of the electronic device.

[0087] In summary, this embodiment can achieve intelligent dialogue guidance based on language awareness classification, can accurately understand the user's intention, provide targeted replies and guidance, and significantly improve the efficiency of human-computer interaction and the user experience. Especially in the legal and administrative field, this guiding mechanism can help non-professional users understand complex legal concepts and procedures and promote the efficient handling of cases.

[0088] This implementation method adopts model compression and quantization technologies to enable complex language models to run efficiently on resource-constrained electronic devices; secondly, the deployment method based on the microservice architecture provides good scalability and fault tolerance, adapting to application scenarios of different scales; the user-friendly design of the front-end interface provides an intuitive and friendly user experience, which is particularly suitable for case handling and user consultation scenarios in the legal and administrative field, and can significantly improve service efficiency and user satisfaction.

[0089] In one implementation manner of this embodiment, preprocessing is performed on the small-sample supervised data set, including the following steps:

[0090] S310. Perform text cleaning on the dialogue text;

[0091] S320. Use a pre-trained tokenizer to perform tokenization on the dialogue text after text cleaning;

[0092] S330. Convert the dialogue text after tokenization into a word vector representation to obtain the preprocessed training data.

[0093] In this embodiment, text cleaning includes special character processing, punctuation normalization, case unification, typo correction, and / or noise data filtering.

[0094] Specifically, special character processing can identify and process special characters in the text through regular expressions, such as HTML tags, XML tags, URL links, email addresses, etc. These special characters can be either deleted or replaced with standardized markers. Punctuation normalization can convert different forms of punctuation marks (such as Chinese and English punctuation, full-width and half-width punctuation) into a standard format. Case unification can convert the text into lowercase or maintain the original case according to requirements, reducing vocabulary inconsistencies caused by case differences. Typo correction can identify and correct common typos and spelling mistakes through dictionary matching or rule-based methods. Noise data filtering can identify and process noise data in the text, such as duplicate text, meaningless strings, spam, etc.

[0095] Using a pre-trained tokenizer to tokenize the cleaned dialogue text can convert continuous text into a sequence of discrete tokens. A pre-trained tokenizer refers to a tokenization model pre-trained on a large-scale corpus that can accurately tokenize text based on the characteristics of the language and context information. For example, for languages without obvious delimiters such as Chinese, a pre-trained BERT tokenizer can be used. This tokenizer is based on the WordPiece algorithm and can split text into sub-word units, effectively handling out-of-vocabulary words.

[0096] Converting the tokenized dialogue text into a word vector representation to obtain preprocessed training data is to convert the discrete token sequence into a continuous vector space representation, enabling machine learning models to effectively process and understand text data. Word vector representation maps words into a high-dimensional vector space, making words with similar semantics closer in the vector space, thereby capturing the semantic relationships between words. In this embodiment, the word vector representation adopts the output of the embedding layer of a pre-trained language model. The specific implementation is as follows: Input the token sequence after tokenization into the embedding layer of the pre-trained language model to obtain the vector representation of each token. The embedding layer of the BERT model maps each token to a vector of a fixed dimension (usually 768 dimensions).

[0097] Through preprocessing, this embodiment effectively eliminates noise and non-standard content in the original text, improving data quality; the application of the pre-trained tokenizer realizes accurate tokenization of different language texts; the word vector representation method based on the pre-trained language model converts the discrete token sequence into a continuous vector space representation, effectively capturing the semantic information and context relationships of words, which is suitable for small-sample learning scenarios. By means of high-quality data representation, it maximally utilizes limited labeled data and significantly improves the model's ability to understand complex semantics and recognize intentions in the legal and administrative fields.

[0098] In one implementation of this embodiment, the application layer calls a language awareness classification model to classify the intent of the real-time dialogue text, and the intent classification result is obtained, including the following steps:

[0099] S410. The application layer preprocesses the real-time dialogue text to obtain the preprocessed dialogue text;

[0100] S420. Input the preprocessed dialogue text into the language awareness classification model to obtain the probability distribution of the intent categories;

[0101] S430. Use the intent category with the highest probability as the intent classification result. When the intent classification result is lower than the preset confidence threshold, mark the intent classification result as an uncertain category.

[0102] The application layer preprocesses the real-time dialogue text to convert the original dialogue text into a standardized and structured form so that the subsequent model can process it effectively. The preprocessing of the real-time dialogue text is consistent with the preprocessing of the training data to ensure the consistency and comparability of the model input. The preprocessing process first performs text cleaning, including removing HTML tags, special characters, and redundant whitespace characters, unifying the punctuation format, and correcting common typos and spelling mistakes. Secondly, context integration is performed to integrate the current user input with the historical dialogue turns to construct a complete dialogue context. In actual applications, it is possible to choose to retain the last N dialogue turns, such as the last 3 turns, to capture the coherence and context information of the dialogue.

[0103] Then, the cleaned text is tokenized using the same pre-trained tokenizer as in the training phase, splitting the continuous text into a sequence of tokens.

[0104] Next, special tokens are added, such as adding the [CLS] token at the beginning of the sequence and the [SEP] token at the end of the sequence. For multi-turn dialogues, the [SEP] token is added between different turns for separation. Finally, the tokenized token sequence is converted into a word vector representation, using the embedding layer of the same pre-trained language model as in the training phase to map each token to a vector of a fixed dimension.

[0105] Input the preprocessed dialogue text into the trained language awareness classification model to obtain the probability distribution of the intent categories.

[0106] Finally, the intent category with the highest probability is taken as the intent classification result. When the intent classification result is lower than the preset confidence threshold, the intent classification result is marked as an uncertain category. After obtaining the probability distribution of the intent categories, it is first necessary to determine the most likely intent category, that is, the category with the highest probability value. To improve the reliability of classification, a confidence threshold mechanism is introduced. The confidence threshold is a preset probability value used to judge the certainty of the model about the prediction result. If the highest probability value is lower than this threshold, it is considered that the model is not confident enough about the prediction result, and the result is marked as an uncertain category.

[0107] In practical applications, for queries marked as uncertain categories, multiple processing strategies can be adopted: one is to request the user to clarify or provide more information; the second is to provide multiple potentially relevant intent options for the user to choose from; the third is to forward the query to a human customer service for processing to ensure that users receive accurate help.

[0108] In this embodiment, the real-time dialogue text preprocessing link maintains consistency with the training data preprocessing. Through text cleaning, context integration, and vector representation, the quality of the input data is ensured. Secondly, the intent classifier based on the pre-trained language model can effectively understand the text semantics and context information, capture long-distance dependencies through the self-attention mechanism, and generate accurate probability distributions of intent categories. Finally, the introduction of the confidence threshold mechanism enables the identification and proper handling of uncertain situations, avoiding misclassification-induced misleading. It is applicable to the complex semantic understanding in the legal and administrative fields, can accurately identify the specific legal problem types consulted by users, and provides a precise guidance for subsequent professional answers.

[0109] In one implementation manner of this embodiment, the language awareness classification model constructed based on the multi-head attention mechanism is a language awareness classification model based on the Transformer architecture. The language awareness classification model includes an encoder and a classification head;

[0110] Among them, the encoder includes multiple layers of multi-head self-attention layers and a feed-forward neural network layer for extracting the context representation of the real-time dialogue text; the classification head includes a fully connected layer and a softmax layer for mapping the context representation to the probability distribution of intent categories.

[0111] The language awareness classification model constructed based on the multi-head attention mechanism is a language awareness classification model based on the Transformer architecture. This model can effectively process sequence data and capture long-distance dependencies. In the legal text understanding scenario, the Transformer architecture is suitable for processing long texts and complex semantics, and can accurately capture the relationships and context dependencies between legal terms. The core of this architecture is the self-attention mechanism, which allows the model to consider the information of all other positions in the sequence when processing each position in the sequence, thereby forming a global perception ability.

[0112] Specifically, for each token in the input sequence, the self-attention mechanism calculates the degree of association between the token and all tokens in the sequence (including itself), and then based on these degrees of association, weights and sums up the representations of all tokens to generate a context-sensitive representation of the token.

[0113] Multi-head attention further extends this mechanism by calculating multiple different attention sets in parallel, enabling the model to simultaneously focus on information patterns in different subspaces. The Transformer architecture can effectively handle complex semantic relationships and long-distance dependencies in legal texts, providing a strong foundation for intent classification.

[0114] The encoder consists of multiple layers of multi-head self-attention layers and feed-forward neural network layers, which are used to extract the context representation of real-time dialogue texts and are responsible for converting the input text into a vector representation with rich semantic information. The encoder adopts a stacked structure, usually containing 6 - 24 identical layers, and each layer consists of two sub-layers: a multi-head self-attention layer and a feed-forward neural network layer. In the multi-head self-attention layer, the input first undergoes layer normalization, then calculates the context representation through the multi-head self-attention mechanism, and finally applies a residual connection to add the original input and the attention output to form the output of this sub-layer. Layer normalization helps to stabilize the training of deep neural networks by normalizing the feature dimensions of each sample and reducing the problem of internal covariate shift. The residual connection alleviates the problem of gradient vanishing in deep neural networks, allowing information to be directly transmitted from the bottom layer to the top layer.

[0115] After the multi-head self-attention layer is the feed-forward neural network layer, which independently applies the same fully connected network to each position in the sequence. This fully connected network usually contains two linear transformations with a ReLU activation function in the middle. The role of the feed-forward network layer is to introduce non-linear transformations, enhance the expressive power of the model, and process patterns that the self-attention layer may not have captured. Similar to the multi-head self-attention layer, the feed-forward network layer also uses layer normalization and residual connections. By stacking multiple such structures, the encoder can gradually extract and refine the semantic representation of the text, from surface features to deep semantics, to construct a rich context representation. In the final layer, the representation of a special token is usually used as the aggregated representation of the entire sequence for subsequent classification tasks. This multi-level encoding structure enables the model to understand complex concepts and relationships in legal texts, laying a foundation for accurate intent classification.

[0116] The classification head includes a fully connected layer and a softmax layer, which are used to map the context representation to the probability distribution of intent categories. It is the output component of the language awareness classification model and is responsible for converting the semantic representation extracted by the encoder into specific intent category predictions. The classification head receives the context representation output by the encoder, especially the representation vector of the special token, and maps it to the predefined intent category space.

[0117] The fully connected layer is the first component of the classification head. It performs a linear transformation on the input vector, mapping the high-dimensional semantic space to the intent category space.

[0118] After the fully connected layer, the softmax layer converts the score vector into a probability distribution, ensuring that the sum of probabilities for all categories is 1. Through the softmax transformation, the model outputs a probability distribution representing the likelihood of the input text belonging to each intent category.

[0119] This embodiment fully utilizes the advantages of the multi-head self-attention mechanism, can effectively capture long-range dependencies and complex semantic structures in legal texts, understand the context information and the associations between professional terms. The multi-layer stacked structure of the encoder realizes the gradual abstraction from surface features to deep semantics. Through the alternating action of the self-attention mechanism and the feed-forward neural network, it generates context representations rich in semantic information. The classification head then accurately maps these high-dimensional semantic representations to the predefined intent category space. Through the combination of the fully connected layer and the softmax layer, it outputs an interpretable probability distribution, which is applicable to the intent recognition task in the legal consultation scenario, can accurately understand the types of legal questions consulted by users, and can maintain a high accuracy even when facing texts with ambiguous wording or rich professional terms.

[0120] In one implementation of this embodiment, based on the preprocessed training data, a dynamic masking training strategy is adopted to train the language awareness classification model, including the following steps:

[0121] S510. During the training process, randomly mask a preset number of tokens in the preprocessed training data, so that the language awareness classification model predicts the masked tokens and the correct intent categories to obtain a prediction result;

[0122] S520. Use the cross-entropy loss function to calculate the loss between the prediction result and the true label, and update the model parameters of the language awareness classification model through backpropagation according to the loss.

[0123] During the training process, randomly mask a preset number of tokens in the preprocessed training data, so that the language awareness classification model predicts the masked tokens and the correct intent categories to obtain a prediction result. This step adopts a dynamic masking training strategy. By randomly masking some tokens in the input text, the model is forced to learn context semantic relationships and improve its language understanding ability.

[0124] In specific implementation, first determine the masking ratio, usually set to 15% of the input sequence length. This ratio has been experimentally verified to achieve a good balance between training efficiency and model performance. For example, for an input sequence of length 100, approximately 15 tokens will be masked. The masking process adopts a random selection strategy. For the tokens selected for masking, there is an 80% probability of being replaced with a special token, a 10% probability of being replaced with a random token from the vocabulary, and the remaining 10% probability of remaining unchanged, aiming to prevent the model from deviating during the inference stage because the [MASK] token does not appear in actual applications.

[0125] In actual implementation, the masking operation is usually dynamically generated for each training batch rather than being fixed in advance. As the name implies, it represents the meaning of dynamic masking. This effectively ensures that the model can see different masked versions of the same text during training, enhancing the model's robustness and generalization ability. For legal text processing, since legal terms and expressions usually have strict context dependencies, by predicting the masked tokens, the model can learn the associations between legal concepts and the usage of professional terms in the context. At the same time, the model also needs to predict the correct intent category, which forms a multi-task learning framework: on the one hand, predicting the masked tokens, and on the other hand, predicting the intent category of the entire text. This multi-task learning method can prompt the model to simultaneously focus on local lexical information and global semantic understanding, which is particularly suitable for the intent recognition requirements in legal consultation scenarios.

[0126] During the implementation process, the masking ratio and strategy can be adjusted according to specific application scenarios. For example, for key legal terms, the masking probability can be appropriately increased to strengthen the model's understanding of professional terms; or different masking strategies can be adopted for different types of legal texts (such as litigation documents, consultation questions, regulations, etc.). Through this dynamic masking training, the model can more effectively learn the internal structure and semantic relationships of the language under the condition of limited labeled data, improving the understanding ability of legal texts and the accuracy of intent classification.

[0127] The cross-entropy loss function is used to calculate the loss between the prediction result and the true label, and the model parameters of the language awareness classification model are updated through backpropagation according to the loss. By defining appropriate loss functions and optimization algorithms, the model parameters can be gradually adjusted to improve the prediction accuracy. In this embodiment, the cross-entropy loss function is used as the optimization objective, which is one of the most commonly used loss functions in classification tasks and is particularly suitable for multi-class classification problems. The cross-entropy loss function measures the difference between the predicted probability distribution and the true label distribution, and its mathematical expression is: ;

[0128] where C is the number of classes, y iis the one-hot encoding of the true label (the target class is 1 and other classes are 0), p i is the probability that the model predicts the i-th class.

[0129] For the intent classification task, the cross-entropy loss calculates the difference between the probability distribution of the intent classes predicted by the model and the true intent labels; for the masked language modeling task, it calculates the difference between the probability distribution of the tokens predicted by the model and the true labels of the masked tokens.

[0130] In the multi-task learning framework, the total loss is usually the weighted sum of these two parts of losses. After calculating the loss, the gradients of the loss function with respect to the model's various parameters are calculated through the backpropagation algorithm, and the optimizer is used to update the model parameters. Backpropagation uses the chain rule to calculate the gradients layer by layer from the output layer to the input layer, achieving efficient updates of all parameters in the network. For the legal text domain, domain adaptation techniques can be adopted, such as pre-training on general corpora and then fine-tuning on legal corpora, or using adversarial training to enhance the robustness of the model. The training process usually sets an early stopping mechanism, that is, when the performance metrics (such as accuracy, F1 score, etc.) on the validation set do not improve for several consecutive rounds, the training is terminated early to prevent overfitting.

[0131] This embodiment randomly masks some tokens in the input text, forcing the model to learn context semantic relationships, which not only enhances the model's ability to understand legal terms and expressions, but also improves its robustness in dealing with incomplete or ambiguous expressions. The multi-task learning framework simultaneously optimizes two objectives of masked language modeling and intent classification, enabling the model to understand the global semantic structure while capturing local lexical information, which is particularly suitable for the complex language understanding requirements in the legal consultation scenario. The cross-entropy loss function accurately measures the difference between the prediction and the true label, providing an accurate optimization direction for the model, while the backpropagation algorithm ensures efficient parameter updates. This training method introduces self-supervised learning elements through the masked prediction task, making full use of the language rules in the unlabeled text, providing a reliable foundation for intent understanding for the intelligent legal consultation system.

[0132] In one implementation of this embodiment, after deploying the language awareness classification model to the electronic device, the following steps are further included:

[0133] S610: Divide the small-sample supervised dataset into K subsets of equal size to train the language awareness classification model K times. Among them, each training selects a different subset as the validation set, and the remaining subsets as the training set, where K is an integer greater than or equal to 2;

[0134] S620: Calculate the average accuracy of the K times of training;

[0135] S630. When the average accuracy rate reaches the preset accuracy threshold, stop training and save the model parameters of the language awareness classification model.

[0136] Divide the small-sample supervised dataset into K subsets of equal size to train the language awareness classification model K times. Among them, for each training, a different subset is selected as the validation set, and the remaining subsets are used as the training set, where K is an integer greater than or equal to 2.

[0137] Specifically, during implementation, first randomly shuffle the entire small-sample supervised dataset, and then evenly divide it into K non-overlapping subsets, with each subset containing approximately the same number of samples. The choice of the K value usually depends on the dataset size and computing resources. K can be 5, 10, or 20. For example, for a dataset containing 1000 legal consultation texts, if K = 10 is selected, each subset will contain 100 samples. During the division process, it is necessary to maintain the distribution of different intent categories in each subset similar to that of the original dataset, that is, perform stratified sampling, to avoid biases caused by too many or too few samples of specific categories in certain subsets. For each fold of validation, select the i-th subset (i ranges from 1 to K) as the validation set, and the remaining K - 1 subsets are combined as the training set. For example, in the first training, the first subset is used as the validation set, and the second to the K-th subsets are used as the training set; in the second training, the second subset is used as the validation set, and the first and the third to the K-th subsets are used as the training set, and so on.

[0138] This rotation method ensures that each sample has the opportunity to be a training sample and a validation sample, making full use of the limited data resources. In each fold of training, the model parameters are either re-initialized or loaded from a pre-trained model, and then trained on the training set of the current fold and evaluated on the corresponding validation set. This method can not only maximize the use of small-sample data for model training, but also obtain a more stable and reliable model performance evaluation through multiple trainings, reducing the randomness impact that may be brought by a single training-validation division. For professional field tasks such as legal text classification, it can provide a more accurate estimate of the model's generalization ability under limited labeled data conditions, and help identify whether there are overfitting or underfitting problems in the model.

[0139] Calculate the average accuracy rate of the K trainings. This step is to perform statistical analysis on the results of K-fold cross-validation to obtain a comprehensive evaluation index of the model performance. In each fold of validation, the prediction results of the model on the validation set are compared with the true labels to calculate the accuracy rate. The accuracy rate is the most intuitive evaluation index in classification tasks, defined as the ratio of the number of correctly classified samples to the total number of samples. For multi-classification problems, the accuracy rate can be simplified as the number of correctly classified samples divided by the total number of samples. For example, if the validation set contains 100 legal consultation texts and the model correctly identifies the intent categories of 85 texts, the accuracy rate of this fold is 85%.

[0140] After K training sessions are completed, K accuracy values will be obtained, corresponding to the results of each fold of validation respectively. The average accuracy is the arithmetic mean of these K accuracy values.

[0141] For example, for 10-fold cross-validation, if the accuracies of the 10 validations are 82%, 85%, 83%, 84%, 86%, 83%, 85%, 84%, 82%, and 86% respectively, then the average accuracy is 84%.

[0142] When the average accuracy reaches the preset accuracy threshold, stop training and save the model parameters of the language awareness classification model, aiming to ensure that the deployed model meets the expected performance requirements. The preset accuracy threshold is a performance standard determined in advance according to specific application scenarios and business requirements, and it represents the minimum accuracy level that the model must reach to be considered acceptable. For example, for tasks with high accuracy requirements such as legal consultation intention classification, a relatively high threshold may be set, such as 85% or 90%; while for some application scenarios with higher fault tolerance, the threshold may be relatively low.

[0143] In actual operation, when the calculated average accuracy of K-fold cross-validation reaches or exceeds the preset threshold, it indicates that the model already has sufficient performance level and further training and tuning processes can be stopped. At this time, the model parameters need to be saved, including network weights, bias terms, and other learnable parameters, for subsequent deployment and application. Model parameters are usually saved as binary files or checkpoint files in a specific format, which contain the complete state of the model and can be reloaded when needed to restore the model. When saving the model, relevant metadata is also recorded, such as model architecture, hyperparameter settings, training data characteristics, evaluation metrics, etc. If the average accuracy fails to reach the preset threshold, the model needs to be further optimized.

[0144] In this embodiment, by dividing the limited labeled data into K equal-sized subsets and performing rotation training, the information value of each sample is fully utilized, overcoming the randomness and bias that may be brought by single training-validation division in the small-sample scenario. K-fold cross-validation not only provides a more reliable model performance evaluation but also can effectively identify overfitting or underfitting problems, which is especially suitable for classification tasks in professional fields such as legal texts. The calculation of the average accuracy provides a comprehensive and objective metric for model performance, reflecting the generalization ability of the model on different data partitions, and the stopping criterion based on the preset accuracy threshold ensures that the finally deployed model meets the performance standards of business requirements. In summary, a good balance can be achieved between computational efficiency and evaluation reliability, avoiding unnecessary waste of computing resources and ensuring model quality. It is not only applicable to the initial model training but also suitable for model update and iterative optimization, providing reliable technical support for the continuous improvement of intelligent legal consultation systems.

[0145] The embodiments of the present application further provide an electronic device, including:

[0146] A memory configured to store instructions; and

[0147] A processor configured to call instructions from the memory and capable of implementing the above-mentioned few-shot language awareness classification method when executing the instructions.

[0148] The embodiments of the present application further provide a few-shot language awareness classification system, including:

[0149] The above-mentioned electronic device.

[0150] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0151] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of flows and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0152] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in one Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0153] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide for implementing the functions in the processFigure 1 one or more processes and / or blocks Figure 1 steps of functions specified in one or more blocks

[0154] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0155] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of computer-readable media.

[0156] Computer-readable media includes permanent and non-permanent, removable and non-removable media and can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0157] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed or elements inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0158] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A small sample language awareness classification method, characterized in that: Applied to an electronic device, the electronic device stores a language awareness classification model constructed based on a multi-head attention mechanism, the method comprising: Receive a small sample supervised dataset uploaded by a user, where the small sample supervised dataset includes annotated dialogue texts and corresponding intent categories; Preprocessing the small sample supervised data set to obtain preprocessed training data; Based on the preprocessed training data, the language awareness classification model is trained using a dynamic mask training strategy, wherein the dynamic mask training strategy is used to randomly mask some words in the input text and control the model to learn semantic relations from the context; During the process of training the language awareness classification model, the accuracy of the language awareness classification model is evaluated in real time, and when the accuracy of the language awareness classification model reaches a preset threshold, the language awareness classification model is deployed to the electronic device, so that the electronic device calls the language awareness classification model to perform intent classification on the acquired real-time conversation text to obtain an intent classification result.

2. The method according to claim 1, characterized in that The electronic device is deployed with an intent classification platform, which includes a data processing layer, a model training layer and an application layer. The model training layer includes the language awareness classification model, and the application layer is connected to at least one database.

3. The method according to claim 2, characterized in that The deploying the language awareness classification model to the electronic device includes: The language awareness classification model is deployed to the application layer of the electronic device, wherein the application layer is used to receive real-time conversation text, and call the language awareness classification model to perform intent classification on the real-time conversation text to obtain intent classification results, and according to the intent classification results, the corresponding reply template is determined from a preset reply template library, and guiding reply content is generated, and the guiding reply content is displayed through the front-end interface of the electronic device to achieve conversation guidance.

4. The method according to claim 1, characterized in that The preprocessing of the small sample supervised data set includes: Performing text cleaning on the dialogue text; Use the pre-trained word segmenter to segment the cleaned conversation text; The conversation text after word segmentation is converted into word vector representation to obtain preprocessed training data.

5. The method according to claim 3, characterized in that: The application layer calls the language awareness classification model to perform intent classification on the real-time conversation text to obtain an intent classification result, including: The application layer preprocesses the real-time conversation text to obtain a preprocessed conversation text; Inputting the preprocessed dialogue text into the language awareness classification model to obtain a probability distribution of intent categories; The intent category with the highest probability is taken as the intent classification result, wherein when the intent classification result is lower than a preset confidence threshold, the intent classification result is marked as an uncertain category.

6. The method according to claim 1, characterized in that The language awareness classification model constructed based on the multi-head attention mechanism is the language awareness classification model based on the Transformer architecture, and the language awareness classification model includes an encoder and a classification head; The encoder includes a multi-layer multi-head self-attention layer and a feedforward neural network layer, which is used to extract the contextual representation of the real-time conversation text; the classification head includes a fully connected layer and a softmax layer, which is used to map the contextual representation to the probability distribution of the intent category.

7. The method according to claim 6, characterized in that The method of training the language awareness classification model based on the preprocessed training data using a dynamic mask training strategy includes: During the training process, the pre-processed training data is randomly masked by a preset number of word units, so that the language awareness classification model predicts the masked word units and the correct intent category to obtain a prediction result; A cross entropy loss function is used to calculate the loss between the prediction result and the true label, and the model parameters of the language awareness classification model are updated through back propagation according to the loss.

8. The method according to claim 1, characterized in that After deploying the language awareness classification model to the electronic device, the method further includes: Dividing the small sample supervised data set into K subsets of equal size to train the language awareness classification model K times, wherein a different subset is selected as a validation set for each training, and the remaining subsets are used as training sets, and K is an integer greater than or equal to 2; Calculate the average accuracy of K training times; When the average accuracy reaches a preset accuracy threshold, the training is stopped and the model parameters of the language awareness classification model are saved.

9. An electronic device, characterized in that: include: a memory configured to store instructions; as well as A processor is configured to call the instructions from the memory and implement the small sample language awareness classification method according to any one of claims 1 to 8 when executing the instructions.

10. A small sample language awareness classification system, characterized in that include: An electronic device according to claim 9.

Citation Information

Patent Citations

  • Voice robot call method and device adopting high-generalization multi-task intention recognition

    CN115662431A

  • Real estate industry dialogue intention recognition method and device

    CN116050422A

  • Language model training method and device, equipment and medium

    CN117743516A

  • Electricity utilization inspection intelligent voice interaction system and method based on large language model

    CN118364046A

Cited By

  • Training method of thinking type recognition model and bad thinking recognition method

    CN120687918A

  • Automobile field article automatic classification method based on large model

    CN121167397A

  • Generative large language model-based intention recognition method and system

    CN121833953A

  • Automatic plug-in generation method and device based on AI assistance

    CN122152292A