Data classification and grading method and system based on large model
Through the data classification and grading method based on the big model, the converging government text classification model is debugged, and the cosine similarity adjustment and distance value correction is used to solve the problem of low accuracy of government text classification in the traditional method, achieving more efficient and accurate government text classification.
Patent Information
- Application Number
- CN202510188104.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-20
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2045-02-20
AI Technical Summary
Traditional data classification methods face problems such as complex content, blurred classification boundaries and category imbalance when processing government texts, resulting in low classification accuracy and inability to meet the needs of government work for accurate information classification.
The data classification and grading method based on the big model is adopted, and the converging target government text classification model is debugged, and the cosine similarity between the text representation array and the centroid array representation is adjusted to determine the value of the distance generation, and the model parameter variables are corrected to improve the classification accuracy.
It improves the classification accuracy of government affairs texts, enhances the model's ability to handle complex and fuzzy classifications, improves the problem of category imbalance, and meets the precise classification needs of government affairs information.
Smart Images

Figure CN119669475B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of text processing and machine learning technology, and in particular to a data classification and grading method and system based on a large model. Background Art
[0002] In the field of government affairs, the classification of government text data is of vital importance. With the rapid development of government informationization, the amount of government text data has increased dramatically, covering various policies and regulations, government service information, public feedback and other contents. Traditional data classification methods face many challenges in processing government texts. On the one hand, the content of government texts is complex and diverse, the classification boundaries are often unclear, and there are many fuzzy areas between different classifications, which makes it difficult for traditional methods to accurately classify government texts. For example, some comprehensive texts involving multiple government service sections are difficult for traditional methods to accurately determine their categories. On the other hand, there is a problem of category imbalance in government text data. The number of training texts for some classifications is large, while the number of training texts for some small number of text instance classifications is sparse. This makes the traditional method inefficient in learning when processing a small number of text instance classifications, and it is difficult to accurately identify the characteristics of these classifications, resulting in low classification accuracy. These problems seriously affect the management and utilization efficiency of government text data, and cannot meet the needs of government work for accurate information classification. Therefore, a more efficient and accurate data classification method is urgently needed to solve the above problems. Summary of the invention
[0003] In view of this, the embodiment of the present invention at least provides a data classification and grading method and system based on a large model. The technical solution of the present invention is implemented as follows:
[0004] In the first aspect, an embodiment of the present invention provides a data classification and grading method based on a large model, the method comprising: obtaining a target government text to be classified, and calling a debugged and converged target government text classification model; loading the target government text into the target government text classification model to perform content classification on the target government text based on the target government text classification model to obtain a target text classification result. The target government text classification model is obtained by debugging the basic government text classification model, and the debugging process includes the following steps: extracting the text representation array of the government training text in the government training text library based on the basic government text classification model; obtaining the first centroid array representation of the prior true classification and the second centroid array representation of the remaining classifications in the connection influencing variables of the fully connected network of the basic government text classification model, wherein the prior true classification is the classification corresponding to the government training text; adjusting the first cosine similarity between the text representation array and the second centroid array representation to obtain a target cosine similarity greater than the first cosine similarity; determining a distance cost value based on the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity; the second cosine similarity is the cosine similarity between the text representation array and the first centroid array representation; correcting the model parameters of the basic government text classification model based on the distance cost value to obtain the target government text classification model.
[0005] On the second side, the present invention provides a computer system, comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, and the processor implements the steps in the above method when executing the program.
[0006] The present invention provides a data classification and grading method and system based on a large model, which adopts a government text classification model to classify government texts that need to be classified. The government text classification model is obtained by debugging a basic government text classification model. During the debugging process, the basic government text classification model extracts the text representation array of the government training text, obtains the first centroid array representation of the prior real classification and the second centroid array representation of the remaining classification, increases the first cosine similarity between the text representation array and the second centroid array representation, obtains the target cosine similarity, determines the distance cost value according to the target cosine similarity, the second cosine similarity between the text representation array and the first centroid array representation, the first centroid array representation and the second centroid array representation, and uses the distance cost value to correct the model parameters of the basic government text classification model. Because the remaining classifications are not the actual classifications corresponding to the government training texts, increasing the first cosine similarity can expand the correlation between the text representation array and the second centroid array representation, and correcting the model parameters of the basic government text classification model by the distance proxy value determined according to the target cosine similarity, making the classification of the basic government text classification model more difficult. This can help improve the training quality of the basic government text classification model for training texts with unclear classification boundaries, difficult classification, and complex training texts, increase the classification effect of the target government text classification model obtained after the corrected parameters, and improve the classification accuracy of the target government text classification model for government texts. BRIEF DESCRIPTION OF THE DRAWINGS
[0007] Figure 1 A schematic diagram of the implementation flow of a large model-based data classification and grading method provided in an embodiment of the present invention.
[0008] Figure 2 A hardware entity schematic diagram of a computer system provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0009] The embodiment of the present invention provides a data classification and grading method based on a large model, which can be executed by a processor of a computer system. The computer system can refer to a device with data processing capabilities such as a server, a laptop, a tablet computer, a desktop computer, a mobile device, etc.
[0010] Figure 1 A schematic diagram of the implementation flow of a data classification and grading method based on a large model provided in an embodiment of the present invention is shown in FIG. Figure 1 As shown, the method includes:
[0011] Step S100: Obtain the target government text to be classified, and call the debugged and converged target government text classification model.
[0012] In step S100, the computer system can extract the target government text to be classified from the government database or file storage system through the data interface. These texts may include policy documents, public service records, citizen consultation records and other types, such as a government service work order text about "citizen medical insurance reimbursement process consultation". In the data acquisition stage, the computer system will perform format verification and encoding conversion to ensure that the text data meets the model input specification, such as converting unstructured text into a UTF-8 encoded string and removing special characters and redundant spaces. Subsequently, the system loads the pre-trained target government text classification model, which is obtained by parameter debugging and convergence training of the basic government text classification model, and its core architecture includes, for example, a text embedding coding network and a fully connected classification network. The text embedding coding network can use a variant of a pre-trained language model such as BERT, and map the input text to a high-dimensional semantic vector through a multi-layer Transformer structure; the fully connected network is responsible for mapping the semantic vector to the classification space, and its weight parameters are optimized by the back propagation algorithm during the debugging process. The model call process involves the dynamic allocation of computing resources, such as loading model parameters into video memory in a GPU cluster and establishing an inference service interface.
[0013] During the implementation process, the computer system performs standardized preprocessing on the input text, including but not limited to word segmentation (using word segmentation tools such as Jieba), stop word filtering (based on a stop word list customized for the government field), and length truncation (unifying the text length to the maximum supported length of the model, such as 512 tokens). Taking the classification task of the government service section as an example, when the system receives the "List of Materials for Renewal of Enterprise Establishment License", it forms a token sequence that meets the model input after preprocessing, generates a 768-dimensional text representation vector through the text embedding network, and finally outputs the classification probability distribution through the fully connected network. To ensure the efficiency of model calls, the computer system adopts a batch inference mechanism to process multiple texts in parallel through matrix operations. For example, 100 government texts are spliced into batch tensors for input into the model, which significantly improves throughput.
[0014] Step S200: Load the target government text into the target government text classification model to perform content classification on the target government text based on the target government text classification model to obtain a target text classification result.
[0015] In step S200, the computer system loads the pre-processed government text data into the model calculation flow, performs end-to-end semantic parsing and feature mapping, and finally outputs the structured classification results. The implementation of this step involves multiple technical links such as text representation generation, feature space mapping, and classification boundary calculation. Taking the classification task of the government service section as an example, when the system processes the text of "List of Materials for Renewal of Enterprise Establishment License", it first maps the text sequence into a 768-dimensional standardized vector representation through the text embedding encoding network. This process is based on the pre-trained multi-layer Transformer architecture. The specific calculation can be expressed as: ;in is the input word embedding matrix, MultiHeadAttention is the multi-head self-attention mechanism, and FFN is the feedforward neural network layer. After the text representation vector is generated, the computer system calls the weight matrix of the fully connected network (d is the embedding dimension, k is the number of classification categories) for linear transformation, the calculation formula is: ; Where b is the bias term and s is the classification score vector. In the debugged model, the parameters of the fully connected layer have been optimized through the distance cost value, so that the distribution of the centroid vectors of different categories in the feature space meets the maximum interval constraint. Taking the sentiment polarity classification task as an example, when processing the text "Citizens complain about untimely road maintenance", the system calculates the cosine similarity between the text representation vector and the centroid c- of the "negative emotion" class as 0.92, and the similarity with the centroid c+ of the "positive emotion" class as 0.15. According to the similarity adjustment factor introduced in the debugging stage = 0.3, the actual classification decision boundary is adjusted to: ;in is the temperature hyperparameter, which is used to control the smoothness of the probability distribution, h is the text representation vector, c i and c j The centroid vector represents different sentiment categories. The computer system achieves flexible control of classification boundaries by dynamically adjusting the cosine similarity threshold. For example, in the classification of government service sections, when the adjusted similarity between the target text and the "social security" category exceeds the preset threshold, =0.85, the system will trigger a high confidence classification mark. For multi-label classification scenarios, the system uses the sigmoid activation function and binary cross entropy loss, and the calculation formula is: ;p i represents the predicted probability of the i-th label, w i is the weight associated with the ith label and sets the dynamic threshold , where p represents the predicted probability distribution of each label by the model, is the entropy sensitivity coefficient, which is used to balance the classification accuracy and recall rate. In the model inference stage, the computer system uses a hierarchical caching mechanism to accelerate the calculation. For example, the embedding vectors of high-frequency government terms are pre-loaded into the GPU memory to reduce the data handling overhead during real-time calculations. Taking the processing of the "Urban and Rural Residents Medical Insurance Participation Process Consultation" text as an example, the system first splits it into a Token sequence through the word segmentation module, converts it into a 512-dimensional vector through the embedding layer, and then extracts contextual semantic features through a 12-layer Transformer encoder, and finally outputs the classification probability distribution through the fully connected layer. In this process, the inter-class distance constraint introduced in the debugging stage plays a key role: if the text actually belongs to the "medical insurance" category, the cosine similarity between its representation vector and the target class centroid is forced to be increased to , while suppressing the similarity with other class centroids to , where c t represents the centroid vector of the target category, c o represents the centroid vectors of other categories, The similarity adjustment amount determined during the debugging phase. The computer system implements an anomaly detection mechanism by quantitatively analyzing the relationship between classification confidence and feature space distribution. For example, when the Mahalanobis distance of the text representation vector deviates from the target class centroid by more than three times the standard deviation, the manual review process is triggered. In real-time service scenarios, the system uses an asynchronous pipeline architecture to distribute text preprocessing, model reasoning, and result post-processing on different computing nodes. For example, Apache Kafka queues are used to buffer input requests, and TensorRT is used to optimize the model calculation graph to achieve millisecond-level response. In response to the needs of long text classification, the system implements a hierarchical processing strategy: first, paragraph-level features are extracted through the BiLSTM network, and then the attention mechanism is used to aggregate the global representation. The calculation formula is:
[0016] ;
[0017] Where q is the learnable query vector, W a is the attention weight matrix, n is the number of paragraph-level features, and h i is the i-th paragraph-level feature, α i is the attention weight. At the quality monitoring level, the computer system continuously collects the difference data between the classification results and the manual annotations, and updates the model parameters through the online learning module. Specifically, the elastic weight consolidation algorithm is used to prevent catastrophic forgetting. The update formula is:
[0018] ; where F i are the diagonal elements of the Fisher information matrix, is the regularization strength, are the current model parameters, The model parameters at the end of the old task learning. For low-resource classification scenarios (such as emerging government affairs categories), the system enables a small amount of text instance learning mode, calculates the centroid vector of the support set training text instance based on the prototypical network, and the classification decision is based on the distance between the query training text instance and various prototypes:
[0019] ;in, It represents the probability of belonging to category c given the input feature vector h. d(·) is the Euclidean distance metric, p c is the prototype vector of category c is the variable indexed by category, For Category The computer system improves classification robustness by deploying a multi-model integration strategy, such as performing weighted voting on the prediction results of the BERT, RoBERTa, and ALBERT models, and the weights are dynamically adjusted according to the F1-score of each model on the validation set. In the output stage, the system not only returns the classification label, but also generates an explainable report, including visualization of key decision features (such as highlighting text fragments that affect classification through the Grad-CAM algorithm) and confidence interval estimation (using Bootstrap sampling to calculate 95% confidence intervals). For example, when processing the text of "query on the progress of renovation of old communities", the system outputs the classification label of "municipal construction" and identifies the influence weights of keywords such as "renovation progress" and "community age" to assist manual review.
[0020] In the embodiment of the present invention, the target government text classification model is obtained by debugging the basic government text classification model, and the debugging process includes the following steps:
[0021] Step S10: extracting a text representation array of the government training text in the government training text library based on the basic government text classification model.
[0022] When the computer system performs this step, the government training text library is a collection of a large number of government training texts with different themes, formats and sources, such as government announcements, policy documents, administrative notices, etc. The government training text library can be stored in a storage medium such as a database or a file system for access and use by the computer system at any time. It should be noted that the premise of collecting the above training texts is within the scope permitted by laws and regulations.
[0023] The basic government text classification model is a pre-built model, for example, composed of a text embedding coding network and a fully connected network. The function of the text embedding coding network is to convert the input government training text into a low-dimensional vector representation so that the computer system can process and analyze it. The fully connected network is used to classify the vector output by the text embedding coding network and determine the category to which the text belongs. The computer system reads the government training text from the government training text library and then inputs the text into the text embedding coding network of the basic government text classification model. The text embedding coding network processes the input text and converts it into a text representation array. The text representation array is a numerical array that can reflect the semantic information and features of the government training text.
[0024] For example, suppose there is a government announcement about urban traffic management in the government training text library, which reads "In order to ease urban traffic congestion, the municipal government has decided to implement traffic control measures on some sections of the road, from 7am to 9am and from 5pm to 7pm every day". The computer system inputs the text into the text embedding coding network of the basic government text classification model, which performs word segmentation and word vector representation on the text. Assume that the word vector representation method is used to convert each word into a vector of fixed length, and then these vectors are concatenated or aggregated to obtain the vector representation of the text.
[0025] In practical applications, computer systems can use a variety of technical means to implement text embedding coding networks. Feasible methods include bag-of-words model, TF-IDF model, Word2Vec, GloVe, BERT, etc. After the computer system inputs the government training text into the text embedding coding network, it obtains the text representation array of each government training text. These text representation arrays will be used as input for subsequent steps to calculate distance cost values, correct model parameters of the basic government text classification model, etc. By extracting the text representation array of the government training text, the computer system can convert the government training text into a numerical form that the computer can process, providing a basis for subsequent model debugging and classification tasks.
[0026] In actual operation, the computer system can also standardize the extracted text representation array so that the norm of the text representation array is equal to 1. The purpose of this is to eliminate the scale differences between different text representation arrays and improve the training effect of the model. Standardization can use a feasible normalization method, such as L2 normalization. For a text representation array x, the calculation formula of its L2 normalized array x' is: x'=x / ||x||, where ||x|| represents the L2 norm of the array x, that is, the square root of the sum of the squares of each element in the array.
[0027] Step S20: Obtain the first centroid array representation of the prior true classification and the second centroid array representation of the remaining classifications from the connection influencing variables of the fully connected network of the basic government text classification model, wherein the prior true classification is the classification corresponding to the government training text.
[0028] The fully connected network of the basic government text classification model consists of multiple neurons and weight parameters connecting these neurons. These weight parameters are connection influencing variables, which determine how the input features are combined and transformed to produce output classification results. Each government training text in the government training text library has its corresponding prior true classification. For example, government texts may involve different categories such as policies and regulations, people's livelihood services, and urban construction, and the prior true classification is the category to which the text actually belongs. The first centroid array representation is the array representation of the centroid corresponding to the prior true classification. The centroid is the center position of all training text instances in a classification, which can be mathematically understood as the average value of all training text instance vectors in the classification. Taking the policy and regulation category government text as an example, the computer system will summarize the text representation arrays obtained after all government training texts belonging to the policy and regulation category are processed by the text embedding coding network, and then calculate the average value of these text representation arrays. The result is the first centroid array representation of the policy and regulation category. Suppose there are three government training texts of policies and regulations. After the text embedding coding network, the text representation arrays obtained are [1, 2, 3], [2, 3, 4], and [3, 4, 5] respectively. Then the first centroid array representation of the policy and regulations category is [(1+2+3) / 3, (2+3+4) / 3, (3+4+5) / 3]=[2, 3, 4]. The second centroid array representation is the array representation of the centroids of the remaining categories. The remaining categories refer to all categories other than the a priori true category corresponding to the current government training text. For example, if the a priori true category of the current government training text is people's livelihood services, then other categories such as policies and regulations, urban construction, etc. belong to the remaining categories. The computer system will calculate the centroid of each remaining category separately, that is, average the text representation arrays of all government training texts under the category to obtain the corresponding second centroid array representation. Assume that the policy and regulation category has the three government training texts mentioned above, and their text representation arrays are [1, 2, 3], [2, 3, 4], and [3, 4, 5] respectively, and the urban construction category has two government training texts, and their text representation arrays are [4, 5, 6] and [5, 6, 7] respectively. Then the second centroid array of the policy and regulation category is [2, 3, 4], and the second centroid array of the urban construction category is [(4+5) / 2, (5+6) / 2, (6+7) / 2]=[4.5, 5.5, 6.5].
[0029] In actual operation, the computer system can use the following technical means to obtain the first centroid array representation and the second centroid array representation. First, the computer system classifies and stores the text representation array according to the a priori true classification of the government training text, and puts the text representation arrays belonging to the same classification together. For example, this is achieved by establishing a classification dictionary, the key of the dictionary is the classification name, and the value is a list of all text representation arrays under the classification. Then, for each classification, the computer system calculates the average value of all text representation arrays under the classification. Suppose there are n text representation arrays under a certain classification, the dimension of each text representation array is m, and the i-th text representation array is xi=[xi1, xi2, …, xim], then the centroid array representation c=[c1, c2, …, cm] of the classification, where cj=(∑i=1n xij) / n, j=1, 2, …, m.
[0030] The first centroid array representation and the second centroid array representation reflect the central features of each classification. In the subsequent steps, the computer system will use these centroid array representations to calculate the cosine similarity between the text representation array and the centroid array representation, and then determine the distance cost value, and finally correct the model parameters of the basic government text classification model. By continuously adjusting the model parameters, the basic government text classification model can better distinguish government texts of different categories, and improve the classification accuracy and generalization ability of the model. For example, when calculating the distance cost value later, the similarity between the text representation array and the first centroid array representation and the second centroid array representation will be compared. If the text representation array has a higher similarity with the first centroid array representation and a lower similarity with the second centroid array representation, it means that the text is more likely to belong to the a priori true classification, otherwise it may be classified incorrectly. In this way, the computer system can adjust the model according to the distance cost value, so that the model is more accurate in classification.
[0031] Step S30: adjusting the first cosine similarity between the text representation array and the second centroid array representation to obtain a target cosine similarity greater than the first cosine similarity.
[0032] When the computer system executes this step, the text representation array is a numerical array obtained after the government training text is processed by the text embedding encoding network of the basic government text classification model, which reflects the semantic information and characteristics of the government training text. The second centroid array representation is an array representation of the centroids of other categories except the prior true categories corresponding to the current government training text, and the first cosine similarity is an indicator to measure the degree of directional similarity between the text representation array and the second centroid array representation.
[0033] The calculation formula for cosine similarity is: ,in Represents an array of text representations, represents the second centroid array representation, is the dot product of two arrays, are the norms of the two arrays respectively.
[0034] The purpose of the computer system adjusting the first cosine similarity is to expand the correlation between the text representation array and the second centroid array representation, making it more difficult for the basic government text classification model to classify, thereby improving the training quality of the model for training texts with unclear classification boundaries, difficult classification, and complex training. The result of the adjustment is to obtain a target cosine similarity that is greater than the first cosine similarity, which means that the directions between the text representation array and the second centroid array representation are more similar.
[0035] In order to achieve this adjustment, the computer system can use a variety of technical means. One feasible method is to modify the first cosine similarity by setting a similarity adjustment factor. The similarity adjustment factor is a pre-set value, which can be adjusted according to the actual situation to achieve the best adjustment effect. Assume that the similarity adjustment factor is adjusted to achieve the best adjustment effect. Assume that the similarity adjustment factor , then the target cosine similarity ,in is the first cosine similarity. For example, if the first cosine similarity is 0.98, the similarity adjustment factor =0.02, then the target cosine similarity =0.98+0.02=1.
[0036] In practical applications, the computer system can select an appropriate similarity adjustment factor based on the specific data set and model performance. The computer system can also consider using a dynamic adjustment method to determine the similarity adjustment factor. For example, the size of the similarity adjustment factor can be dynamically adjusted based on factors such as the classification difficulty and number of categories of the government training text. For government training texts that are more difficult to classify, the similarity adjustment factor can be appropriately increased to enhance the model's learning ability for these texts; for situations where there are a large number of categories, the similarity adjustment factor can be dynamically adjusted based on the degree of difference between different categories, so that the model can better distinguish government texts of different categories.
[0037] By increasing the correlation between the text representation array and the second centroid array representation, the computer system can make the model pay more attention to the subtle differences between different classifications and improve the model's ability to classify complex government texts. In subsequent steps, the target cosine similarity will be used to determine the distance cost, and the distance cost will be used to correct the model parameters of the basic government text classification model, thereby continuously optimizing the performance of the model. For example, when the target cosine similarity increases, the calculated distance cost will also change accordingly. The model will adjust the model parameters according to the new distance cost, making the model more accurate in classification.
[0038] Step S40: determining a distance cost value according to the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity; the second cosine similarity is the cosine similarity between the text representation array and the first centroid array representation.
[0039] When the computer system executes this step, the target cosine similarity is obtained by adjusting the first cosine similarity between the text representation array and the second centroid array representation, and its value is greater than the first cosine similarity, with the purpose of expanding the correlation between the text representation array and the second centroid array representation; the text representation array is a numerical array obtained after the government training text is processed by the text embedding encoding network of the basic government text classification model, reflecting the semantic information and characteristics of the government training text; the first centroid array representation is the array representation of the centroid of the prior true classification; the second centroid array representation is the array representation of the centroid of the remaining classifications; the second cosine similarity is the cosine similarity between the text representation array and the first centroid array representation, which measures the directional similarity between the text representation array and the centroid of the prior true classification.
[0040] The distance cost value is used to measure the difference between the text representation array and the prior true classification and the remaining classifications. It can reflect the error of the model when classifying government training texts. The computer system determines the distance cost value by comprehensively considering the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity, which can more comprehensively evaluate the classification effect of the model. For example, if the target cosine similarity is large, it means that the correlation between the text representation array and the second centroid array representation is high. At this time, the text is more likely to be misclassified into the remaining classifications, and the corresponding distance cost value will be large; if the second cosine similarity is large, it means that the direction of the text representation array and the first centroid array representation is relatively similar, and the text is more in line with the prior true classification, and the distance cost value will be relatively small.
[0041] In order to determine the distance cost value, the computer system can adopt a calculation method based on cosine similarity and array norm. First, the computer system clarifies that the text representation array, the first centroid array representation and the second centroid array representation are, for example, standardized array representations, and their norms are all equal to 1, which helps to eliminate the scale differences between different arrays and improve the accuracy of the calculation. The computer system will determine a correlation index between the text representation array and the second centroid array representation based on the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation; at the same time, based on the second cosine similarity, the norm of the text representation array and the norm of the first centroid array representation, another correlation index between the text representation array and the first centroid array representation is determined. Then, by further calculating and processing these two correlation indexes, the distance cost value is obtained.
[0042] Assume that the text representation array is , the first centroid array is represented as , the second centroid array is represented as , the target cosine similarity is , the second cosine similarity is Since the array is a normalized array, its norm The computer system can first calculate the dot product between the text representation array and the second centroid array representation, and the text representation array and the first centroid array representation based on the target cosine similarity and the second cosine similarity, and then use the natural constant e as the base value and the dot product as the power to obtain the corresponding correlation index. Assuming that the dot product calculated based on the target cosine similarity is dot1, and the dot product calculated based on the second cosine similarity is dot2, the correlation indexes are respectively Finally, by Perform specific operations, such as calculating the ratio, logarithmic operation, etc., to obtain the distance cost value D. For example, if there is only one remaining category, the distance cost value D can be expressed as .
[0043] In practical applications, the computer system selects an appropriate calculation method based on the characteristics of the government training text and the performance requirements of the model. When the number of other categories is multiple, the computer system calculates the correlation index between each second centroid array representation and the text representation array, and then combines these indicators to determine the distance cost value. The computer system can also consider introducing an influencing variable of the prior true classification, which is inversely correlated with the ratio between the number of training texts corresponding to the prior true classification and the total amount of training texts corresponding to the government training text library. By considering the influencing variables to optimize the calculation of the distance cost value, improve the problem of category imbalance, and improve the learning efficiency of the model for classifying a small number of text instances.
[0044] After determining the distance cost value, the computer system will use the distance cost value to correct the model parameters of the basic government text classification model. By continuously adjusting the model parameters, the distance cost value is gradually reduced, so that the model can be more accurate in classifying government texts. For example, if the distance cost value is large, it means that the classification error of the model is large. The computer system will adjust the model parameters according to the size and direction of the distance cost value, so that the model is more inclined to classify the text into the prior true classification, reducing the situation of misclassification.
[0045] Step S50: Modify the model parameters of the basic government text classification model according to the distance cost value to obtain the target government text classification model.
[0046] The purpose of the computer system to correct the model parameters is to make the basic government text classification model more accurate in classifying government texts and reduce classification errors. Specifically, when the distance cost is large, it means that there is a large difference between the current classification result of the model and the true classification, and the model parameters need to be adjusted to reduce the distance cost; when the distance cost is small, it means that the classification effect of the model is good, but the classification accuracy can still be further improved by fine-tuning the model parameters.
[0047] In order to modify the model parameters according to the distance cost value, the computer system can use the gradient descent method, a feasible optimization algorithm. The basic idea of the gradient descent method is to update the model parameters along the negative gradient direction of the objective function (the objective function here can be understood as the distance cost value), thereby gradually reducing the value of the objective function. Assume that the model parameters are , the distance cost is , then the update formula of the gradient descent method is ,in is the learning rate, which is used to control the step size of each update, is the distance cost exist The gradient at .
[0048] The computer system will continue to repeat this update process until the distance cost converges to a smaller value or reaches a preset number of iterations.
[0049] When the computer system corrects the model parameters, it also considers the overfitting and underfitting problems of the model. In order to avoid overfitting, the computer system can use regularization methods, such as L1 regularization and L2 regularization. The distance cost value after adding the regularization term can be expressed as ,in is the regularization coefficient, is the regularization term.
[0050] After multiple iterations of updating the model parameters, the basic government text classification model obtained by the computer system becomes the target government text classification model. The target government text classification model includes the target text embedding encoding network and the target fully connected network. After adjusting their model parameters, they can classify government texts more accurately.
[0051] As an implementation manner, step S30, adjusting the first cosine similarity between the text representation array and the second centroid array representation to obtain a target cosine similarity greater than the first cosine similarity, includes:
[0052] Step S31: Obtaining the first cosine similarity between the text representation array and the second centroid array representation;
[0053] Step S32: Calculate the difference between the first cosine similarity and the first similarity adjustment factor to obtain a target cosine similarity greater than the first cosine similarity.
[0054] In step S31, the text representation array is a numerical array obtained after the government training text is processed by the text embedding coding network of the basic government text classification model, which reflects the semantic information and features of the government training text, and the second centroid array representation is an array representation of the centroids of other categories except the a priori true classification corresponding to the current government training text. The computer system can understand the similarity between the government training text and the centroids of other categories by calculating the cosine similarity between the text representation array and the second centroid array representation.
[0055] The computer system can calculate the first cosine similarity using the calculation formula of cosine similarity: ,in Represents an array of text representations, represents the second centroid array representation, is the dot product of two arrays, that is, the corresponding elements are multiplied and summed. are the norms of the two arrays. In practical applications, in order to eliminate the scale differences between different arrays, for example, the text representation array and the second centroid array representation are standardized so that their norms are equal to 1. At this time, the calculation formula of cosine similarity is simplified to .
[0056] Step S32 is to calculate the difference between the first cosine similarity and the first similarity adjustment factor to obtain a target cosine similarity greater than the first cosine similarity. The first similarity adjustment factor is a pre-set value, which is a parameter introduced by the computer system to expand the correlation between the text representation array and the second centroid array representation. By adjusting the first cosine similarity, the target cosine similarity is greater than the first cosine similarity, which means that the computer system will pay more attention to the similarity between the text representation array and the second centroid array representation in the subsequent model debugging process, thereby improving the training quality of the model for training texts with unclear classification boundaries and difficult and complex classification.
[0057] Assume the first cosine similarity is , the first similarity adjustment factor is , then the target cosine similarity The calculation formula is as follows. The negative sign here is because we want to get a value greater than the first cosine similarity, so in fact, we add a positive number to the first cosine similarity.
[0058] When determining the first similarity adjustment factor, the computer system needs to adjust it according to the specific data set and model performance. Through experimental methods, different values of the first similarity adjustment factor can be tried, the classification effect of the model can be observed, and the first similarity adjustment factor that makes the model classification accuracy the highest can be selected. It is also possible to consider using a method of dynamically adjusting the first similarity adjustment factor. For example, the size of the first similarity adjustment factor can be dynamically adjusted according to factors such as the classification difficulty and the number of categories of the government training text. For government training texts that are more difficult to classify, the first similarity adjustment factor can be appropriately increased to enhance the model's learning ability for these texts; for situations where there are a large number of categories, the first similarity adjustment factor can be dynamically adjusted according to the degree of difference between different categories, so that the model can better distinguish government texts of different categories.
[0059] As an implementation mode, step S40, determining the distance cost value according to the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity, includes:
[0060] Step S401: determining a first correlation coefficient between the text representation array and the second centroid array representation according to the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation;
[0061] Step S402: determining a second correlation coefficient between the text representation array and the first centroid array representation according to the second cosine similarity, the norm of the text representation array and the norm of the first centroid array representation;
[0062] Step S403: Determine a distance cost value according to the first correlation coefficient and the second correlation coefficient.
[0063] Step S401 is to determine the first correlation coefficient between the text representation array and the second centroid array representation based on the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation. The target cosine similarity is obtained after the computer system adjusts the first cosine similarity between the text representation array and the second centroid array representation. Its purpose is to expand the correlation between the text representation array and the second centroid array representation so that the model pays more attention to the subtle differences between different categories during classification. The text representation array is a numerical array obtained after the government training text is processed by the text embedding coding network of the basic government text classification model, which reflects the semantic information and features of the government training text. The second centroid array representation is an array representation of the centroids of other categories except the a priori true classification corresponding to the current government training text. The norm is a measurement method for measuring the length of a vector. In practical applications, for example, the text representation array and the second centroid array representation will be standardized so that their norms are both equal to 1.
[0064] The first correlation coefficient is used to indicate the degree of correlation between the text representation array and the second centroid array representation. The computer system calculates the first correlation coefficient through the target cosine similarity and the array norm, which can more accurately reflect the similarity between the two. The specific calculation method is to determine the first dot product of the text representation array and the second centroid array representation based on the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation, and then use the natural constant e as the base value and the first dot product as the power to obtain the first correlation coefficient. Since the array norm is 1, the first dot product is equal to the target cosine similarity at this time, so the calculation formula of the first correlation coefficient can be expressed as, where K1 represents the first correlation coefficient and represents the target cosine similarity.
[0065] Step S402 is to determine the second correlation coefficient between the text representation array and the first centroid array representation based on the second cosine similarity, the norm of the text representation array and the norm of the first centroid array representation. The second cosine similarity is the cosine similarity between the text representation array and the first centroid array representation, which measures the directional similarity between the text representation array and the a priori true classification centroid. The first centroid array representation is an array representation of the centroid of the a priori true classification, and its norm is 1 after normalization.
[0066] The second correlation coefficient is used to indicate the degree of correlation between the text representation array and the first centroid array representation. The computer system calculates the second correlation coefficient through the second cosine similarity and the array norm, which can reflect the degree of matching between the text representation array and the prior true classification. The specific calculation method is similar to the first correlation coefficient. Based on the second cosine similarity, the norm of the text representation array and the norm of the first centroid array representation, the second dot product of the text representation array and the first centroid array representation is determined, and then the second correlation coefficient is determined using the natural constant e as the base value and the second dot product as the power. Since the array norm is 1, the second dot product is equal to the second cosine similarity, so the calculation formula of the second correlation coefficient can be expressed as: , where K2 represents the second correlation coefficient, Represents the second cosine similarity.
[0067] The computer system can calculate the second correlation coefficient by using a technical means similar to step S401. The value of the second cosine similarity is stored in a numerical calculation library, and then an exponential function is used to calculate the result with the natural constant e as the base value and the second cosine similarity as the power to obtain the second correlation coefficient.
[0068] Step S403 is to determine the distance cost value based on the first correlation coefficient and the second correlation coefficient. The distance cost value is used to measure the difference between the text representation array and the prior true classification and the remaining classifications. It reflects the error of the model when classifying the government training text. The computer system determines the distance cost value by comprehensively considering the first correlation coefficient and the second correlation coefficient, which can more comprehensively evaluate the classification effect of the model.
[0069] When calculating the distance cost value, the computer system considers the proportional relationship between the first correlation coefficient and the second correlation coefficient. If the first correlation coefficient is large, it means that the correlation between the text representation array and the second centroid array representation is high, and the text is more likely to be misclassified into other categories, and the distance cost value will increase accordingly; if the second correlation coefficient is large, it means that the correlation between the text representation array and the first centroid array representation is high, and the text is more consistent with the prior true classification, and the distance cost value will be relatively small.
[0070] The distance cost value is calculated based on the logarithmic operation of the proportional relationship between the first correlation coefficient and the second correlation coefficient. Its purpose is to convert the proportional relationship into a numerical value that can more intuitively reflect the classification error of the model. The logarithmic operation can compress the proportional relationship so that the distance cost value is comparable in different situations. For example, when the proportional difference between the first correlation coefficient and the second correlation coefficient is large, the logarithmic operation can avoid the distance cost value being too large; when the proportional difference is small, the logarithmic operation can make the distance cost value more sensitive to reflect this difference.
[0071] The calculation formula of distance cost is: , where D represents the distance cost value, K1 represents the first correlation coefficient, and K2 represents the second correlation coefficient.
[0072] The determined distance cost is crucial for debugging the basic government text classification model. The computer system will correct the model parameters of the basic government text classification model based on the distance cost, making the model more accurate in classifying government texts. When the distance cost is large, it means that the classification error of the model is large. The computer system will adjust the model parameters according to the size and direction of the distance cost, making the model more inclined to classify the text into the prior true classification and reducing the misclassification.
[0073] The implementation of steps S401-S403 can help the basic government text classification model better handle government training texts with unclear classification boundaries and difficult and complex classification. By calculating the correlation coefficient and distance proxy value, the model can pay more attention to the subtle differences between different classifications and improve the classification ability of complex government texts. In the task of government text classification, many government texts may involve content in multiple fields at the same time. By accurately calculating the distance proxy value, the model can more accurately determine the category to which the text belongs.
[0074] As an implementation mode, the text representation array, the first centroid array representation, and the second centroid array representation are standardized array representations, and the norm of the text representation array, the norm of the first centroid array representation, and the norm of the second centroid array representation are all equal to 1. Then, the process of determining the first correlation coefficient and the second correlation coefficient may include the following steps:
[0075] Step S40A: Based on the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation, determine the first dot product of the text representation array and the second centroid array representation, take the natural constant as the base value and the first dot product as the power, and obtain the first correlation coefficient;
[0076] Step S40B: Based on the second cosine similarity, the norm of the text representation array and the norm represented by the first centroid array, determine the second dot product of the norm of the text representation array and the first centroid array, take the natural constant as the base value and the second dot product as the power, and determine the second correlation coefficient.
[0077] In some implementations, if there are multiple remaining categories, then step S403 determines the distance cost value according to the first correlation coefficient and the second correlation coefficient, including:
[0078] Step S4031: obtaining a plurality of first correlation coefficients;
[0079] Step S4032: Determine the sum of the first correlation coefficient and the second correlation coefficient, determine the ratio between the second correlation coefficient and the sum, perform logarithmic operation on the ratio, and obtain the distance cost value.
[0080] Step S40A is to determine the first dot product of the text representation array and the second centroid array representation based on the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation, and obtain the first correlation coefficient by taking the natural constant as the base value and the first dot product as the power. The target cosine similarity is obtained by adjusting the first cosine similarity between the text representation array and the second centroid array representation, and its purpose is to expand the correlation between the two so that the model pays more attention to the subtle differences between different classifications during classification. The text representation array is a numerical array obtained after the government training text is processed by the text embedding coding network of the basic government text classification model, which reflects the semantic information and features of the government training text. The second centroid array representation is an array representation of the centroids of other classifications except the a priori true classification corresponding to the current government training text. In practical applications, for example, the text representation array and the second centroid array representation will be standardized so that their norms are both equal to 1.
[0081] Since the array norms are all 1, the first dot product is equal to the target cosine similarity. The first correlation coefficient is used to indicate the degree of correlation between the text representation array and the second centroid array representation. The computer system obtains the first correlation coefficient by performing exponential operation with the natural constant e as the base value and the target cosine similarity as the power.
[0082] For example, suppose there is a government training text about social security policy in the government training text library. After being processed by the text embedding coding network, the text representation array obtained is , after normalization, its norm , the second centroid array of environmental protection classification is expressed as , its norm , the target cosine similarity calculated by the computer system , then the first correlation coefficient .
[0083] Step S40B is to determine the second dot product of the norm of the text representation array and the first centroid array representation based on the second cosine similarity, the norm of the text representation array and the norm of the first centroid array representation, and determine the second correlation coefficient with the natural constant as the base value and the second dot product as the power. The second cosine similarity is the cosine similarity between the text representation array and the first centroid array representation, which measures the directional similarity between the text representation array and the a priori true classification centroid. The first centroid array representation is the array representation of the centroid of the a priori true classification, and the norm is 1 after the same normalization process.
[0084] Since the array norm is 1, the second dot product is equal to the second cosine similarity. The second correlation coefficient is used to indicate the degree of correlation between the text representation array and the first centroid array representation. The computer system obtains the second correlation coefficient by performing exponential calculations based on the natural constant e as the base value and the second cosine similarity as the power. The calculation formula for the second correlation coefficient is: , where K2 represents the second correlation coefficient, Represents the second cosine similarity.
[0085] The computer system can calculate the second correlation coefficient by using a technical means similar to step S40A, using a numerical calculation library to store the value of the second cosine similarity, and then using an exponential function to calculate the result with the natural constant e as the base value and the second cosine similarity as the power to obtain the second correlation coefficient.
[0086] Step S4031 is to obtain multiple first correlation coefficients. In the actual government text classification task, the number of other categories is, for example, multiple, and each of the other categories has a corresponding second centroid array representation. The computer system will calculate the first correlation coefficient between the text representation array and the second centroid array representation according to the method of step S40A for each second centroid array representation. Therefore, multiple first correlation coefficients will eventually be obtained, and these first correlation coefficients reflect the degree of correlation between the text representation array and each of the other classification centroids.
[0087] The computer system can loop through each second centroid array representation, calculate and store the corresponding first correlation coefficients in sequence. These first correlation coefficients can be stored in a data structure such as a list or an array for subsequent processing.
[0088] Step S4032 is to determine the sum of the first correlation coefficient and the second correlation coefficient, determine the ratio between the second correlation coefficient and the sum, perform logarithmic operation on the ratio, and obtain a distance cost value. The purpose of this step is to comprehensively consider the degree of association between the text representation array and the prior true classification and each of the remaining classifications, and obtain a distance cost value that can accurately measure the model classification error by calculating the ratio and logarithmic operation.
[0089] First, the computer system adds all the first correlation coefficients, and then adds the second correlation coefficients to obtain the sum. Assume that there are n first correlation coefficients in total. , the second correlation coefficient is K2, then the addition result is Then, the ratio between the second correlation coefficient and the summation result is calculated. Finally, a logarithmic operation is performed on the ratio. Since the result of the logarithmic operation may be a negative number, in order to make the distance cost value a positive number, for example, a negative logarithm is taken, that is, the distance cost value D = -ln(P).
[0090] The computer system will modify the model parameters of the basic government text classification model based on the distance cost value, making the model more accurate in classifying government texts. When the distance cost value is large, it means that the classification error of the model is large. The computer system will adjust the model parameters according to the size and direction of the distance cost value, making the model more inclined to classify the text into the prior true classification and reducing the misclassification.
[0091] The implementation of steps S40A-S40B and steps S4031-S4032 can help the basic government text classification model better handle government training texts with unclear classification boundaries and difficult and complex classification. By calculating the correlation coefficient and distance proxy value, the model can pay more attention to the subtle differences between different classifications and improve the classification ability of complex government texts. In the task of government text classification, many government texts may involve content in multiple fields at the same time. By accurately calculating the distance proxy value, the model can more accurately determine the category to which the text belongs.
[0092] In some other embodiments, the method provided by the present invention may further include:
[0093] Step S33: Determine the influencing variables of the prior true classification according to the ratio between the number of training texts corresponding to the prior true classification and the total amount of training texts corresponding to the government training text library; the influencing variables are inversely correlated with the ratio.
[0094] Based on this, as another implementation of step S40, that is, determining the distance cost value according to the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity may include:
[0095] Step S41: determining a distance cost value according to the influencing variable, the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity.
[0096] Step S33 is to determine the influencing variable of the prior true classification based on the ratio between the number of training texts corresponding to the prior true classification and the total amount of training texts corresponding to the government training text library, and the influencing variable is inversely correlated with the above ratio. In the government text classification task, the number of training texts under different classifications often differs. Some classifications may have a large number of training texts (a large number of text instance classifications), while some classifications have fewer training texts (a small number of text instance classifications). This category imbalance will cause the model to be more inclined to learn the features of a large number of text instance classifications during the training process, while ignoring the features of a small number of text instance classifications, thereby affecting the classification accuracy of the model for a small number of text instance classifications. By determining the influencing variables of the prior true classification, the computer system can perform differentiated processing on different classifications and increase the model's attention to the classification of a small number of text instances.
[0097] The a priori true classification is the classification to which the government training text actually belongs. The number of training texts corresponding to the a priori true classification is the total number of training texts under this classification, and the total amount of training texts corresponding to the government training text library is the total number of training texts under all classifications. The computer system calculates the ratio between the two and then determines the influencing variables of the a priori true classification according to certain rules. For example, assuming that there are 1,000 training texts in the government training text library, of which 100 are in the social security policy classification, then the ratio of the number of training texts corresponding to the social security policy classification to the total amount of training texts is 100 / 1000=0.1. According to the principle that the influencing variable is inversely correlated with the ratio, when the ratio is small, the influencing variable will be larger; when the ratio is large, the influencing variable will be smaller. This means that the influencing variable of the classification of a small number of text instances is greater than the influencing variable of the classification of a large number of text instances, thereby giving more weight to the classification of a small number of text instances in subsequent model training.
[0098] In order to determine the influencing variables of the prior true classification, the computer system can use a variety of technical means. One feasible method is to first count the number of training texts corresponding to the prior true classification and the total amount of training texts corresponding to the government training text library, and then calculate the ratio between the two. Then, the ratio is converted into an influencing variable according to a preset rule. For example, a mapping function can be used to map the ratio to a suitable influencing variable value.
[0099] Step S41 is to determine the distance cost value based on the influencing variables, the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity. In the traditional distance cost value calculation, the problem of class imbalance is not taken into account, resulting in poor learning effect of the model for the classification of a small number of text instances. By introducing the influencing variables, the computer system can optimize the calculation of the distance cost value, so that the classification of a small number of text instances gets more attention in the model training, thereby improving the problem of class imbalance and improving the learning efficiency of the model for the classification of a small number of text instances.
[0100] The target cosine similarity is obtained by adjusting the first cosine similarity between the text representation array and the second centroid array representation, which reflects the degree of association between the text representation array and the remaining classification centroids; the second cosine similarity is the cosine similarity between the text representation array and the first centroid array representation, which reflects the degree of association between the text representation array and the prior true classification centroid. When calculating the distance cost value, the computer system will comprehensively consider these factors and influencing variables. For example, when the influencing variable is large, it means that the classification is a classification of a small number of text instances. The computer system will give it a greater weight when calculating the distance cost value, so that the model pays more attention to the training text of this classification.
[0101] The computer system can determine the distance cost value based on these factors in the following manner. First, determine the first correlation coefficient between the text representation array and the second centroid array representation according to the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation, and determine the second correlation coefficient between the text representation array and the first centroid array representation according to the second cosine similarity, the norm of the text representation array and the norm of the first centroid array representation. Then, these correlation coefficients are processed in combination with the influencing variables to finally obtain the distance cost value. Specifically, the influencing variable can be operated with a certain combination of the first correlation coefficient and the second correlation coefficient to obtain a new value, which is then used to calculate the distance cost value. For example, the influencing variable can be multiplied by the ratio of the first correlation coefficient to the second correlation coefficient, and then logarithmic operations and other operations can be performed to obtain the distance cost value.
[0102] The computer system can also further optimize these two steps. For example, in step S33, a method of dynamically adjusting the influencing variables can be adopted to adjust the size of the influencing variables in real time according to the training progress and classification performance of the model. In the early stage of model training, a large influencing variable can be given to the classification of a small number of text instances, so that the model can learn the characteristics of the classification of a small number of text instances more quickly; in the later stage of model training, the influencing variables can be appropriately reduced to avoid the model from overfitting the classification of a small number of text instances. In step S41, more factors can be introduced to calculate the distance cost value, such as considering the correlation between different classifications, etc., to further improve the classification accuracy and generalization ability of the model.
[0103] As an implementation mode, step S33, according to the ratio between the number of training texts corresponding to the priori true classification and the total amount of training texts corresponding to the government training text library, determines the influencing variables of the priori true classification, including:
[0104] Step S331: determining the ratio between the number of training texts corresponding to the a priori true classification and the total amount of training texts corresponding to the government training text library;
[0105] Step S332: determining a power correction coefficient based on the proportion; wherein the power correction coefficient is used to represent the existence proportion of the training texts of the remaining categories in the government training text library, the larger the influencing variable of the prior true classification is, the smaller the power correction coefficient is, the smaller the influencing variable of the prior true classification is, the larger the power correction coefficient is, and the preset basic value is a natural constant;
[0106] Step S333: determining the influencing variables of the a priori true classification according to the power correction coefficient and the preset basic value;
[0107] Specifically, determine the ratio between the number of training texts and the total amount of training texts, and use the difference between the constant 1 and the ratio as the power correction coefficient; determine the natural constant as the basic value, and use the power correction coefficient as the power to obtain a temporary variable; calculate the division result between the temporary variable and the set variable to obtain the influencing variable of the prior true classification.
[0108] In the implementation method of step S33, steps S331-S333 are the specific operation procedures for the computer system to determine the influencing variables of the prior true classification based on the ratio between the number of training texts corresponding to the prior true classification and the total amount of training texts corresponding to the government training text library. The influencing variables are used to improve the category imbalance problem during the debugging process of the basic government text classification model and improve the model's learning efficiency for classifying a small number of text instances.
[0109] Step S331 is to determine the ratio between the number of training texts corresponding to the prior true classification and the total amount of training texts corresponding to the government training text library. The prior true classification is the classification to which the government training text actually belongs. The number of training texts corresponding to the prior true classification is the total number of training texts under this classification, and the total amount of training texts corresponding to the government training text library is the total number of training texts under all classifications. By counting these two quantities and calculating the ratio between them, the computer system can understand the proportion of each classification in the entire training text library. This ratio is the basis for the subsequent determination of the influencing variables of the prior true classification, and it reflects the size of the training text instance scale of this classification relative to the size of the entire data set.
[0110] In order to implement step S331, the computer system may adopt a data statistics method. First, the computer system sorts and counts the training texts in the government training text library according to categories to obtain the number of training texts corresponding to each category. Then, the number of training texts in all categories is added together to obtain the total amount of training texts. Finally, the number of training texts corresponding to each category is divided by the total amount of training texts to obtain the corresponding ratio.
[0111] Step S332 is to determine the power correction coefficient based on the ratio obtained in step S331. The power correction coefficient is used to represent the existence ratio of the training texts of the remaining categories in the government training text library. The larger the influencing variable of the prior true classification, the smaller the power correction coefficient. The smaller the influencing variable of the prior true classification, the larger the power correction coefficient. The preset basic value is a natural constant. The power correction coefficient is an intermediate variable. Its function is to convert the ratio between the number of training texts corresponding to the prior true classification and the total amount of training texts into a value related to the influencing variable, so as to calculate the influencing variable later.
[0112] The specific method for the computer system to determine the power correction coefficient based on the proportion is to use the difference between the constant 1 and the proportion as the power correction coefficient. This is because the proportion reflects the proportion of the a priori true classification in the training text library. Subtracting the proportion from 1 gives the proportion of the remaining classifications in the training text library, which is the power correction coefficient. For example, in the above example, the proportion of the policy and regulation classification is 0.3, so its power correction coefficient is 1-0.3=0.7; the proportion of the people's livelihood service classification is 0.5, and its power correction coefficient is 1-0.5=0.5; the proportion of the urban construction classification is 0.2, and its power correction coefficient is 1-0.2=0.8.
[0113] The inverse correlation between the power correction coefficient and the influencing variables of the prior true classification is to achieve the emphasis on the classification of a small number of text instances. When the number of training texts for a classification is small, that is, the proportion is small, the power correction coefficient is large, which means that the training texts of other classifications account for a large proportion, and this classification belongs to the classification of a small number of text instances, and a larger influencing variable is needed to improve the model's learning ability for it; conversely, when the number of training texts for a classification is large, that is, the proportion is large, the power correction coefficient is small, and this classification belongs to the classification of a large number of text instances, and the influencing variable is relatively small.
[0114] Step S333 is to determine the influencing variables of the a priori true classification based on the power correction coefficient and the preset base value. The preset base value is the natural constant e. The computer system uses the natural constant as the base value and the power correction coefficient as the power to obtain a temporary variable, and then calculates the division result between the temporary variable and the set variable to obtain the influencing variable of the a priori true classification. The calculation formula of the influencing variable is , where I represents the influencing variable, k represents the power correction coefficient, and s represents the setting variable. The setting variable is a preset constant that scales the influencing variable to ensure that the influencing variable is in an appropriate range.
[0115] To summarize, steps S331-S333 determine the ratio between the number of training texts corresponding to the prior true classification and the total amount of training texts through a computer system, determine the power correction coefficient based on the ratio, and then determine the influencing variables of the prior true classification based on the power correction coefficient and the preset basic value, which provides an effective method for debugging the basic government text classification model to solve the category imbalance problem, improve the model's learning ability and classification accuracy for classifying a small number of text instances, and enable the model to better cope with various challenges in actual government text classification tasks.
[0116] As an implementation mode, step S41, determining the distance cost value according to the influencing variable, the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity, includes:
[0117] Step S411: determining a first correlation coefficient between the text representation array and the second centroid array representation according to the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation; the first correlation coefficient is used to indicate the degree of correlation between the text representation array and the second centroid array representation;
[0118] Step S412: determining a second correlation coefficient between the text representation array and the first centroid array representation according to the second cosine similarity, the norm of the text representation array and the norm of the first centroid array representation; the second correlation coefficient is used to indicate the degree of correlation between the text representation array and the first centroid array representation;
[0119] Step S413: Determine the distance cost value according to the influencing variable, the first correlation coefficient and the second correlation coefficient.
[0120] Step S411 determines the first correlation coefficient between the text representation array and the second centroid array representation based on the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation. The first correlation coefficient is used to indicate the degree of correlation between the text representation array and the second centroid array representation. The target cosine similarity is obtained by adjusting the first cosine similarity between the text representation array and the second centroid array representation. Its purpose is to expand the correlation between the two and allow the model to pay more attention to the subtle differences between different categories when classifying. The text representation array is a numerical array obtained after the government training text is processed by the text embedding coding network of the basic government text classification model, which reflects the semantic information and features of the government training text. The second centroid array representation is an array representation of the centroids of other categories except the a priori true classification corresponding to the current government training text. In practical applications, for example, the text representation array and the second centroid array representation will be standardized so that their norms are both equal to 1.
[0121] Since the array norms are all 1, the first dot product of the text representation array and the second centroid array can be determined based on the target cosine similarity, and the first dot product is equal to the target cosine similarity. Then, the natural constant e is used as the base value and the first dot product is used as the power to obtain the first correlation coefficient.
[0122] Step S412 determines the second correlation coefficient between the text representation array and the first centroid array representation based on the second cosine similarity, the norm of the text representation array and the norm of the first centroid array representation, and the second correlation coefficient is used to indicate the degree of correlation between the text representation array and the first centroid array representation. The second cosine similarity is the cosine similarity between the text representation array and the first centroid array representation, which measures the directional similarity between the text representation array and the a priori true classification centroid. The first centroid array representation is the array representation of the centroid of the a priori true classification, and the norm is also 1 after normalization.
[0123] Since the array norms are all 1, the second dot product of the text representation array and the first centroid array can be determined based on the second cosine similarity, and the second dot product is equal to the second cosine similarity. Then, the second correlation coefficient is determined using the natural constant e as the base value and the second dot product as the power.
[0124] The computer system can calculate the second correlation coefficient by using a technical means similar to step S411. The value of the second cosine similarity is stored in a numerical calculation library, and then an exponential function is used to calculate the result with the natural constant e as the base value and the second cosine similarity as the power to obtain the second correlation coefficient.
[0125] Step S413 determines the distance cost value based on the influencing variable, the first correlation coefficient and the second correlation coefficient. The influencing variable is determined by step S33 based on the ratio between the number of training texts corresponding to the prior true classification and the total amount of training texts corresponding to the government training text library. It is inversely correlated with the ratio, that is, the influencing variable for the classification of a small number of text instances is greater than the influencing variable for the classification of a large number of text instances. The purpose of introducing the influencing variable is to optimize the distance cost value obtained by the government training text classified based on a small number of text instances, improve the problem of category imbalance, and improve the learning efficiency of the basic government text classification model for the classification of a small number of text instances.
[0126] When determining the distance cost value, the computer system comprehensively considers the influencing variables, the first correlation coefficient and the second correlation coefficient. The first correlation coefficient reflects the degree of correlation between the text representation array and the rest of the classification centroids, and the second correlation coefficient reflects the degree of correlation between the text representation array and the a priori true classification centroids. The influencing variables adjust the distance cost value according to the scale of the classified training text instances, so that the classification of a small number of text instances receives more attention in the model training.
[0127] The calculation of the distance cost value is based on a certain combination relationship between the first correlation coefficient, the second correlation coefficient and the influencing variable. Specifically, the computer system will first consider the proportional relationship between the first correlation coefficient and the second correlation coefficient, and then combine the influencing variable with the ratio to obtain the distance cost value through logarithmic operations and other methods. The core idea is that when the text representation array has a high degree of correlation with the prior true classification centroid (the second correlation coefficient is large), and a low degree of correlation with the other classification centroids (the first correlation coefficient is small), the distance cost value should be small, indicating that the model classification is more accurate; conversely, when the text representation array has a high degree of correlation with the other classification centroids (the first correlation coefficient is large), and a low degree of correlation with the prior true classification centroid (the second correlation coefficient is small), the distance cost value should be large, indicating that the model classification may have errors. The role of the influencing variable is to increase the distance cost value corresponding to this error in the case of a small number of text instance classifications, thereby prompting the model to pay more attention to the training texts for the classification of a small number of text instances.
[0128] Assuming the influencing variable is I, the first correlation coefficient is K1, and the second correlation coefficient is K2, a possible distance cost calculation formula is: .
[0129] The computer system will modify the model parameters of the basic government text classification model based on the distance cost value, making the model more accurate in classifying government texts. When the distance cost value is large, it means that the classification error of the model is large. The computer system will adjust the model parameters according to the size and direction of the distance cost value, making the model more inclined to classify the text into the prior true classification and reducing the number of misclassifications. Especially for the classification of a small number of text instances, due to the effect of the influencing variables, the adjustment of the distance cost value will be more obvious, thereby prompting the model to better learn the characteristics of the classification of a small number of text instances.
[0130] In summary, steps S411-S413 determine the distance cost value through a computer system according to the influencing variables, the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity, which provides an effective method for debugging the basic government text classification model to solve the category imbalance problem, improve the model's learning ability and classification accuracy for classifying a small number of text instances, and enable the model to better cope with various challenges in actual government text classification tasks.
[0131] As an implementation mode, the text representation array, the first centroid array representation and the second centroid array representation are standardized array representations, and the norm of the text representation array, the norm of the first centroid array representation and the norm of the second centroid array representation are all equal to 1.
[0132] The process of determining the first correlation coefficient and the second correlation coefficient can refer to the following steps:
[0133] Step S41a: Based on the target cosine similarity, the norm of the text representation array and the norm of the second centroid array, determine the first dot product of the text representation array and the second centroid array, take the natural constant as the base value and the first dot product as the power, and obtain the first correlation coefficient;
[0134] Step S41b: Based on the second cosine similarity, the norm of the text representation array and the norm represented by the first centroid array, determine the second dot product of the norm of the text representation array and the first centroid array, take the natural constant as the base value and the second dot product as the power, and determine the second correlation coefficient.
[0135] In some implementations, if the number of the remaining categories is multiple, then step S413, determining the distance cost value according to the influencing variable, the first correlation coefficient and the second correlation coefficient, may specifically include:
[0136] Step S4131: obtaining a plurality of first correlation coefficients;
[0137] Step S4132: determine the sum of the first correlation coefficient and the second correlation coefficient, determine the ratio between the second correlation coefficient and the sum, perform logarithmic operation on the ratio, and obtain the first distance cost value; and take the product of the first distance cost value and the influencing variable as the distance cost value.
[0138] Step S41a determines the first dot product of the text representation array and the second centroid array representation based on the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation, and takes the natural constant as the base value and the first dot product as the power to obtain the first correlation coefficient. The target cosine similarity is obtained by adjusting the first cosine similarity between the text representation array and the second centroid array representation, and its purpose is to expand the correlation between the two, so that the model can more keenly capture the subtle differences between different classifications during the classification process. The text representation array is a numerical array obtained after the government training text is processed by the text embedding encoding network of the basic government text classification model. It contains the semantic information and features of the government training text, and the second centroid array representation is an array representation of the centroids of other classifications except the a priori true classification corresponding to the current government training text. In practical applications, for example, the text representation array and the second centroid array representation will be standardized so that their norms are both equal to 1.
[0139] Since the array norms are all 1, the first dot product is equal to the target cosine similarity. The first correlation coefficient is used to intuitively indicate the degree of correlation between the text representation array and the second centroid array representation. The computer system obtains the first correlation coefficient by performing exponential operation with the natural constant e as the base value and the target cosine similarity as the power.
[0140] Step S41b determines the second dot product of the norm of the text representation array and the first centroid array representation based on the second cosine similarity, the norm of the text representation array, and the norm of the first centroid array representation, and determines the second correlation coefficient with the natural constant as the base value and the second dot product as the power. The second cosine similarity is the cosine similarity between the text representation array and the first centroid array representation, which accurately measures the directional similarity between the text representation array and the a priori true classification centroid. The first centroid array representation is the array representation of the centroid of the a priori true classification, and the norm is also 1 after normalization.
[0141] Since the array norms are all 1, the second dot product is equal to the second cosine similarity. The second correlation coefficient is used to clearly indicate the degree of correlation between the text representation array and the first centroid array representation. The computer system obtains the second correlation coefficient by performing exponential operation with the natural constant e as the base value and the second cosine similarity as the power.
[0142] The computer system can calculate the second correlation coefficient by using a technical means similar to step S41a. The value of the second cosine similarity is stored in a numerical calculation library, and then an exponential function is used to calculate the result with the natural constant e as the base value and the second cosine similarity as the power to obtain the second correlation coefficient.
[0143] Step S4131 is to obtain multiple first correlation coefficients. In the actual government text classification task, the number of other categories is, for example, multiple, and each of the other categories has a corresponding second centroid array representation. The computer system will calculate the first correlation coefficient between the text representation array and the second centroid array representation according to the method of step S41a for each second centroid array representation. Therefore, multiple first correlation coefficients will be obtained in the end, and these first correlation coefficients fully reflect the degree of correlation between the text representation array and each of the other classification centroids.
[0144] The computer system can loop through each second centroid array representation, calculate and store the corresponding first correlation coefficients in sequence. These first correlation coefficients can be stored in a data structure such as a list or an array for subsequent processing.
[0145] Step S4132 determines the sum of the first correlation coefficient and the second correlation coefficient, determines the ratio between the second correlation coefficient and the sum, performs a logarithmic operation on the ratio, obtains a first distance proxy value, and uses the product of the first distance proxy value and the influencing variable as the distance proxy value. The influencing variable is determined by the ratio between the number of training texts corresponding to the prior true classification and the total amount of training texts corresponding to the government training text library in step S33, and is inversely correlated with the ratio, that is, the influencing variable for the classification of a small number of text instances is greater than the influencing variable for the classification of a large number of text instances. The purpose of introducing the influencing variable is to optimize the distance proxy value obtained by the government training text classified based on a small number of text instances, improve the problem of category imbalance, and improve the learning efficiency of the basic government text classification model for the classification of a small number of text instances.
[0146] First, the computer system adds all the first correlation coefficients, and then adds the second correlation coefficients to obtain the sum. Assume that there are n first correlation coefficients in total. , the second correlation coefficient is K2, then the addition result is Then, the ratio between the second correlation coefficient and the summation result is calculated. Next, perform a logarithmic operation on the ratio to obtain the first distance cost value D1 = -ln(P). Finally, multiply the first distance cost value by the influencing variable I to obtain the final distance cost value .
[0147] The determined distance cost value plays a vital role in the debugging of the basic government text classification model. The computer system will finely modify the model parameters of the basic government text classification model based on the distance cost value, so that the model is more accurate in classifying government texts. When the distance cost value is large, it means that the classification error of the model is large. The computer system will make detailed adjustments to the model parameters according to the size and direction of the distance cost value, so that the model is more inclined to accurately classify the text into the prior true classification, significantly reducing the misclassification. Especially for the classification of a small number of text instances, due to the role of the influencing variables, the adjustment of the distance cost value will be more obvious, thereby prompting the model to better learn the characteristics of the classification of a small number of text instances and effectively solve the problem of category imbalance.
[0148] As an implementation manner, step S402 or step S412, determining a second correlation coefficient between the text representation array and the first centroid array representation based on the second cosine similarity, the norm of the text representation array, and the norm of the first centroid array representation, may include:
[0149] Step S42a1: determining the addition result between the second cosine similarity and the second similarity adjustment factor to obtain a modified cosine similarity that is smaller than the second cosine similarity;
[0150] Step S42a2: Determine the second correlation coefficient between the text representation array and the first centroid array representation based on the corrected cosine similarity, the norm of the text representation array and the norm of the first centroid array representation; specifically, the process can be based on the corrected cosine similarity, the norm of the text representation array and the norm of the first centroid array representation, determine the third dot product between the text representation array and the first centroid array representation, take the natural constant as the base value, and use the third dot product as the power to determine the second correlation coefficient.
[0151] Step S42a1 determines the addition result between the second cosine similarity and the second similarity adjustment factor, and obtains a modified cosine similarity that is smaller than the second cosine similarity. The second cosine similarity is used to measure the directional similarity between the text representation array and the first centroid array representation, and the second similarity adjustment factor is an interval angle, and its introduction is intended to increase the difference requirement between the text representation array and the first centroid array representation. The computer system adds the second cosine similarity to the second similarity adjustment factor, and since the second similarity adjustment factor is, for example, a negative number, the obtained modified cosine similarity will be smaller than the second cosine similarity. The significance of this operation is that the text representation arrays of the training texts of the same classification need to be closer to each other in order to be correctly classified, and the distance between the text representation arrays of the training texts of different classifications needs to be larger to prevent being misclassified, so that the training texts of the same classification are closer to each other and the training texts of different classifications are farther away from each other.
[0152] Step S42a2 determines the second correlation coefficient between the text representation array and the first centroid array representation based on the corrected cosine similarity, the norm of the text representation array, and the norm of the first centroid array representation. Since the text representation array and the first centroid array representation are, for example, standardized array representations, their norms are both equal to 1, the computer system determines the dot product of the text representation array and the first centroid array representation based on the corrected cosine similarity (the dot product is equal to the corrected cosine similarity at this time), and then determines the second correlation coefficient with the natural constant e as the base value and the dot product as the power.
[0153] Step S4131 obtains multiple first correlation coefficients. In actual government text classification tasks, the number of other categories is often multiple, and each of the other categories has a corresponding second centroid array representation. For each second centroid array representation, the computer system calculates the first correlation coefficient between the text representation array and the second centroid array representation according to the method previously used to determine the first correlation coefficient (based on the target cosine similarity, the norm of the text representation array, and the norm of the second centroid array representation, with the natural constant e as the base value and the target cosine similarity as the power to obtain the first correlation coefficient). In the end, multiple first correlation coefficients will be obtained, which reflect the degree of correlation between the text representation array and each of the other classification centroids.
[0154] Step S4132 determines the sum of the first correlation coefficient and the second correlation coefficient, determines the ratio between the second correlation coefficient and the sum, performs a logarithmic operation on the ratio, obtains a first distance proxy value, and uses the product of the first distance proxy value and the influencing variable as the distance proxy value. The influencing variable is determined based on the ratio between the number of training texts corresponding to the prior true classification and the total amount of training texts corresponding to the government training text library, and is inversely correlated with the ratio, that is, the influencing variable for the classification of a small number of text instances is greater than the influencing variable for the classification of a large number of text instances. The influencing variable is introduced to optimize the distance proxy value obtained by the government training text classified based on a small number of text instances, and improve the problem of category imbalance.
[0155] In practical applications, when executing these steps, the computer system needs to pay attention to the values of the second similarity adjustment factor and the influencing variable. The value of the second similarity adjustment factor needs to be adjusted according to the specific data set and model performance. If the value is too large, it may make it difficult to correctly classify the same type of training texts; if the value is too small, it will not be able to effectively increase the distance between different classification training texts. The value of the influencing variable also needs to be carefully determined. Too large an influencing variable may cause the model to over-focus on the classification of a small number of text instances, resulting in overfitting; too small an influencing variable cannot effectively improve the problem of class imbalance.
[0156] The determined distance cost is crucial for debugging the basic government text classification model. The computer system will correct the model parameters of the basic government text classification model based on the distance cost, making the model more accurate in classifying government texts. When the distance cost is large, it means that the classification error of the model is large. The computer system will adjust the model parameters according to the size and direction of the distance cost, making the model more inclined to classify the text into the prior true classification and reducing the number of misclassifications. Through the synergy of this series of steps, the basic government text classification model can better handle complex government text classification tasks and improve classification accuracy and generalization ability.
[0157] In summary, steps S42a1-S42a2 and steps S4131-S4132 correct the second cosine similarity through the computer system to determine the second correlation coefficient, and then calculate the distance cost value in combination with multiple first correlation coefficients and influencing variables, which provides an effective method for debugging the basic government text classification model, can improve the category imbalance problem, enhance the model's classification performance for government texts, and make it more adaptable to the needs of practical applications.
[0158] As an embodiment, the method further includes:
[0159] Step S40A: Determine the classification boundary cost value according to the commonality measurement results of the first centroid array representation and the second centroid array representation.
[0160] Based on this, step S50, the model parameters of the basic government text classification model are modified according to the distance cost value to obtain the target government text classification model, including:
[0161] Step S51: adding the distance cost value and the classification boundary cost value to obtain a first comprehensive cost;
[0162] Step S52: Modify the model parameters of the basic government text classification model based on the first comprehensive cost to obtain the target government text classification model.
[0163] The derivation step S40A determines the classification boundary cost value, also called the inter-class cost, based on the commonality measurement results of the first centroid array representation and the second centroid array representation. The first centroid array representation is an array representation of the centroid of the a priori true classification, and the second centroid array representation is an array representation of the centroid of other classifications except the a priori true classification corresponding to the current government training text. The commonality measurement result reflects the similarity between the centroids of different classifications. The computer system determines the classification boundary cost value through this result. The cost value is used to measure the degree of distinction between different classifications, which helps the model to more clearly divide the boundaries between different categories during classification. When the number of remaining classifications is one, the computer system determines the dot product between the first centroid array representation and the second centroid array representation, and uses the absolute value of the dot product as the classification boundary cost value. The dot product is the result of summing the corresponding elements of two vectors after multiplication, which can reflect the directional relationship between the two vectors. The larger the absolute value of the dot product, the more similar the directions between the two centroid array representations are, and the more blurred the classification boundary is; the smaller the absolute value of the dot product, the greater the directional difference between the two centroid array representations, and the clearer the classification boundary is.
[0164] When there are multiple remaining categories, the computer system determines the dot products between the first centroid array representation and each second centroid array representation, determines the addition result between each dot product, and uses the absolute value of the addition result as the category boundary cost value.
[0165] Step S51 adds the distance cost value and the classification boundary cost value to obtain the first comprehensive cost. The distance cost value is determined based on the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity, etc. It reflects the degree of correlation between the text representation array and the centroids of different classifications and the error of the model during classification. The classification boundary cost value measures the degree of distinction between different classifications. Adding the two together to obtain the first comprehensive cost can comprehensively consider the classification error of the model at the training text instance level and the distinction at the category level, and more comprehensively evaluate the performance of the model.
[0166] Step S52 modifies the model parameters of the basic government text classification model based on the first comprehensive cost to obtain the target government text classification model. The basic government text classification model includes, for example, a text embedding coding network and a fully connected network. The model parameters are variables such as weight parameters in these networks, which determine how the model processes and classifies the input government text. The goal of the computer system is to make the first comprehensive cost as small as possible by adjusting these model parameters, thereby improving the classification performance of the model.
[0167] The computer system can use optimization algorithms such as gradient descent to correct model parameters. The basic idea of gradient descent is to update the model parameters along the negative gradient direction of the objective function (the objective function here is the first comprehensive cost), thereby gradually reducing the value of the objective function.
[0168] As an implementation mode, the number of government affairs training texts is multiple, then the method further includes:
[0169] Step S40B: Determine the classification cohesion cost based on the text representation arrays of multiple government training texts.
[0170] Based on this, step S52, based on the first comprehensive cost, corrects the model parameter of the basic government text classification model to obtain the target government text classification model, including:
[0171] Step S521: Add the first comprehensive cost and the classification cohesion cost to obtain a second comprehensive cost;
[0172] Step S522: Modify the model parameters of the basic government text classification model based on the second comprehensive cost to obtain the target government text classification model.
[0173] In step S40B, the computer system determines the classification cohesion cost value based on the text representation arrays of multiple government training texts. The classification cohesion cost value reflects the closeness between the government training texts under the same classification. The smaller the value, the more closely the texts under the same classification are clustered together, and the better the cohesion of the classification. In order to determine this value, the computer system determines the classification cluster centroid array representation of the a priori true classification. The classification cluster centroid array representation can be understood as the center position of the text representation array of all government training texts under the same classification, similar to the center position of a group of points in a multidimensional space. The technical means for determining the classification cluster centroid array representation is to calculate the average value of the text representation array of all government training texts under the same classification.
[0174] Next, the computer system determines the classification cohesion cost value based on the text representation arrays and classification cluster centroid array representations of multiple government training texts. Specifically, for each government training text, the computer system calculates the gap between its text representation array and the classification cluster centroid array representation. This gap is measured, for example, by array distance, and a feasible method is Euclidean distance. The array distance of each government training text is the sub-classification cohesion cost value corresponding to the text representation array. Finally, the computer system determines the average of the sub-classification cohesion cost values of all text representation arrays as the classification cohesion cost value.
[0175] After completing step S40B to determine the classification cohesion cost, the computer system enters step S52, which is to modify the model parameters of the basic government text classification model based on the first comprehensive cost and the classification cohesion cost calculated previously to obtain the target government text classification model. Step S52 is further divided into step S521 and step S522.
[0176] In step S521, the computer system adds the first comprehensive cost and the classification cohesion cost to obtain the second comprehensive cost. The first comprehensive cost is obtained by adding the distance cost and the classification boundary cost in the previous step. It comprehensively considers the relationship between the text representation array and the different classification centroid array representations and the boundary conditions between different classifications. The classification cohesion cost reflects the degree of aggregation of texts under the same classification. The second comprehensive cost obtained by adding these two costs can more comprehensively reflect the performance of the basic government text classification model in the classification process.
[0177] In step S522, the computer system modifies the model parameters of the basic government text classification model based on the second comprehensive cost to obtain the target government text classification model. In the basic government text classification model, for example, it includes a text embedding coding network and a fully connected network, and its model parameters are mainly network parameters of the text embedding coding network and connection influence variables of the fully connected network. The computer system adjusts these model parameters through an optimization algorithm (such as a stochastic gradient descent algorithm) to minimize the second comprehensive cost. The basic idea of the stochastic gradient descent algorithm is to randomly select a training text instance or a small batch of training text instances in each iteration, calculate the gradients of these training text instances, and then update the model parameters according to the direction of the gradient.
[0178] By determining the classification cohesion cost value through step S40B and correcting the model parameters of the basic government text classification model based on the cost value through steps S521-S522, the computer system can comprehensively consider the classification cohesion, the boundaries between different classifications, and the relationship between the text representation array and the centroid array representation, thereby improving the performance of the basic government text classification model, so that the target government text classification model has higher accuracy and generalization ability in the government text classification task. In actual government text classification scenarios, such as government service section classification tasks or sentiment polarity classification tasks, the target government text classification model after such optimization can more accurately classify government texts into appropriate categories, providing strong support for the management and processing of government information.
[0179] As an implementation mode, step S40B, based on the text representation arrays of the plurality of government training texts, determines the classification cohesion cost, including:
[0180] Step S40B1: determining a classification cluster centroid array representation of a priori true classification based on a text representation array of a plurality of government training texts;
[0181] Step S40B2: Determine the classification cohesion cost based on the text representation array and classification cluster centroid array representation of multiple government training texts.
[0182] In step S40B1, the computer system determines the classification cluster centroid array representation of the a priori true classification based on the text representation arrays of multiple government training texts. The classification cluster centroid array representation refers to the center position of the text representation arrays of all government training texts under the same classification in a multidimensional space. In the scenario of government text classification, the classification cluster centroid array representation reflects the overall characteristics of the government training texts under the same classification. In order to determine the classification cluster centroid array representation, the computer system processes the text representation arrays of all government training texts under the same classification. The specific method is to calculate the average value of these text representation arrays.
[0183] In step S40B2, the computer system determines the classification cohesion cost value based on the text representation arrays and classification cluster centroid array representations of multiple government training texts. The classification cohesion cost value is an indicator to measure the closeness between government training texts under the same classification. The smaller its value is, the more closely the government training texts under the same classification are clustered around the classification cluster centroid array representation in the multidimensional space, and the better the cohesion of the classification. In order to determine the classification cohesion cost value, the computer system calculates the gap between the text representation array and the classification cluster centroid array representation of each government training text. This gap is measured, for example, by array distance, and the feasible method is Euclidean distance. The array distance of each government training text is the sub-classification cohesion cost value corresponding to the text representation array. After calculating the sub-classification cohesion cost value of each government training text, the computer system determines the average of the sub-classification cohesion cost values of all text representation arrays as the classification cohesion cost value.
[0184] In the actual government text classification scenario, by determining the classification cluster centroid array representation and classification cohesion cost, the computer system can better understand the distribution of government training texts under the same category. The classification cluster centroid array representation can be used as a representative feature of the classification for subsequent classification judgment. The classification cohesion cost can help evaluate the quality of the classification. If the classification cohesion cost is large, it means that the government training texts under the same category are relatively scattered, and it may be necessary to further adjust the classification strategy or add more training texts. For example, in a government text collection containing multiple government service section classifications, for the "administrative approval" classification, if the calculated classification cohesion cost is large, it may mean that the government training texts under this category contain some texts that are not very relevant to the core features of "administrative approval", and these texts need to be reviewed and classified again.
[0185] In addition, step S40B1 and step S40B2 can also be combined with other steps to jointly optimize the basic government text classification model. In subsequent steps, the classification cohesion cost will be combined with other costs (such as distance cost, classification boundary cost) to form a comprehensive cost, which is used to correct the model parameters of the basic government text classification model. By continuously adjusting the model parameters, the comprehensive cost is minimized, thereby improving the performance of the basic government text classification model, so that it can classify government texts more accurately. When performing classification training on a data set containing a large number of government texts, the computer system gradually optimizes the model by continuously iterating these steps, and finally obtains a target government text classification model that can effectively distinguish different government text classifications.
[0186] When the computer system executes step S40B1 and step S40B2, by determining the classification cluster centroid array representation and classification cohesion generation value, it can deeply understand the characteristics and distribution of government training texts under the same classification, providing an important basis for government text classification. These steps not only help to evaluate the quality of classification, but also work together with other steps to jointly optimize the basic government text classification model, improve the accuracy and efficiency of government text classification, and provide strong support for the management and processing of government information.
[0187] As an implementation method, the basic government text classification model includes a text embedding coding network and a fully connected network. Then, step S50, according to the distance cost value, corrects the model parameter variables of the basic government text classification model to obtain the target government text classification model, including:
[0188] Step S501: Modify the network parameters of the text embedding coding network and the connection influencing variables of the fully connected network according to the distance cost value to obtain a target government text classification model; the target government text classification model includes a target text embedding coding network and a target fully connected network.
[0189] In step S501, the computer system modifies the network parameter variables of the text embedding coding network of the basic government text classification model and the connection influencing variables of the fully connected network according to the distance cost value, and then obtains the target government text classification model, which includes the target text embedding coding network and the target fully connected network. In the government text classification process, the text embedding coding network of the basic government text classification model is responsible for converting the input government text into a vector representation suitable for model processing, that is, a text representation array, and the fully connected network performs classification prediction based on these text representation arrays, and the connection influencing variables (that is, the weight parameters of the fully connected network) determine the mapping relationship between input and output.
[0190] To achieve this goal, the computer system can use optimization algorithms, such as the stochastic gradient descent (SGD) algorithm. The basic idea of the stochastic gradient descent algorithm is to randomly select a training text instance or a small batch of training text instances in each iteration, calculate the gradient of these training text instances, and then update the model parameters according to the direction of the gradient.
[0191] In the specific operation, for updating the network parameters of the text embedding coding network, the computer system first calculates the gradient of the distance cost value with respect to the network parameters. For updating the connection influence variable of the fully connected network, the computer system also calculates the gradient of the distance cost value with respect to the connection influence variable. Assuming that there is a connection weight matrix V as the connection influence variable in the fully connected network, the computer system calculates , and then update the value of the connection weight matrix V according to the stochastic gradient descent formula. For example, if the initial connection weight matrix V is a The matrix of the calculated gradient For one The matrix, learning rate = 0.01, then each element of the updated connection weight matrix V' The calculation formula is: ,in is the element in the kth row and lth column of the original connection weight matrix V, (( is the gradient matrix The element at the kth row and lth column of .
[0192] By correcting the network parameters of the text embedding encoding network and the connection influence variables of the fully connected network of the basic government text classification model according to the distance cost value, the computer system can continuously optimize the performance of the model, so that the target government text classification model has higher accuracy and generalization ability in the government text classification task, thereby providing more reliable support for the management and processing of government information.
[0193] Figure 2 A hardware entity diagram of a computer system provided by an embodiment of the present invention is as follows Figure 2 As shown, the hardware entity of the computer system 1000 includes: a processor 1001 and a memory 1002, wherein the memory 1002 stores a computer program that can be run on the processor 1001, and the processor 1001 implements the steps in the method of any of the above embodiments when executing the program.
[0194] The above description is only an implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A data classification method based on a large model, characterized in that: The method comprises: Obtain the target government text to be classified, and call the debugged and converged target government text classification model; Loading the target government text into the target government text classification model to classify the target government text based on the target government text classification model to obtain a target text classification result; The target government text classification model is obtained by debugging the basic government text classification model, and the debugging process includes the following steps: Extracting a text representation array of government training texts in a government training text library based on a basic government text classification model; In the connection influencing variables of the fully connected network of the basic government text classification model, a first centroid array representation of a priori true classification and a second centroid array representation of other classifications are obtained, wherein the priori true classification is the classification corresponding to the government training text; Adjusting a first cosine similarity between the text representation array and the second centroid array representation to obtain a target cosine similarity greater than the first cosine similarity; Determining a distance cost value according to the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation, and a second cosine similarity; the second cosine similarity is a cosine similarity between the text representation array and the first centroid array representation; Modifying the model parameter variables of the basic government text classification model according to the distance cost value to obtain the target government text classification model; The determining of the distance cost value according to the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity comprises: Determining a first correlation coefficient between the text representation array and the second centroid array representation according to the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation; the first correlation coefficient is used to indicate the degree of correlation between the text representation array and the second centroid array representation; Determining a second correlation coefficient between the text representation array and the first centroid array representation according to a second cosine similarity, a norm of the text representation array, and a norm of the first centroid array representation; the second correlation coefficient is used to indicate a degree of correlation between the text representation array and the first centroid array representation; Determining a distance cost value according to the influencing variable, the first correlation coefficient and the second correlation coefficient; The process of determining the influencing variables is as follows: determining the ratio between the number of training texts corresponding to the a priori true classification and the total amount of training texts corresponding to the government training text library; determining a power correction factor based on the ratio; Determining the influencing variables of the a priori true classification according to the power correction coefficient and the preset basic value; Among them, the power correction coefficient is used to represent the existence ratio of the training texts of the remaining categories in the government training text library. The larger the influencing variable of the priori true classification, the smaller the power correction coefficient, the smaller the influencing variable of the priori true classification, the larger the power correction coefficient, and the preset basic value is a natural constant; specifically, determine the ratio between the number of training texts and the total amount of training texts, and take the difference between the constant 1 and the ratio as the power correction coefficient; determine the natural constant as the basic value, and take the power correction coefficient as the power to obtain a temporary variable; calculate the division result between the temporary variable and the set variable to obtain the influencing variable of the priori true classification; the influencing variable is inversely correlated with the ratio.
2. The method according to claim 1, characterized in that: The step of adjusting the first cosine similarity between the text representation array and the second centroid array representation to obtain a target cosine similarity greater than the first cosine similarity includes: Obtaining a first cosine similarity between the text representation array and the second centroid array representation; Calculating a difference between the first cosine similarity and a first similarity adjustment factor to obtain a target cosine similarity greater than the first cosine similarity; The basic government text classification model includes a text embedding coding network and a fully connected network; the model parameter variables of the basic government text classification model are corrected according to the distance cost value to obtain the target government text classification model, including: The network parameters of the text embedding coding network and the connection influencing variables of the fully connected network are modified according to the distance cost value to obtain the target government text classification model; the target government text classification model includes a target text embedding coding network and a target fully connected network.
3. The method according to claim 1, characterized in that in, The text representation array, the first centroid array representation and the second centroid array representation are standardized array representations, and the norm of the text representation array, the norm of the first centroid array representation and the norm of the second centroid array representation are all equal to 1; The process of determining the first correlation coefficient and the second correlation coefficient includes: Based on the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation, determine the first dot product of the text representation array and the second centroid array representation, take the natural constant as the base value and the first dot product as the power, and obtain the first correlation coefficient; Based on the second cosine similarity, the norm of the text representation array and the norm represented by the first centroid array, determine a second dot product of the norm of the text representation array and the first centroid array, and determine the second correlation coefficient by taking a natural constant as a base value and the second dot product as a power; If the number of the remaining categories is multiple, determining the distance cost value according to the influencing variable, the first correlation coefficient and the second correlation coefficient includes: Obtaining a plurality of the first correlation coefficients; Determine the addition result of each of the first correlation coefficients and the second correlation coefficients, determine the ratio between the second correlation coefficient and the addition result, perform logarithmic operation on the ratio to obtain a first distance cost value; and use the product of the first distance cost value and the influencing variable as the distance cost value.
4. The method according to claim 1, characterized in that: The determining of the distance cost value according to the target cosine similarity, the text representation array, the first centroid array representation, the second centroid array representation and the second cosine similarity comprises: Determining a first correlation coefficient between the text representation array and the second centroid array representation based on the target cosine similarity, the norm of the text representation array, and the norm of the second centroid array representation; Determining a second correlation coefficient between the text representation array and the first centroid array representation based on a second cosine similarity, a norm of the text representation array, and a norm of the first centroid array representation; A distance cost value is determined according to the first correlation coefficient and the second correlation coefficient.
5. The method according to claim 4, characterized in that in, The text representation array, the first centroid array representation and the second centroid array representation are standardized array representations, and the norm of the text representation array, the norm of the first centroid array representation and the norm of the second centroid array representation are all equal to 1; The process of determining the first correlation coefficient and the second correlation coefficient includes: Based on the target cosine similarity, the norm of the text representation array and the norm of the second centroid array representation, determine the first dot product of the text representation array and the second centroid array representation, take the natural constant as the base value and the first dot product as the power, and obtain the first correlation coefficient; Based on the second cosine similarity, the norm of the text representation array and the norm represented by the first centroid array, determine a second dot product of the norm of the text representation array and the first centroid array, and determine the second correlation coefficient by taking a natural constant as a base value and the second dot product as a power; If the number of the remaining categories is multiple, determining the distance cost value according to the first association coefficient and the second association coefficient includes: Obtaining a plurality of the first correlation coefficients; Determine the sum of the first correlation coefficient and the second correlation coefficient, determine the ratio between the second correlation coefficient and the sum, perform a logarithmic operation on the ratio, and obtain the distance cost value.
6. The method according to claim 1 or 4, characterized in that: The determining, based on the second cosine similarity, the norm of the text representation array and the norm of the first centroid array representation, a second correlation coefficient between the text representation array and the first centroid array representation comprises: Determine an addition result between the second cosine similarity and the second similarity adjustment factor to obtain a modified cosine similarity that is smaller than the second cosine similarity; Determine a second correlation coefficient between the text representation array and the first centroid array representation based on the corrected cosine similarity, the norm of the text representation array, and the norm of the first centroid array representation; specifically, determine a third dot product between the text representation array and the first centroid array representation based on the corrected cosine similarity, the norm of the text representation array, and the norm of the first centroid array representation, and determine the second correlation coefficient by taking a natural constant as a base value and the third dot product as a power; The method further comprises: Determining a classification boundary cost value according to a commonality measurement result of the first centroid array representation and the second centroid array representation; The step of modifying the model parameter of the basic government text classification model according to the distance cost value to obtain the target government text classification model includes: Adding the distance cost value and the classification boundary cost value to obtain a first comprehensive cost; The model parameters of the basic government text classification model are modified based on the first comprehensive cost to obtain the target government text classification model.
7. The method according to claim 6, characterized in that The number of the government affairs training texts is multiple; the method further includes: Determining a classification cohesion cost based on a text representation array of a plurality of government training texts; The step of modifying the model parameter of the basic government text classification model based on the first comprehensive cost to obtain the target government text classification model includes: Adding the first comprehensive cost and the classification cohesion cost value to obtain a second comprehensive cost; Modifying the model parameter variables of the basic government text classification model based on the second comprehensive cost to obtain the target government text classification model; The step of determining the classification cohesion cost based on the text representation arrays of the plurality of government training texts includes: Determine the classification cluster centroid array representation of the prior true classification based on the text representation arrays of the plurality of government training texts; Based on the text representation arrays of the plurality of government training texts and the classification cluster centroid array representations, a classification cohesion cost is determined.
8. A computer system comprising a memory and a processor, wherein the memory stores a computer program executable on the processor, wherein: When the processor executes the program, the steps in the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Text query method and device and storage medium
CN114416979A
Classification model training method and device, electronic equipment and storage medium
CN118172641A