Data classification and grading method and device, equipment and storage medium
By summarizing the divided data and using pre-trained data classification and grading models and rules engines, the data is classified and grading, which solves the problems of low efficiency and strong subjectivity in the existing technology, and achieves efficient and accurate data classification and grading.
Patent Information
- Application Number
- CN202510194918.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-10
AI Technical Summary
The existing data classification and grading methods rely on manual labeling, which is inefficient and subjective, making it difficult to ensure the accuracy of labeling while ensuring the efficiency of labeling, especially for large-scale and high-complex data.
The target data is summarized to obtain the target data; the target data and classification rating prompt words are input into the pre-trained data classification rating model, combined with the rule engine that incorporates classification rating rules for each application field, the target data is classified and graded, and finally the final classification and rating results are determined based on multiple results.
It improves the efficiency and accuracy of data classification and grading, reduces manual intervention, reduces labor costs, avoids the subjectivity and uncertainty of manual classification and grading, and enhances the reliability of results.
Smart Images

Figure CN120123818A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention relate to the field of big data technology, and in particular, to a method, apparatus, device, and storage medium for data classification and grading. Background Art
[0002] With the rapid development of Internet, big data, and cloud computing technologies, data has become an extremely valuable resource in modern society. However, the emergence of massive data has brought unprecedented challenges to data management and effective utilization. In order to efficiently store, retrieve, analyze, and protect this data, how to accurately and efficiently classify and grade it has become a key problem that urgently needs to be solved in the current data processing field.
[0003] Currently, existing data classification and grading methods mainly rely on manual annotation, which is not only inefficient, but also subjective and error-prone. Especially for large-scale and high-complexity data, the workload of manual annotation is extremely heavy, and it is difficult to ensure the accuracy of annotation while guaranteeing the annotation efficiency.
[0004] Therefore, there is an urgent need to propose a new method to solve the above problems. Summary of the Invention
[0005] The present invention provides a method, apparatus, device, and storage medium for data classification and grading, which improves the efficiency and accuracy of data classification and grading.
[0006] In a first aspect, embodiments of the present invention provide a method for data classification and grading, the method comprising:
[0007] Obtaining target data corresponding to the data to be divided by generalizing the data to be divided;
[0008] Inputting the target data and classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model divides the target data based on the classification and grading prompt words to obtain a first classification result and a first grading result of the data to be divided;
[0009] Classifying and grading the target data by using a rule engine with classification and grading rules for each application field to obtain a second classification result and a second grading result of the data to be divided;
[0010] Determining a classification result of the data to be divided according to the first classification result and the second classification result, and determining a grading result of the data to be divided according to the first grading result and the second grading result.
[0011] The technical solution of the present invention obtains the target data corresponding to the data to be classified by generalizing the data to be classified; inputs the target data and the classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model classifies the target data based on the classification and grading prompt words to obtain the first classification result and the first grading result of the data to be classified; uses a rule engine with classification and grading rules for each application field to classify and grade the target data to obtain the second classification result and the second grading result of the data to be classified; determines the classification result of the data to be classified according to the first classification result and the second classification result, and determines the grading result of the data to be classified according to the first grading result and the second grading result. The above technical solution obtains the target data corresponding to the data to be classified by generalizing the data to be classified, which can effectively remove redundant information, focus on key information, make the target data more refined, thereby improving the efficiency of data processing, and the generalized target data is more concise, reducing the storage cost and the transmission time, and further improving the overall performance. Then, the target data and the classification and grading prompt words are input into a pre-trained data classification and grading model, so that the data classification and grading model classifies the target data based on the classification and grading prompt words to obtain the first classification result and the first grading result of the data to be classified, which not only reduces manual intervention, reduces labor costs, but also improves the accuracy and efficiency of classification and grading. Using a rule engine with classification and grading rules for each application field to classify and grade the target data to obtain the second classification result and the second grading result of the data to be classified avoids the subjectivity and uncertainty of manual classification and grading, improves the accuracy of classification and grading, and the automated processing of the rule engine reduces the workload of manual classification and grading, reducing labor costs and time costs. Finally, determining the classification result of the data to be classified according to the first classification result and the second classification result, and determining the grading result of the data to be classified according to the first grading result and the second grading result can comprehensively evaluate the data to be classified, reduce the one-sidedness brought by a single classification result and a single grading result, thereby improving the accuracy and reliability of data classification and grading. Therefore, the present invention can ensure the accuracy of data classification and grading while ensuring the efficiency of data classification and grading.
[0012] In a second aspect, an embodiment of the present invention further provides a data classification and grading device, which includes:
[0013] A generalization module for obtaining the target data corresponding to the data to be classified by generalizing the data to be classified;
[0014] A first classification and grading module for inputting the target data and the classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model classifies the target data based on the classification and grading prompt words to obtain the first classification result and the first grading result of the data to be classified;
[0015] A second classification and grading module, configured to classify and grade the target data by using a rule engine with classification and grading rules for each application field built therein, so as to obtain a second classification result and a second grading result of the data to be partitioned;
[0016] A determination module, configured to determine a classification result of the data to be partitioned according to the first classification result and the second classification result, and determine a grading result of the data to be partitioned according to the first grading result and the second grading result.
[0017] In a third aspect, an embodiment of the present invention further provides an electronic device, which includes:
[0018] At least one processor; and a memory communicatively connected to the at least one processor;
[0019] Wherein, the memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to implement any one of the data classification and grading methods in the first aspect.
[0020] In a fourth aspect, an embodiment of the present invention further provides a storage medium containing computer-executable instructions,
[0021] The computer-executable instructions are used to implement any one of the data classification and grading methods in the first aspect when executed by a computer processor.
[0022] In a fifth aspect, the present application provides a computer program product, which includes computer instructions. When the computer instructions run on a computer, the computer is enabled to execute the data classification and grading method provided in the first aspect.
[0023] It should be noted that the above computer instructions may be stored in whole or in part on a computer-readable storage medium. Wherein, the computer-readable storage medium may be packaged together with the processor of the data classification and grading device, or may be separately packaged from the processor of the data classification and grading device, and the present application does not make any limitation thereto.
[0024] The descriptions of the second aspect, the third aspect, the fourth aspect, and the fifth aspect in the present application may refer to the detailed description of the first aspect; and, the beneficial effects of the descriptions of the second aspect, the third aspect, the fourth aspect, and the fifth aspect may refer to the analysis of the beneficial effects of the first aspect, and will not be elaborated herein.
[0025] In this application, the names of the above data classification and grading devices do not constitute limitations on the devices or functional modules themselves. In actual implementation, these devices or functional modules may appear under other names. As long as the functions of each device or functional module are similar to those of this application and fall within the scope of the claims of this application and their equivalent technologies.
[0026] These aspects or other aspects of this application will be more clearly understood in the following description. Brief Description of the Drawings
[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0028] Figure 1 It is a flowchart of a data classification and grading method provided by an embodiment of the present invention;
[0029] Figure 2 It is a flowchart of another data classification and grading method provided by an embodiment of the present invention;
[0030] Figure 3 It is a schematic structural diagram of a data classification and grading device provided by an embodiment of the present invention;
[0031] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed Embodiments
[0032] The following will further elaborate on the present invention in conjunction with the drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the present invention, rather than limiting the present invention. Additionally, it should be noted that for the sake of convenience of description, only parts related to the present invention rather than all structures are shown in the drawings.
[0033] The term "and / or" in this document is only a description of the association relationship between associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.
[0034] The terms "first" and "second" in the specification and drawings of this application are used to distinguish different objects or different processes for the same object, rather than to describe the specific order of the objects.
[0035] In addition, the terms "including" and "having" and any variations thereof mentioned in the description of the present application are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally further include other steps or units not listed, or may optionally further include other steps or units inherent to these processes, methods, products or devices.
[0036] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts depict the operations (or steps) as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it may also have additional steps not included in the figures. The process can correspond to a method, function, procedure, subroutine, subprogram, etc. In addition, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.
[0037] It should be noted that in the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0038] In the description of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more.
[0039] Figure 1 The flowchart of a data classification and grading method provided for the embodiments of the present invention. This embodiment is applicable to situations where data needs to be classified and graded. This method can be executed by a data classification and grading device, and specifically includes the following steps:
[0040] Step 110: Obtain the target data corresponding to the data to be divided by generalizing the data to be divided.
[0041] Among them, the data to be divided refers to the data that needs to be classified and graded. For example, the data to be divided can be a document. The target data refers to the data obtained after generalizing the data to be divided.
[0042] Specifically, first, the data to be partitioned is preprocessed by removing special characters, stop words, etc. to improve the quality of the data to be partitioned. Then, a keyword extraction algorithm (such as TF-IDF or TextRank) is used to extract keywords from the preprocessed data. Next, the keywords can be screened and sorted according to their importance and relevance to obtain a refined keyword list. Finally, the screened keywords are processed based on a pre-trained language model to generate an abstract of the data to be partitioned. Finally, the generated abstract data is vectorized to obtain the target data.
[0043] In this embodiment, by summarizing the data to be partitioned, redundant information can be effectively removed, and the core points in the data can be refined, making the target data more refined and focused on key information, thus significantly improving the efficiency of data processing. Moreover, the summarized target data is more concise, so less resources are required during storage and transmission, which helps to reduce storage costs and transmission time, and further improves the overall performance.
[0044] Step 120: Input the target data and classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model partitions the target data based on the classification and grading prompt words to obtain the first classification result and the first grading result of the data to be partitioned.
[0045] Among them, the classification and grading prompt words refer to the additional information provided by the user for the data classification and grading model, and its function is to guide the model to classify and grade the target data more accurately. For example, the classification and grading prompt words include classification and grading criteria, terms in specific fields, rules, etc. The data classification and grading model refers to a pre-trained large language model used to classify and grade data according to the input data and classification and grading prompt words. The first classification result refers to the classification result obtained by the data classification and grading model after classifying the target data. For example, the first classification result can be finance. The first grading result refers to the grading result obtained by the data classification and grading model after grading the sensitivity of the target data. For example, the first grading result can be level 5.
[0046] Specifically, first obtain the classification and grading prompt words input by the user through an interface (such as a graphical user interface or a network application programming interface), and input the target data and the classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model divides the target data based on the classification and grading prompt words to obtain the first classification result and the first grading result of the data to be divided. Specifically, the data classification and grading model can first determine the classification and grading rules according to the classification and grading prompt words, and extract features from the target data through a self-attention mechanism to obtain target features. Then, the model will divide the target data according to the classification and grading rules and the target features to obtain the first classification result and the first grading result of the data to be divided.
[0047] In this embodiment, using a pre-trained data classification and grading model to classify and grade the target data not only reduces manual intervention and labor costs, but also improves the accuracy and efficiency of classification and grading.
[0048] Step 130: Use a rule engine with classification and grading rules for each application field built in to classify and grade the target data to obtain the second classification result and the second grading result of the data to be divided.
[0049] Among them, the rule engine refers to a system with classification and grading rules for various application fields built in. It can use these rules to process the target data, so as to realize the classification and grading of the data to be divided. The second classification result refers to the classification result obtained after the rule engine classifies the target data. The second grading result refers to the grading result obtained after the rule engine grades the target data.
[0050] Specifically, before using a rule engine with classification and grading rules for each application field built in to classify and grade the target data, first, it is necessary to select a suitable rule engine (such as Drools, Easy Rules, etc.) according to the actual situation or requirements. Then, load the pre-written classification and grading rules into the rule engine to obtain a rule engine with classification and grading rules for each application field built in. Among them, the pre-written classification and grading rules refer to the classification and grading rules defined and written in detail using a specific rule engine language (such as the DRL language of Drools) according to the specific situation and requirements of the actual application field (such as finance, medical care, education, etc.).
[0051] After obtaining the target data, input the target data into the rule engine, so that the rule engine with classification and grading rules for each application field built-in classifies and grades the target data to obtain the second classification result and the second grading result of the data to be divided. Specifically, the rule engine will traverse all the built-in classification and grading rules and check one by one whether the target data meets the conditions of a certain rule. If the target data meets the conditions of a certain rule, classify and grade the target data according to that rule; if the target data meets the conditions of multiple rules at the same time, the rule engine will determine the finally used rule according to the preset rule priority and classify and grade the target data based on the determined rule, so as to obtain the second classification result and the second grading result of the data to be divided.
[0052] In this embodiment, by using the rule engine with classification and grading rules for each application field built-in, accurate matching, classification and grading can be performed on the target data, avoiding the subjectivity and uncertainty of manual classification and grading, improving the accuracy of classification and grading, and the automated processing of the rule engine reduces the workload of manual classification and grading, reducing the labor cost and time cost.
[0053] Step 140, determine the classification result of the data to be divided according to the first classification result and the second classification result, and determine the grading result of the data to be divided according to the first grading result and the second grading result.
[0054] Among them, the classification result refers to the final category to which the data to be divided belongs determined according to the first classification result and the second classification result. The grading result refers to the final level at which the data to be divided is located determined according to the first grading result and the second grading result.
[0055] Specifically, first, determine whether the first classification result and the second classification result are consistent. If the first classification result and the second classification result are inconsistent, determine whether the model weight of the data classification and grading model is greater than the engine weight of the rule engine; if the model weight is greater than the engine weight, determine the first classification result as the classification result; if the model weight is not greater than the engine weight, determine the second classification result as the classification result. If the first classification result and the second classification result are consistent, determine the first classification result or the second classification result as the classification result.
[0056] At the same time, it can be determined whether the first grading result and the second grading result are consistent. If the first grading result and the second grading result are inconsistent, determine the product of the model weight of the data classification and grading model and the first grading result as the first grading weighted result, determine the product of the engine weight of the rule engine and the second grading result as the second grading weighted result, and then determine the sum value of the first grading weighted result and the second grading weighted result as the grading result of the data to be divided. If the first grading result and the second grading result are consistent, determine the first grading result or the second grading result as the grading result.
[0057] In this embodiment, by determining the classification result of the data to be classified according to the first classification result and the second classification result, and determining the grading result of the data to be classified according to the first grading result and the second grading result, a more comprehensive evaluation of the data to be classified can be performed, reducing the one-sidedness brought by a single classification result and a single grading result, thereby improving the accuracy and reliability of data classification and grading.
[0058] The data classification and grading method provided by the embodiment of the present invention includes: obtaining the target data corresponding to the data to be classified by generalizing the data to be classified; inputting the target data and the classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model classifies the target data based on the classification and grading prompt words to obtain the first classification result and the first grading result of the data to be classified; classifying and grading the target data by using a rule engine with classification and grading rules for each application field to obtain the second classification result and the second grading result of the data to be classified; determining the classification result of the data to be classified according to the first classification result and the second classification result, and determining the grading result of the data to be classified according to the first grading result and the second grading result. In the above technical solution, by generalizing the data to be classified to obtain the target data corresponding to the data to be classified, redundant information can be effectively removed, focusing on key information, making the target data more refined, thereby improving the efficiency of data processing, and the generalized target data is more concise, reducing the storage cost and the transmission time, and further improving the overall performance. Then, inputting the target data and the classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model classifies the target data based on the classification and grading prompt words to obtain the first classification result and the first grading result of the data to be classified, not only reduces manual intervention, reduces the labor cost, but also improves the accuracy and efficiency of classification and grading. Using a rule engine with classification and grading rules for each application field to classify and grade the target data to obtain the second classification result and the second grading result of the data to be classified avoids the subjectivity and uncertainty of manual classification and grading, improves the accuracy of classification and grading, and the automated processing of the rule engine reduces the workload of manual classification and grading, reducing the labor cost and time cost. Finally, determining the classification result of the data to be classified according to the first classification result and the second classification result, and determining the grading result of the data to be classified according to the first grading result and the second grading result, a more comprehensive evaluation of the data to be classified can be performed, reducing the one-sidedness brought by a single classification result and a single grading result, thereby improving the accuracy and reliability of data classification and grading. Therefore, the present invention can ensure the accuracy of data classification and grading while ensuring the efficiency of data classification and grading.
[0059] Figure 2The flowchart of another data classification and grading method provided by an embodiment of the present invention. This embodiment is a concretization based on the above embodiment. In this embodiment, the method may further include:
[0060] Step 210: Perform vectorization processing on each data block that makes up the data to be divided, and obtain the data vectors corresponding to each data block.
[0061] Among them, a data block refers to an independent unit that makes up the data to be divided. A data vector refers to a numerical vector representation obtained by converting a data block through a specific vectorization method (such as word embedding, feature extraction, etc.).
[0062] Specifically, a suitable text vectorization method (such as the bag-of-words model, Word2Vec, or FastText) can be selected according to the actual situation or requirements, and then the selected text vectorization method is used to perform vectorization processing on each data block that makes up the data to be divided, and obtain the data vectors corresponding to each data block.
[0063] In this embodiment, by performing vectorization processing on each data block that makes up the data to be divided and obtaining the data vectors corresponding to each data block, not only can the speed and efficiency of subsequent data processing be improved, but also the unified storage, management, and analysis of data are greatly facilitated.
[0064] Step 211: Summarize the clustering data of each cluster obtained by performing soft clustering on each data vector to obtain at least one summary data.
[0065] Among them, soft clustering refers to a clustering method that allows a data point to belong to multiple clusters simultaneously. In soft clustering, each data point has a probability of belonging to each cluster.
[0066] Specifically, first, a suitable soft clustering method (such as fuzzy clustering, Gaussian mixture model, Bayesian Gaussian mixture model, etc.) can be selected according to the actual situation or requirements, and then the selected soft clustering algorithm is used to perform clustering processing on each data vector to obtain the clustering data of multiple clusters. After that, a pre-trained large language model (such as GPT1, GPT2, or BLOOM, etc.) is selected according to the requirements. Finally, the clustering data of each cluster is input into the selected large language model to obtain at least one summary data.
[0067] In this embodiment, by summarizing the clustering data of each cluster obtained by performing soft clustering on each data vector to obtain at least one summary data, not only is the effective compression of the data set realized, but also the complete retention of key information is ensured while reducing the data volume. And in the subsequent data analysis and processing links, using the summary data to replace the original data vectors for calculation can greatly reduce the computational complexity and the amount of processing required, thereby improving the efficiency of data processing.
[0068] Further, step 211 may specifically include: performing soft clustering on each data vector based on a Gaussian mixture model to obtain at least one cluster of clustering data; and summarizing each cluster of clustering data based on a summarization model to obtain at least one summary data.
[0069] Among them, the Gaussian mixture model refers to a probability model. The model parameters of the Gaussian mixture model include preset categories, and the preset categories refer to the categories predefined in the Gaussian mixture model. Each preset category has a corresponding mean vector, covariance matrix, and mixing coefficient. The summary data refers to the data obtained after summarizing the clustering data. The summarization model refers to a large model used to generate summary data. For example, the summarization model can be GPT1, GPT2, or GPT3.
[0070] Specifically, first set the initial values of the model parameters of the Gaussian mixture model (including the number of preset categories, the mean vector, covariance matrix, and mixing coefficient of each preset category) according to the actual situation or requirements. Then, based on the current model parameters and the expectation step of the expectation-maximization (EM) algorithm, calculate the probability values of each data vector belonging to each preset category. Next, use the maximization step of the EM algorithm and the probability values of each data vector belonging to each preset category to update the model parameters, and determine whether the updated model parameters meet the preset parameter convergence condition (such as the difference between the updated model parameters and the pre-updated model parameters is less than a preset threshold). If not, return to execute the step of calculating the probability values of each data vector belonging to each preset category based on the current model parameters and the expectation step of the EM algorithm; if so, divide the data vectors into at least one cluster of clustering data according to the probability values of each data vector belonging to each preset category and a preset probability threshold. Finally, input each cluster of clustering data and a preset summary prompt word into a pre-trained summarization model, so that the summarization model summarizes each cluster of clustering data based on the preset summary prompt word to obtain at least one summary data.
[0071] In addition, in addition to setting the number of preset categories according to the actual situation or requirements, it is also possible to set multiple numbers of preset categories, and then select the optimal number from the multiple numbers of preset categories using the Bayesian information criterion, that is, the number with the smallest Bayesian information criterion is determined as the final number of preset categories. The specific calculation formula of the Bayesian information criterion is:
[0072] BIC = -2ln(L) + k * ln(n)
[0073] Among them, BIC is the Bayesian information criterion, k is the number of preset categories, n is the number of data blocks, and L is the maximum likelihood estimate value of the model on the data.
[0074] In this embodiment, first perform soft clustering on each data vector, which can divide the data into multiple clusters with statistical significance. Then, process the clustering data of each cluster through a summary model, which can extract the key information or features of each cluster, thereby removing redundant information and retaining the information that can best represent the data within the cluster, providing a more concise and efficient data set for subsequent data classification and grading.
[0075] Further, perform soft clustering on each data vector based on the Gaussian mixture model to obtain at least one cluster of clustering data, including: calculating the probability density function values of each data vector being assigned to the Gaussian distributions of each preset category based on the mean vectors and covariance matrices of each preset category; determining the probability values of each data vector being assigned to each preset category according to the product of the probability density function values of each data vector being assigned to the Gaussian distributions of each preset category and the mixing coefficients of each preset category; and dividing the data vectors into at least one cluster of clustering data according to the probability values of each data vector being assigned to each preset category and a preset probability threshold.
[0076] Among them, the mean vector and covariance matrix refer to the parameters describing the shape and position of the Gaussian distribution. The probability density function value refers to the probability density of a certain data point appearing under the Gaussian distribution. The mixing coefficient refers to the weight or contribution degree of each Gaussian distribution in the Gaussian mixture model. The preset probability threshold refers to the probability boundary set in advance according to the actual situation or requirements for determining whether a data vector belongs to a certain preset category. In practical applications, each preset category has a corresponding preset probability threshold, and they can be the same or different. For example: when there are three preset categories, the preset probability thresholds of each preset category can all be 0.5 or can be 0.8, 0.6, and 0.7 respectively.
[0077] Specifically, first calculate the probability density function values of each data vector being assigned to the Gaussian distributions of each preset category based on the mean vectors and covariance matrices of each preset category. The specific calculation formula is as follows:
[0078]
[0079] Among them, x represents the data vector, μ k represents the mean vector of the kth preset category, ∑ k represents the covariance matrix of the kth preset category, d represents the dimension of the data vector. P(x|μ k ,∑ k ) represents the probability density function value of the data vector x in the kth preset category, and T represents the transpose symbol.
[0080] Then, multiply the probability density function values of the Gaussian distributions to which each data vector is assigned to each preset category by the mixing coefficients of each preset category to obtain the unnormalized probability values of each data vector being assigned to each preset category. Next, sum up all the product results to obtain the overall probability density function value. Then, divide the unnormalized probability values of each data vector being assigned to each preset category by the overall probability density function value to obtain the probability values of each data vector being assigned to each preset category. After that, based on the probability values of each data vector being assigned to each preset category, the mixing coefficients, mean vectors, and covariance matrices of each preset category can be updated. The specific update formulas are as follows:
[0081]
[0082] Among them, a k represents the mixing coefficient of the k-th preset category, r ik represents the probability value of the i-th data vector in the k-th preset category, and n represents the total number of data vectors.
[0083]
[0084] Among them, r ik represents the probability value of the i-th data vector in the k-th preset category, and x i represents the i-th data vector.
[0085]
[0086] Then, calculate the differences between the updated mixing coefficients, mean vectors, and covariance matrices of each preset category and their corresponding values before the update, and determine whether all the differences are less than a preset difference threshold. If all the differences are less than the preset difference threshold, then divide the data vectors into at least one cluster of clustering data based on the probability values of each data vector being assigned to each preset category and a preset probability threshold. Otherwise, return to execute the step of calculating the probability density function values of the Gaussian distributions to which each data vector is assigned to each preset category based on the mean vectors and covariance matrices of each preset category.
[0087] Exemplarily, assume there are three preset categories A, B, and C, and four data vectors x1, x2, x3, x4. The preset probability thresholds for A, B, and C are all 0.7. The probability value of x1 for A is 0.8, for B is 0.2, and for C is 0.1. The probability value of x2 for B is 0.75, for A is 0.2, and for C is 0.15. The probability value of x3 for C is 0.9, for A is 0.7, and for B is 0.15. The probability value of x4 for A is 0.4, for B is 0.3, and for C is 0.15. Therefore, the clustering result is: The 1st cluster: contains data vectors x1 and x3; The 2nd cluster: contains data vector x2; The 3rd cluster: contains data vector x3.
[0088] In this embodiment, through the above steps, the accuracy and flexibility of clustering are improved.
[0089] Furthermore, based on the summary model, summarize the clustering data of each cluster to obtain at least one summary data, including: input the clustering data of each cluster and the preset summary prompt words into the summary model, so that the summary model summarizes the clustering data of each cluster based on the preset summary prompt words to obtain at least one summary data.
[0090] Among them, the preset summary prompt words refer to the prompt words set in advance according to the actual situation or requirements and used in the summary model to guide the summary process.
[0091] Specifically, after obtaining the clustering data of each cluster, input the clustering data of each cluster and the preset summary prompt words into the summary model, so that the summary model summarizes the clustering data of each cluster based on the preset summary prompt words to obtain at least one summary data. Specifically, the summary model first identifies the core features of each cluster according to the input data (including the preset prompt words and clustering data). Then, according to the indication of the preset summary prompt words, automatically extract key information (such as the distribution of the cluster, the central tendency of the data, the variability represented by the standard deviation, etc.). Finally, the summary model generates at least one summary data based on the key information and the core features of each cluster.
[0092] In this embodiment, through the above steps, the accuracy and efficiency of summarization are improved.
[0093] Step 212, in the case where it is determined that the data length of the summary data exceeds the preset data length, determine the summary data as the target data.
[0094] Among them, the preset data length refers to the data length threshold set in advance according to the actual situation or requirements.
[0095] Specifically, after obtaining the summary data, determine whether there is at least one summary data whose data length exceeds the preset data length. If there is at least one summary data whose data length exceeds the preset data length, then determine all the summary data as the target data. If the data lengths of all the summary data do not exceed the preset threshold, continue to perform soft clustering processing on all the summary data, and summarize the clustered data again to obtain at least one summary data.
[0096] In this embodiment, through the above steps, important information can be prevented from being lost due to excessive simplification, ensuring the accuracy of subsequent data classification and grading.
[0097] Step 213: Input the target data and the classification and grading prompt words into the pre-trained data classification and grading model, so that the data classification and grading model divides the target data based on the classification and grading prompt words to obtain the first classification result and the first grading result of the data to be divided.
[0098] Specifically, the training process of the data classification and grading model is as follows: First, collect the data classification and grading result sets of each application field and use them as the training sample set. Then, formulate the data classification and grading rules for each application field according to the actual situation or the requirements of a specific application scenario. Then, select a suitable pre-trained large language model according to the requirements (such as ChatGLM, PaLM, LLaMa, etc.), and import the knowledge base containing data classification and grading domain knowledge into the selected large language model so that it learns the data classification and grading domain knowledge, thereby realizing the construction of the data classification and grading model. Then, input the previously obtained training sample set and the data classification and grading rules of each application field into the constructed data classification and grading model, so that the data classification and grading model calculates the loss function value by comparing its output with the expected results in the training sample set, and then iteratively adjusts the model parameters according to the loss function value until the preset stop condition is reached (such as reaching the maximum number of iterations). After completing the above training process, the model can be further fine-tuned according to the loss function value of the last time (such as full parameter fine-tuning, Adapter fine-tuning, incremental learning, etc.) to obtain the finally trained data classification and grading model.
[0099] In practical applications, the training process of the summary model is similar to that of the data classification and grading model. First, collect the data summary result sets in each application field as the training set. Then, select a suitable pre-trained large language model according to the requirements, and import the knowledge base of the data summary field into the selected large language model so that it can learn the knowledge of the data summary field, thereby realizing the construction of the summary model. Next, input the aforementioned obtained training set into the constructed summary model, so that the summary model calculates the loss function value by comparing its output with the expected results in the training set, and then iteratively adjusts the model parameters according to the loss function value until the preset stop condition is reached. After completing the aforementioned training process, the model can be further fine-tuned according to the loss function value of the last time to obtain the finally trained summary model.
[0100] It should be noted that after obtaining the data classification and grading result sets in each application field, the data can be screened first using a two-stage strategy, and the screened data is determined as the training sample set. The specific implementation process is as follows:
[0101] First, calculate the zero-shot score of the preselected pre-trained large language model for the preset task set. The specific calculation formula is: S 0 = g(T, LLM); where S 0 is the zero-shot score, T is the preset task set, and LLM is the preselected pre-trained large language model.
[0102] Next, determine the entries in the data classification and grading result set in each application field, and determine the one-shot score of each entry based on the pre-trained large language model. The specific calculation formula is: S 1 = g(P i , T, LLM); where S 1 is the one-shot score, and P i is an entry in the data classification and grading result set of a certain application field.
[0103] Then calculate the gold score of the entries in the data classification and grading result set in each application field. The specific calculation formula is: G = |S 1 - S 0 |; where G is the gold score.
[0104] Finally, use the data entries with a gold score exceeding the preset score threshold as the training sample set. In addition, if there are at least two data classification and grading result sets in the same application field, after obtaining the gold scores of all data, select the data classification and grading result set with the highest score as the "gold subset" and use it as part of the training sample set. The selection process can be represented by the following formula: where G subsetis the golden subset, and argmax represents selecting the set of prompts that maximizes the golden score.
[0105] Step 214: Use a rule engine with classification and grading rules for each application field built-in to classify and grade the target data, obtaining a second classification result and a second grading result of the data to be partitioned.
[0106] Step 215: Determine the classification result of the data to be partitioned based on the first classification result and the second classification result, and determine the grading result of the data to be partitioned based on the first grading result and the second grading result.
[0107] Furthermore, determining the classification result of the data to be partitioned based on the first classification result and the second classification result includes: in the case where it is determined that the first classification result and the second classification result are inconsistent, obtaining a comparison result by comparing the model weight of the data classification and grading model and the engine weight of the rule engine; if the comparison result is that the model weight is greater than the engine weight, then determine the first classification result as the classification result; if the comparison result is that the model weight is not greater than the engine weight, then determine the second classification result as the classification result.
[0108] Among them, the model weight of the data classification and grading model refers to the weight assigned to the data classification and grading model in advance according to the actual situation or requirements. The engine weight of the rule engine refers to the weight assigned to the rule engine in advance according to the actual situation or requirements.
[0109] Exemplarily, if the first classification result is: medical technology document, the second classification result is: artificial intelligence technology document, the model weight of the data classification and grading model is 0.55, and the engine weight of the rule engine is 0.45, then the classification result is a medical technology document.
[0110] In this embodiment, through the above steps, the accuracy of the classification result is improved.
[0111] Furthermore, determining the grading result of the data to be partitioned based on the first grading result and the second grading result includes: in the case where it is determined that the first grading result and the second grading result are inconsistent, determining the product of the model weight of the data classification and grading model and the first grading result as the first grading weighted result, and determining the product of the engine weight of the rule engine and the second grading result as the second grading weighted result; determining the sum value of the first grading weighted result and the second grading weighted result as the grading result of the data to be partitioned.
[0112] Among them, the grading weighted result refers to the result obtained by multiplying the grading result by the corresponding weight (such as the model weight or the engine weight).
[0113] Exemplarily, if the first classification result is: 4, the second classification result is: 3, the model weight of the data classification and grading model is 0.6, and the engine weight of the rule engine is 0.4, then the classification result is 3.6.
[0114] In this embodiment, through the above steps, the accuracy of the classification result is improved.
[0115] The data classification and grading method provided by the embodiment of the present invention first performs vectorization processing on each data block constituting the data to be divided to obtain data vectors corresponding to each data block, which not only improves the speed and efficiency of subsequent data processing, but also greatly facilitates the unified storage, management, and analysis of data. Then, by summarizing the cluster data obtained by soft clustering of each data vector, at least one summary data is obtained, which not only realizes the effective compression of the data set, but also ensures the complete retention of key information while reducing the data volume. And in the subsequent data analysis and processing links, using the summary data to replace the original data vector for calculation can greatly reduce the calculation complexity and the amount of processing required, thereby improving the efficiency of data processing. When it is determined that the data length of the summary data exceeds the preset data length, the summary data is determined as the target data, which can prevent important information from being lost due to excessive reduction, and ensure the accuracy of subsequent data classification and grading. Then, the target data and the classification and grading prompt words are input into the pre-trained data classification and grading model, so that the data classification and grading model divides the target data based on the classification and grading prompt words to obtain the first classification result and the first grading result of the data to be divided, which not only reduces manual intervention, reduces labor costs, but also improves the accuracy and efficiency of classification and grading. Using a rule engine with classification and grading rules for each application field to classify and grade the target data to obtain the second classification result and the second grading result of the data to be divided, avoiding the subjectivity and uncertainty of manual classification and grading, improving the accuracy of classification and grading, and the automated processing of the rule engine reduces the workload of manual classification and grading, reducing labor costs and time costs. Finally, the classification result of the data to be divided is determined according to the first classification result and the second classification result, and the grading result of the data to be divided is determined according to the first grading result and the second grading result, which can comprehensively evaluate the data to be divided, reduce the one-sidedness brought by a single classification result and a single grading result, and thus improve the accuracy and reliability of data classification and grading. Therefore, the present invention can ensure the accuracy of data classification and grading while guaranteeing the efficiency of data classification and grading.
[0116] In addition, soft clustering of each data vector based on the Gaussian mixture model can divide the data into multiple statistically significant clusters. Then, by processing the clustered data of each cluster through the summarization model, the key information or features of each cluster can be refined, thereby removing redundant information and retaining the information that best represents the data within the cluster, providing a more concise and efficient dataset for subsequent data classification and grading. Moreover, the probability density of the data vector in the Gaussian distribution of each preset category is calculated based on the mean vector and covariance matrix of each preset category, and then combined with the mixing coefficient of the preset category to determine the probability value of the data vector belonging to each category. Then, according to the probability value and the preset threshold, the data vector is divided into at least one cluster, improving the accuracy and flexibility of clustering. Next, the clustered data of each cluster and the preset summarization prompt words are input into the summarization model, so that the summarization model summarizes the clustered data of each cluster based on the preset summarization prompt words to obtain at least one summary data, improving the accuracy and efficiency of summarization.
[0117] On the other hand, in the case where the two classification results are inconsistent, if the model weight is greater than the engine weight, the first classification result is determined as the classification result; if the model weight is not greater than the engine weight, the second classification result is determined as the classification result, improving the accuracy of the classification result. In the case where the two grading results are inconsistent, the sum value of the two graded weighted results is determined as the grading result, improving the accuracy of the grading result.
[0118] Figure 3 FIG. is a schematic structural diagram of a data classification and grading device provided by an embodiment of the present invention. This device and the data classification and grading method of the above embodiments belong to the same inventive concept. For the details not described in detail in the embodiment of the data classification and grading device, reference can be made to the embodiments of the above data classification and grading method.
[0119] As Figure 3 shown, the device includes:
[0120] A generalization module 310, configured to obtain target data corresponding to the data to be divided by generalizing the data to be divided;
[0121] A first classification and grading module 320, configured to input the target data and classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model divides the target data based on the classification and grading prompt words to obtain a first classification result and a first grading result of the data to be divided;
[0122] A second classification and grading module 330, configured to classify and grade the target data by using a rule engine that incorporates classification and grading rules for each application field to obtain a second classification result and a second grading result of the data to be divided;
[0123] A determination module 340 is configured to determine the classification result of the data to be partitioned according to the first classification result and the second classification result, and determine the grading result of the data to be partitioned according to the first grading result and the second grading result.
[0124] Based on the above embodiments, a summarization module 310 is specifically configured to: perform vectorization processing on each data block constituting the data to be partitioned to obtain data vectors corresponding to the data blocks; summarize the cluster clustering data obtained by performing soft clustering on the data vectors to obtain at least one summary data; and when it is determined that the data length of the summary data exceeds a preset data length, determine the summary data as the target data.
[0125] Based on the above embodiments, summarizing the cluster clustering data obtained by performing soft clustering on the data vectors to obtain at least one summary data includes: performing soft clustering on the data vectors based on a Gaussian mixture model to obtain at least one cluster of clustering data; and summarizing the cluster clustering data based on a summary model to obtain at least one summary data.
[0126] Based on the above embodiments, performing soft clustering on the data vectors based on a Gaussian mixture model to obtain at least one cluster of clustering data includes: calculating the probability density function values of the Gaussian distributions to which the data vectors are assigned to the preset categories based on the mean vectors and covariance matrices of the preset categories; determining the probability values of the data vectors being assigned to the preset categories according to the product of the probability density function values of the Gaussian distributions to which the data vectors are assigned to the preset categories and the mixing coefficients of the preset categories; and partitioning the data vectors into at least one cluster of clustering data according to the probability values of the data vectors being assigned to the preset categories and a preset probability threshold.
[0127] Based on the above embodiments, summarizing the cluster clustering data based on a summary model to obtain at least one summary data includes: inputting the cluster clustering data and a preset summary prompt word into the summary model, so that the summary model summarizes the cluster clustering data based on the preset summary prompt word to obtain at least one summary data.
[0128] Based on the above embodiments, determining the classification result of the data to be classified according to the first classification result and the second classification result includes: when it is determined that the first classification result and the second classification result are inconsistent, determining the product of the model weight of the data classification and grading model and the first classification result as the first classification weighted result, and determining the product of the engine weight of the rule engine and the second classification result as the second classification weighted result; determining the sum value of the first classification weighted result and the second classification weighted result as the classification result of the data to be classified.
[0129] Based on the above embodiments, determining the classification result of the data to be classified according to the first classification result and the second classification result includes: when it is determined that the first classification result and the second classification result are inconsistent, obtaining a comparison result by comparing the model weight of the data classification and grading model and the engine weight of the rule engine; if the comparison result is that the model weight is greater than the engine weight, determining the first classification result as the classification result; if the comparison result is that the model weight is not greater than the engine weight, determining the second classification result as the classification result.
[0130] The data classification and grading device provided by the embodiments of the present invention can execute the data classification and grading method provided by any embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method.
[0131] It should be noted that in the embodiments of the above data classification and grading device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of the functional units are only for the convenience of mutual distinction and do not limit the protection scope of the present invention.
[0132] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Figure 4 It shows a block diagram of an exemplary electronic device 4 suitable for implementing the embodiments of the present invention. Figure 4 The shown electronic device 4 is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present invention.
[0133] As Figure 4 shown, the electronic device 4 is presented in the form of a general-purpose computing electronic device. The components of the electronic device 4 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).
[0134] Bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, an Accelerated Graphics Port, a processor bus, or a local bus using any of a variety of bus architectures. By way of example, and not limitation, these architectures include Industry Standard Architecture (ISA) buses, Micro Channel Architecture (MAC) buses, Enhanced ISA buses, Video Electronics Standards Association (VESA) local buses, and Peripheral Component Interconnect (PCI) buses.
[0135] Electronic device 4 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by electronic device 4, including both volatile and nonvolatile media, removable and non-removable media.
[0136] System memory 28 can include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Electronic device 4 can further include other removable / non-removable, volatile / nonvolatile computer system storage media. By way of example only, storage system 34 can be used for reading and writing non-removable, nonvolatile magnetic media ( Figure 4 not shown and typically called a "hard disk drive"). Although Figure 4 not shown in the figures, a disk drive for reading and writing removable nonvolatile disks (such as a "floppy disk"), and an optical disk drive for reading and writing removable nonvolatile optical disks (such as a CD-ROM, DVD-ROM, or other optical media) can be provided. In these instances, each drive can be connected to bus 18 by one or more data media interfaces. System memory 28 can include at least one program product having a set (e.g., at least one) of program modules that are configured to carry out the functions of embodiments of the present invention.
[0137] A program / utility 40 having a set (at least one) of program modules 42 can be stored, for example, in system memory 28, such program modules 42 including, but not limited to, an operating system, one or more application programs, other program modules, and program data, each of which examples or some combination thereof may include an implementation of a networking environment. Program modules 42 generally carry out the functions and / or methods of the embodiments described herein.
[0138] The electronic device 4 can also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 4, and / or communicate with any device that enables the electronic device 4 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication can be carried out through the input / output (I / O) interface 22. Moreover, the electronic device 4 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through the network adapter 20. As Figure 4 shown, the network adapter 20 communicates with other modules of the electronic device 4 through the bus 18. It should be understood that although Figure 4 not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 4, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0139] The processing unit 16 executes various functional applications and page displays by running programs stored in the system memory 28. For example, it implements the data classification and grading method provided by the embodiments of the present invention. The method includes: obtaining target data corresponding to the data to be classified by generalizing the data to be classified; inputting the target data and classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model divides the target data based on the classification and grading prompt words to obtain a first classification result and a first grading result of the data to be classified; using a rule engine with classification and grading rules for each application field to classify and grade the target data to obtain a second classification result and a second grading result of the data to be classified; determining the classification result of the data to be classified according to the first classification result and the second classification result, and determining the grading result of the data to be classified according to the first grading result and the second grading result.
[0140] Of course, those skilled in the art can understand that the processor can also implement the technical solutions of the data classification and grading method provided by any embodiment of the present invention.
[0141] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements, for example, the data classification and grading method provided by the embodiment of the present invention. The method includes: obtaining target data corresponding to the data to be classified by generalizing the data to be classified; inputting the target data and classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model classifies the target data based on the classification and grading prompt words to obtain a first classification result and a first grading result of the data to be classified; using a rule engine with classification and grading rules for each application field to classify and grade the target data to obtain a second classification result and a second grading result of the data to be classified; determining the classification result of the data to be classified according to the first classification result and the second classification result, and determining the grading result of the data to be classified according to the first grading result and the second grading result.
[0142] The computer storage medium of the embodiment of the present invention can adopt any combination of one or more computer-readable media. The computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. The computer-readable storage medium can be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the computer-readable storage medium include: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this document, the computer-readable storage medium can be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device.
[0143] The computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries the computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and the computer-readable medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device.
[0144] The program code included on the computer-readable medium can be transmitted by any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination of the above.
[0145] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0146] Those of ordinary skill in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. Optionally, they can be implemented with program code executable by a computer device, so that they can be stored in a storage device and executed by the computing device, or they can be separately made into individual integrated circuit modules, or multiple of them or steps can be made into a single integrated circuit module for implementation. Thus, the present invention is not limited to any specific combination of hardware and software.
[0147] In addition, in the technical solution of the present invention, the acquisition, storage, use, processing, etc. of data all comply with the relevant provisions of national laws and regulations.
[0148] Note that the above is only the preferred embodiment of the present invention and the technical principles applied. Those skilled in the art will understand that the present invention is not limited to the specific embodiments here. Various obvious changes, re-adjustments, and substitutions can be made by those skilled in the art without departing from the protection scope of the present invention. Therefore, although the present invention has been described in more detail through the above embodiments, the present invention is not limited to the above embodiments. Without departing from the concept of the present invention, it may include more other equivalent embodiments, and the scope of the present invention is determined by the scope of the appended claims.
Claims
1. A data classification and grading method, characterized in that: include: By summarizing the data to be divided, target data corresponding to the data to be divided is obtained; Inputting the target data and the classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model divides the target data based on the classification and grading prompt words to obtain a first classification result and a first grading result of the data to be divided; Classify and classify the target data using a rule engine with built-in classification and grading rules for each application field to obtain a second classification result and a second grading result of the data to be classified; The classification result of the data to be divided is determined according to the first classification result and the second classification result, and the grading result of the data to be divided is determined according to the first grading result and the second grading result.
2. The data classification and grading method according to claim 1, characterized in that: By summarizing the data to be divided, the target data corresponding to the data to be divided is obtained, including: Performing vectorization processing on each data block constituting the data to be divided to obtain a data vector corresponding to each data block; Summarizing each cluster data obtained by performing soft clustering on each of the data vectors to obtain at least one summary data; In the case where it is determined that the data length of the summary data exceeds the preset data length, the summary data is determined as the target data.
3. The data classification and grading method according to claim 2, characterized in that: Summarizing the cluster data obtained by performing soft clustering on the data vectors to obtain at least one summary data, including: Performing soft clustering on each of the data vectors based on a Gaussian mixture model to obtain at least one cluster of clustered data; The cluster data are summarized based on the summary model to obtain at least one summary data.
4. The data classification and grading method according to claim 3, characterized in that: Soft clustering is performed on each of the data vectors based on a Gaussian mixture model to obtain at least one cluster of clustered data, including: Based on the mean vector and covariance matrix of each preset category, calculating the probability density function value of the Gaussian distribution of each of the data vectors being divided into the preset categories; Determine the probability value of each of the data vectors being classified into each of the preset categories according to the product of the probability density function value of the Gaussian distribution of each of the preset categories and the mixing coefficient of each of the preset categories; The data vectors are divided into at least one cluster of clustered data according to the probability value of each of the data vectors being divided into each of the preset categories and a preset probability threshold.
5. The data classification and grading method according to claim 3, characterized in that: Summarizing the cluster data based on the summary model to obtain at least one summary data, including: The cluster data and the preset summary prompt words are input into the summary model, so that the summary model summarizes the cluster data based on the preset summary prompt words to obtain at least one summary data.
6. The data classification and grading method according to claim 1, characterized in that: Determining a classification result of the data to be classified according to the first classification result and the second classification result includes: In the case where it is determined that the first grading result and the second grading result are inconsistent, the product of the model weight of the data classification and grading model and the first grading result is determined as a first grading weighted result, and the product of the engine weight of the rule engine and the second grading result is determined as a second grading weighted result; The sum of the first classification weighted result and the second classification weighted result is determined as the classification result of the data to be divided.
7. The data classification and grading method according to claim 1, characterized in that: Determining the classification result of the data to be divided according to the first classification result and the second classification result includes: When it is determined that the first classification result and the second classification result are inconsistent, obtaining a comparison result by comparing the model weight of the data classification and grading model with the engine weight of the rule engine; If the comparison result is that the model weight is greater than the engine weight, determining the first classification result as the classification result; If the comparison result is that the model weight is not greater than the engine weight, the second classification result is determined as the classification result.
8. A data classification and grading device, characterized in that: include: A summarizing module, used for summarizing the data to be divided to obtain target data corresponding to the data to be divided; A first classification and grading module is used to input the target data and the classification and grading prompt words into a pre-trained data classification and grading model, so that the data classification and grading model divides the target data based on the classification and grading prompt words to obtain a first classification result and a first grading result of the data to be divided; A second classification and grading module is used to classify and grade the target data using a rule engine with built-in classification and grading rules for each application field, to obtain a second classification result and a second grading result of the data to be divided; A determination module is used to determine a classification result of the data to be divided according to the first classification result and the second classification result, and to determine a grading result of the data to be divided according to the first grading result and the second grading result.
9. An electronic device, characterized in that: The computer device comprises: at least one processor; and a memory communicatively coupled to the at least one processor; Wherein, the memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the data classification and grading method described in any one of claims 1-7.
10. A storage medium containing computer executable instructions, characterized in that: The computer executable instructions are used to execute the data classification and grading method described in any one of claims 1-7 when executed by a computer processor.