Data processing method and device and electronic equipment
By combining clustering and large models, policy document categories can be automatically identified, solving the problem of low interpretation efficiency by professionals and achieving efficient and accurate identification in a multilingual environment.
Patent Information
- Application Number
- CN202510864996.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-09-26
AI Technical Summary
In existing technologies, the category identification of policy documents needs to rely on professional interpretation, resulting in low efficiency and lack of accuracy.
By obtaining the characteristic data of the target file, cluster analysis is performed, and the file categories are generated using a large model, and the recognition accuracy is improved by combining sample data.
The accuracy and efficiency of automated recognition of policy document categories have been improved, especially in multilingual environments, enabling flexible processing of policy documents in different languages.
Smart Images

Figure CN120705324A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of natural language processing technology, and in particular to a data processing method, device and electronic device. Background Art
[0002] The release of a policy document will affect changes in the attributes of content objects of related types. For example, the release of a tax policy will affect changes in the tax attributes of related tax types.
[0003] Currently, it is necessary for professionals to interpret policy documents to identify the type of content objects and to identify the category of policy documents. Summary of the Invention
[0004] In view of this, the present application provides a data processing method, device, and electronic device as follows:
[0005] A data processing method, comprising:
[0006] Obtain characteristic data of the target file;
[0007] Clustering the characteristic data of the target file and the characteristic data of at least one sample file to obtain a target cluster, wherein the target cluster includes the characteristic data of the target file;
[0008] Determining example data based on the target cluster, the example data including feature data of at least one target sample file in the target cluster and a category label of the target sample file;
[0009] The characteristic data of the target file and the example data are input into a large model, and the large model is guided to generate a category of the target file.
[0010] The above method preferably further comprises:
[0011] Acquire an original sample data set, wherein the original sample data includes original data items and category labels of the sample files, and the original data items include titles, publishing source identifiers, and content;
[0012] Unifying the languages of the original data items and category labels of all the sample files into a target language, where the target language is the language with the largest number of sample files;
[0013] Obtaining content summaries of each of the sample files respectively;
[0014] For each of the sample files, the target data items of the sample files are spliced to obtain the original features of the sample files, wherein the target data items include the title in a unified language, the source identifier, and the content summary;
[0015] Vectorized encoding is performed on the original features of each of the sample files to obtain feature data of the sample files.
[0016] In the above method, preferably, obtaining characteristic data of the target file includes:
[0017] Obtaining original data items of the target file;
[0018] If the language of the original data item of the target file is not the target language, translating the original data item of the target file into the target language;
[0019] Obtaining a content summary of the target file;
[0020] splicing target data items of the target file to obtain original features of the target file;
[0021] Vectorized encoding is performed on the original features of the target file to obtain feature data of the target file.
[0022] Preferably, the above method clusters the characteristic data of the target file and the characteristic data of at least one sample file to obtain a target cluster, including:
[0023] Get the total number of categories;
[0024] Taking the total number of the categories as the number of clusters, clustering the feature data of the target file and the feature data of at least one of the sample files to obtain a plurality of clusters;
[0025] The cluster to which the characteristic data of the target file belongs is used as the target cluster.
[0026] In the above method, preferably, clustering the characteristic data of the target file and the characteristic data of at least one sample file using the total number of categories as the number of clusters includes:
[0027] After principal component analysis dimensionality reduction is performed on the feature data of the target file and the feature data of at least one of the sample files, the feature data of the target file and the feature data of at least one of the sample files are clustered using the total number of categories as the number of clusters.
[0028] The above method preferably determines example data based on the target cluster, including:
[0029] Respectively obtaining feature similarities between feature data of each cluster sample file and feature data of the target file, wherein the cluster sample file is the sample file in the target cluster;
[0030] Based on feature similarity, obtaining at least one target sample file from the target cluster;
[0031] The sample data is obtained based on at least one of the target sample files.
[0032] The above method preferably obtains at least one target sample file from the target cluster based on feature similarity, including:
[0033] Sort the sample files of each cluster according to feature similarity from high to low to obtain a target file sequence;
[0034] A cluster sample file that meets a similarity condition is selected as the target sample file. The similarity condition includes that the feature similarity with the target file is greater than a similarity threshold, and / or the rank in the target file sequence is less than a rank threshold.
[0035] In the above method, preferably, the data processing method further comprises:
[0036] For each category label, a random sample file of an example number is randomly obtained; wherein the example data also includes feature data of the random sample file and the corresponding category label.
[0037] A data processing device, comprising:
[0038] A first feature acquisition unit, configured to acquire feature data of a target file;
[0039] a feature clustering unit, configured to cluster the feature data of the target file and the feature data of at least one sample file to obtain a target cluster, wherein the target cluster includes the feature data of the target file;
[0040] An example acquisition unit, configured to determine example data based on the target cluster, wherein the example data includes feature data of at least one target sample file in the target cluster and a category label of the target sample file;
[0041] The model prediction unit is used to input the characteristic data of the target file and the example data into the large model, and guide the large model to generate the category of the target file.
[0042] An electronic device comprises: a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the data processing method described above.
[0043] It can be seen from the above technical solution that the embodiments of the present application provide a data processing method, device, and electronic device, which obtain feature data of a target file, cluster the feature data of the target file and the feature data of at least one sample file to obtain a target cluster, determine example data based on the target cluster, input the feature data of the target file and the example data into a large model, and guide the large model to generate a category of the target file. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0045] Figure 1 A flowchart of a data processing method provided in an embodiment of the present application;
[0046] Figure 2 A flow chart of a method for obtaining characteristic data provided in an embodiment of the present application;
[0047] Figure 3 A flowchart of another data processing method provided in an embodiment of the present application;
[0048] Figure 4 A flowchart of another data processing method provided in an embodiment of the present application;
[0049] Figure 5 A flowchart of another data processing method provided in an embodiment of the present application;
[0050] Figure 6 A schematic diagram of a specific implementation flow of the data processing method provided in an embodiment of the present application;
[0051] Figure 7 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application;
[0052] Figure 8 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0053] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0054] The present invention provides a data processing method, device, and electronic device for improving the accuracy of automatic identification of file categories. The data processing method of the present invention is described in detail below with reference to the accompanying drawings.
[0055] Reference Figure 1 , Figure 1 A flow chart of a data processing method provided in an embodiment of the present application is shown as follows: Figure 1 As shown, this method includes steps 101 to 104, as follows:
[0056] Step 101: Obtain feature data of a target file.
[0057] In this embodiment, the feature data of the target file is obtained by performing feature engineering on at least one raw data item of the target file. The feature engineering includes, but is not limited to, solvent data cleaning, normalization, and vectorized encoding. For example, the file data of the target file includes a file identifier and multiple raw data items. The feature vector of the target file is obtained by concatenating the multiple raw data items and then performing vectorized encoding. The feature data includes the feature vector.
[0058] Step 102: Cluster the characteristic data of the target file and the characteristic data of at least one sample file to obtain a target cluster.
[0059] In this embodiment, the target cluster includes the feature data of the target file, that is, the target cluster is the cluster to which the feature data of the target file belongs, wherein the feature data of each sample file is obtained by performing feature engineering processing on at least one original data item of the sample file, and the method for obtaining the feature data of each sample file and the target file is the same.
[0060] In this embodiment, the clustering method for clustering the feature data of the target file and the feature data of at least one sample file includes but is not limited to the K-means clustering algorithm and the K-center point clustering algorithm, etc., and the purpose is to divide the feature data of the target file and the at least one sample file to obtain multiple clusters, where the feature data in the same cluster has a high similarity, and the feature data in different clusters has a low similarity.
[0061] It should be noted that by clustering, the feature data of the target file is divided into multiple categories. The feature vectors of files belonging to the same cluster are more likely to belong to the same category, that is, the target file and the sample files in the target cluster are more likely to belong to the same category.
[0062] Step 103: Determine sample data based on the target cluster.
[0063] In this embodiment, the sample data includes feature data of at least one target sample file in the target cluster and a category label of the target sample file.
[0064] In this embodiment, feature data of at least one target sample file is selected from the target cluster, and the selection method includes random selection or selection based on feature similarity. For example, feature data of a target sample file whose similarity to the feature data of the target file is greater than a similarity threshold is selected.
[0065] Step 104: Input the characteristic data and sample data of the target file into the large model to guide the large model to generate the category of the target file.
[0066] In this embodiment, the large model is constructed based on a diversity example learning prompt project. The large model can learn the feature data of the sample data in the example data, and under the guidance of the example data, generate a classification result that matches the feature data of the target file, that is, the category of the target file.
[0067] It can be seen from the above technical solution that the data processing method provided by the embodiment of the present application obtains the feature data of the target file, clusters the feature data of the target file and the feature data of at least one sample file to obtain a target cluster, determines example data based on the target cluster, inputs the feature data of the target file and the example data into the large model, and guides the large model to generate the category of the target file, wherein the target cluster includes the feature data of the target file, and the example data includes the feature data of at least one target sample file in the target cluster and the category label of the target sample file. Since the example data includes the feature data and category label of the target sample file that belongs to the same cluster as the target file, the feature similarity between the target file and the target sample file is high, and the feature similarity with the features in other clusters is low. It can be seen that by inputting example data, the learning ability and sensitivity of the large model to the similar feature data of the target file are improved, thereby strengthening the guidance of the large model in predicting the category of the target file, thereby improving the accuracy of the category prediction.
[0068] Furthermore, the present application also includes the step of obtaining characteristic data of multiple sample files, see Figure 2 , Figure 2 A flow chart of a method for obtaining characteristic data provided in an embodiment of the present application is shown as follows: Figure 2 As shown, this method specifically includes steps 201 to 205.
[0069] Step 201: Obtain an original sample data set.
[0070] In this embodiment, the original sample data includes the original data items and category labels of the sample file. The original data items include the title, the publication source identifier, and the content. The title may include the file title and subtitles at various levels. The publication source identifier includes the name of the publisher, for example, the name of the issuing organization. The content includes the file content, for example, the file text.
[0071] Step 202: Unify the languages of the original data items and category labels of all sample files into the target language.
[0072] In this embodiment, the target language is the language with the largest number of sample files. Specifically, language identification is performed on the raw data items and category labels of all sample files to determine the language of each sample file. Generally, language identification of a single data item in the title can be used to determine the language of a sample file. Furthermore, the number of sample files corresponding to each language is determined. After the language of the sample files with the largest number is selected as the target language, the language of the raw data items and category labels of all sample files is unified as the target language.
[0073] In this embodiment, the method for unifying the language of the original data items and category labels of all sample files into the target language includes: for sample files in a non-target language, inputting the original data items and category labels of the sample files into a translation macro model, and obtaining the original data items and category labels in the target language output by the translation macro model, wherein the translation macro model is a macro model used for multilingual translation.
[0074] Step 203: Obtain the content summary of each sample file.
[0075] In this embodiment, a summary algorithm or a summary macro model is used to obtain a content summary of each sample file, wherein the summary macro model is a macro model used to summarize and generalize text. The number of characters in the content summary can be pre-configured to make the number of characters in the content summary of each sample file uniform.
[0076] Step 204: For each sample file, the target data items of the sample file are spliced together to obtain the original features of the sample file.
[0077] In this embodiment, the target data item includes the title, publishing source identifier and content summary after unifying the language, that is, the target data item includes the title, publishing source identifier and content summary in the target language.
[0078] Optionally, the method of splicing the target data items of the sample file includes sequentially splicing the title, the publication source identifier, and the content summary to obtain the original features of the sample file.
[0079] Step 205: vectorize and encode the original features of each sample file to obtain feature data of the sample file.
[0080] In this embodiment, vectorized encoding is performed using a large encoding model for the original features of each sample file to obtain a feature vector of the sample file, that is, the feature data includes the feature vector.
[0081] It can be seen from the above technical solution that the characteristic data of the sample file is obtained by quantizing the original feature vector of the sample file. The original feature of the sample file is obtained by splicing the title, publication source identifier and content summary. Therefore, the characteristic data of the sample file is used to characterize the content characteristics of the sample file. Moreover, since the title, publication source identifier and content summary are unified in language and summarized, the consistency and standardization of the characteristic data of the sample file are improved, thereby improving the efficiency and accuracy of subsequent data processing.
[0082] Based on the above embodiment, see Figure 3 , Figure 3 A flowchart of another data processing method provided in an embodiment of the present application is shown below. Figure 3 The specific implementation method of step 101, i.e. obtaining the characteristic data of the target file, is shown as follows: Figure 3 As shown, this method specifically includes steps 301 to 306. These steps are described in detail below. Obtaining the characteristic data of the target file includes:
[0083] Step 301: Obtain the original data items of the target file.
[0084] In this embodiment, the original data items include a title, a publishing source identifier, and content.
[0085] Step 302 : Determine whether the language of the original data item of the target file is the target language. If so, directly execute step 304 . If not, execute step 303 and then step 304 .
[0086] In this embodiment, the target language is the language with the largest number of sample files.
[0087] Step 303: If the language of the original data item of the target file is not the target language, translate the original data item of the target file into the target language.
[0088] In this embodiment, for a target file in a non-target language, the original data items of the target file are input into the translation macro model, and the original data items in the target language are output by the translation macro model.
[0089] Step 304: Obtain a content summary of the target file.
[0090] In this embodiment, a summary algorithm or a summary model is used to obtain a content summary of each sample file.
[0091] Step 305: Concatenate the target data items of the target file to obtain the original features of the target file.
[0092] In this embodiment, the target data item includes the title, publishing source identifier and content summary after unifying the language, that is, the target data item includes the title, publishing source identifier and content summary in the target language.
[0093] Step 306: vectorize and encode the original features of the target file to obtain feature data of the target file.
[0094] In this embodiment, a large coding model is used to perform vectorized coding on the original features of the target file to obtain a feature vector of the target file.
[0095] It can be seen from the above technical solution that the feature data of the target file is obtained using the same feature data acquisition method as the sample file, ensuring the consistency of the expression of the feature data, and by converting the language of the target file into the language of the feature data of the sample file, the flexibility and accuracy of the category recognition of multilingual files are improved.
[0096] See also Figure 4 , Figure 4 A flowchart of another data processing method provided in an embodiment of the present application is shown below. Figure 4 The specific implementation method of step 102, i.e. clustering the characteristic data of the target file and the characteristic data of at least one sample file to obtain the target cluster, is shown as follows: Figure 4 As shown, this method specifically includes steps 401 to 403, and these steps are described in detail below.
[0097] Step 401: Get the total number of categories.
[0098] In this embodiment, the total number of categories is the number of category labels of all sample files without duplicates.
[0099] Step 402: clustering the feature data of the target file and the feature data of at least one sample file with the total number of categories as the number of clusters to obtain a plurality of clusters.
[0100] In an optional embodiment, after principal component analysis (PCA) is performed on the feature data of the target file and the feature data of at least one sample file for dimensionality reduction, the feature data of the target file and the feature data of at least one sample file are clustered, with the total number of categories serving as the number of clusters. It should be noted that dimensionality reduction through PCA improves the efficiency and accuracy of clustering.
[0101] Step 403: The cluster to which the characteristic data of the target file belongs is used as the target cluster.
[0102] As can be seen from the above technical solution, the characteristic data of the target file and the characteristic data of at least one sample file are clustered with the total number of categories as the number of clusters, and the number of clusters obtained is the total number of categories. That is, through the clustering algorithm with a fixed number of clusters, the categories to which the characteristic data of the target file and the characteristic data of at least one sample file belong are preliminarily identified. It can be understood that the probability that the characteristic data of target files with the same category label belong to the same cluster cluster is high, and the probability that the characteristic data of target files with different category labels belong to the same cluster cluster is low. It can be seen that, taking the cluster cluster to which the characteristic data of the target file belongs as the target cluster cluster, the probability that the target file and the sample files within the target cluster cluster (recorded as cluster sample files) belong to the same category is greater than the probability that the target file and the sample files within other cluster clusters belong to the same category.
[0103] See also Figure 5 , Figure 5 A flowchart of another data processing method provided in an embodiment of the present application is shown below. Figure 5 FIG1 shows a specific implementation method of step 103, i.e., determining the sample data based on the target cluster, such as Figure 5 As shown, this method specifically includes steps 501 to 505, and these steps are described in detail below.
[0104] Step 501: Obtain the feature similarity between the feature data of each cluster sample file and the feature data of the target file.
[0105] In this embodiment, the cluster sample files are sample files in the target cluster. Optionally, the feature similarity includes cosine similarity, and the similarity is determined by calculating the cosine similarity between the feature vectors of each cluster sample file and the feature vector of the target file.
[0106] Step 502: Sort the sample files of each cluster from high to low according to feature similarity to obtain a target file sequence.
[0107] Step 503: Select cluster sample files that meet the similarity condition as target sample files.
[0108] In this embodiment, the similarity condition includes a feature similarity with the target file greater than a similarity threshold, and a rank within the target file sequence less than a rank threshold. That is, if the feature similarity between the cluster sample file and the target file is greater than the similarity threshold, and the cluster sample file is ranked before m in the target file sequence, then the cluster sample file is the target sample file, and m is the rank threshold.
[0109] Steps 501 to 503 illustrate a specific method for obtaining at least one target sample file from a target cluster based on feature similarity. In other optional embodiments, obtaining at least one target sample file from a target cluster based on feature similarity can also be achieved through other steps. For example, feature similarity can also be obtained based on similarity measurement methods such as Euclidean distance and Manhattan distance. For another example, the similarity condition only includes that the rank in the target file sequence is less than the rank threshold, or only includes that the feature similarity with the target file is greater than the similarity threshold.
[0110] Step 504: For each category label, randomly obtain a random sample file of the same number of examples to obtain a sample set of all categories.
[0111] In this embodiment, for each file category, a sample number of sample files is obtained based on the category labels in the original sample dataset as random sample files. The full category sample set includes the feature data and corresponding category labels of all random sample files. In other words, the full category sample set includes the feature data and corresponding category labels of the random sample files corresponding to each category label.
[0112] Step 505: Add at least one target sample file to the full-category sample set to obtain sample data based on the at least one target sample file.
[0113] In this embodiment, the feature data and corresponding category labels of all sample files are collected to obtain sample data. The sample files of all categories include at least one target sample file and a random sample file. The random sample file is obtained by randomly selecting sample files for each category label. The sample data also includes the feature data and corresponding category labels of the random sample files.
[0114] It can be seen from the above technical solution that the sample data in the target cluster is screened by feature similarity, and the feature data of multiple target sample files that are most similar to the feature data of the target file are selected. The feature data of the multiple target sample files and all randomly selected random sample files and the corresponding category labels are used as example data, which increases the proportion of sample files with similar features to the target file in the example data, improves the sensitivity and ability of the large model to learn the feature data of the target file, and further improves the accuracy of the large model in identifying the category of the target file.
[0115] Based on the above embodiments, a data processing method provided in an embodiment of the present application can be specifically applied to the tax rate calculation scenario to automatically identify the tax attributes affected by the newly released policy documents, where the tax attributes include tax types (i.e., tax categories), taxpayer groups, and tax payment methods, etc.
[0116] Taking tax types as an example, adjustments to tax policies can affect the corresponding tax rates. In the traditional tax rate calculation process, after obtaining the tax policy, financial professionals use their expertise to match tax types, identifying the tax types affected by the policy and then determining the tax rates based on the tax types. However, manually determining tax types based on tax policies by financial professionals is extremely inefficient and inaccurate.
[0117] In this regard, a data processing method provided in an embodiment of the present application is used to perform tax type matching, that is, to classify policy documents, and the classification result is one of multiple tax types. Figure 6 , Figure 6 A schematic diagram of a specific implementation flow of the data processing method provided in the embodiment of the present application is shown in FIG. Figure 6 As shown, this method specifically includes steps 601 to 611, as follows:
[0118] Step 601: Obtain a training set and an inference set.
[0119] In this embodiment, the training set includes the original data items and category labels of multiple sample files, and the training set is denoted as Dtrain(x1, x2, x3, x4, y). The inference set includes the original data items of multiple files to be processed, and the training set is denoted as Dtest(x1, x2, x3, x4).
[0120] Among them, the original data items x1, x2, x3 and x4 are the file name (i.e., id), title, issuing unit and content respectively, and the category label y represents the category of the sample file, i.e., the tax type.
[0121] Step 602: Use the large translation model to translate the training set and the inference set respectively, unify the language of the training set and the inference set, and obtain a training translation set and a inference translation set.
[0122] In this embodiment, the translation model is used for language translation. Specifically, the translation model is used to translate the original data items and category labels in the training set and the inference set into the target language. The target language is the language of the largest number of sample files.
[0123] In this embodiment, the training translation set is denoted as D'train(x'1, x'2, x'3, x'4, y'), and the inference translation set is denoted as D'test(x'1, x'2, x'3, x'4), where x'1, x'2, x'3, x'4, and y' are x1, x2, x3, x4, and y after the language is unified.
[0124] Step 603: Use the summary model to summarize the contents of the training translation set and the inference translation set to obtain corresponding content summaries.
[0125] In this embodiment, x'4 is taken from D'train (x'1, x'2, x'3, x'4, y') and D'test (x'1, x'2, x'3, x'4), and x'4 is input into the summary model, and the content summary x''4 corresponding to x'4 is generated by the summary model.
[0126] In this embodiment, the content is replaced with the content summary to update the training translation set and the inference translation set. The updated training translation set includes multiple target data items and target category labels for multiple sample files. The updated inference translation set includes multiple target data items for multiple files to be processed. The target data items include the original data items after language unification and / or summary.
[0127] Step 604: Use the large encoding model in the training translation set to vectorize the target data items of the sample file and the file to be processed to obtain feature vectors of the sample file and the file to be processed.
[0128] In this example, the target data items of the sample files in the training translation set are concatenated to obtain the original features of the sample files. The target data items of the to-be-processed files in the inference translation set are concatenated to obtain the original features of the to-be-processed files. Specifically, for the sample files and the to-be-processed files, x'2, x'3, and x''4 are connected with a period to obtain the original features.
[0129] Furthermore, the original features are vectorized and encoded using the encoding model to obtain the training feature data set Dtrainembedding(X) corresponding to the training translation set and the unprocessed feature data set Dtestembedding(X) corresponding to the inference translation set.
[0130] It should be noted that, in another optional implementation, vectorized encoding of the original features corresponding to the training translation set is performed in batches to obtain Dtrainembedding(X), and in step 606, vectorized encoding of the original features of a file to be processed is performed again.
[0131] Step 605: Randomly select the feature vectors of 5 random sample files for each category label from the training feature data set to form a full category sample set.
[0132] In this embodiment, for yi∈y, the feature vectors Di1~Di5 of 5 random sample files corresponding to yi are selected from the training feature data set Dtrainembedding(X) to obtain the full category sample set Ditrain, Ditrain={Di1~Di5}, i∈n, i∈n, n is the total number of tax types.
[0133] Next, the feature vector of each file to be processed in the feature data set to be processed is traversed, and steps 606 to 611 are executed in a loop until the tax types matching all the files to be processed are predicted.
[0134] Step 606: Merge the feature vector of the file to be processed with the training feature data set to obtain a set of feature vectors to be clustered of the file to be processed.
[0135] In this embodiment, the set of feature vectors to be clustered of the file to be processed includes the feature vector of the file to be processed and the feature vectors of all sample files.
[0136] Step 607: After reducing the dimension of the set of feature vectors to be clustered using principal component analysis (PCA), cluster the reduced set of feature vectors to be clustered to obtain target clusters.
[0137] In this embodiment, a k-means clustering algorithm is used to cluster the reduced-dimensionality set, resulting in multiple clusters. The cluster to which the feature vector of the file to be processed belongs is the target cluster for the file to be processed. It is understood that the target cluster includes the feature vector of the file to be processed and the feature vector of at least one cluster sample file. The target number of clusters, k, is equal to n, which is the total number of tax types.
[0138] Step 608: Calculate the cosine similarity between the feature vector of each cluster sample file in the target cluster and the feature vector of the file to be processed.
[0139] Step 609 : Sort the cluster sample files from high to low based on the cosine similarity of the feature vectors, collect the feature vectors of the 15 cluster sample files with the highest cosine similarity, and obtain a target sample file set of files to be processed.
[0140] Step 610: Merge the target sample file set and the full-category sample set to obtain sample data.
[0141] In this example, the full-category sample set includes the feature vectors of five sample files corresponding to each tax type, and the target sample file set includes the feature vectors of the 15 target sample files with the highest similarity to the file to be processed. The sample data then includes the feature vectors of 20 sample files corresponding to candidate tax types, as well as the feature vectors of five sample files corresponding to each other tax type. The candidate tax types are the tax types represented by the target clusters obtained after clustering. Other categories are tax types other than the candidate categories.
[0142] Step 611: Input the feature vector and sample data of the document to be processed into the large model, so that the large model uses the diversity example learning prompt engineering to generate the tax type of the document to be processed.
[0143] It can be seen from the above technical solution that based on the data processing method provided in the embodiment of this application, the tax type matching process is made intelligent, and the training set is processed by clustering and cosine similarity technology to obtain example data provided for the diversity example learning prompt project. The tax type is predicted by a large model, and the tax rate matching process is more flexible and faster.
[0144] Furthermore, the sample data is diverse and targeted, thereby improving the standardization and accuracy of the prompting project. Using the sample set as the test set, the large model was tested, and the performance test results showed that the performance index F1 was equal to 84.27%.
[0145] Furthermore, this method unifies the language of the training set and the inference set to achieve matching for global multilingual tax types.
[0146] The present application also provides a data processing device, referring to Figure 7 , Figure 7 A structural diagram of a data processing device provided in an embodiment of the present application is shown in FIG. Figure 7 As shown, the device includes:
[0147] A first feature acquisition unit 701 is used to acquire feature data of a target file;
[0148] A feature clustering unit 702 is configured to cluster the feature data of the target file and the feature data of at least one sample file to obtain a target cluster, wherein the target cluster includes the feature data of the target file;
[0149] An example acquisition unit 703 is configured to determine example data based on the target cluster, where the example data includes feature data of at least one target sample file in the target cluster and a category label of the target sample file;
[0150] The model prediction unit 704 is configured to input the characteristic data of the target file and the example data into a large model, and guide the large model to generate a category of the target file.
[0151] Preferably, the device also includes a second feature acquisition unit, which is used to: obtain an original sample data set, wherein the original sample data includes the original data items and category labels of the sample files, and the original data items include titles, publication source identifiers and content; unify the languages of the original data items and category labels of all the sample files into a target language, and the target language is the language with the largest number of sample files; obtain content summaries of each of the sample files respectively; for each of the sample files, splice the target data items of the sample files to obtain the original features of the sample files, and the target data items include the titles, publication source identifiers and content summaries after the unified language; vectorize and encode the original features of each of the sample files respectively to obtain feature data of the sample files.
[0152] Preferably, when the first feature acquisition unit is used to acquire feature data of the target file, it is specifically used to:
[0153] Obtaining original data items of the target file;
[0154] If the language of the original data item of the target file is not the target language, translating the original data item of the target file into the target language;
[0155] Obtaining a content summary of the target file;
[0156] splicing target data items of the target file to obtain original features of the target file;
[0157] Vectorized encoding is performed on the original features of the target file to obtain feature data of the target file.
[0158] Preferably, the feature clustering unit is used to cluster the feature data of the target file and the feature data of at least one sample file to obtain a target cluster, specifically for:
[0159] Get the total number of categories;
[0160] Taking the total number of the categories as the number of clusters, clustering the feature data of the target file and the feature data of at least one of the sample files to obtain a plurality of clusters;
[0161] The cluster to which the characteristic data of the target file belongs is used as the target cluster.
[0162] Preferably, the feature clustering unit is used to cluster the feature data of the target file and the feature data of at least one of the sample files using the total number of categories as the number of clusters, specifically to:
[0163] After principal component analysis dimensionality reduction is performed on the feature data of the target file and the feature data of at least one of the sample files, the feature data of the target file and the feature data of at least one of the sample files are clustered using the total number of categories as the number of clusters.
[0164] Preferably, when the example acquisition unit is used to determine the example data based on the target cluster, it is specifically used to:
[0165] Respectively obtaining feature similarities between feature data of each cluster sample file and feature data of the target file, wherein the cluster sample file is the sample file in the target cluster;
[0166] Based on feature similarity, obtaining at least one target sample file from the target cluster;
[0167] The sample data is obtained based on at least one of the target sample files.
[0168] Preferably, when the example acquisition unit is used to acquire at least one target sample file from the target cluster based on feature similarity, it is specifically used to:
[0169] Sort the sample files of each cluster according to feature similarity from high to low to obtain a target file sequence;
[0170] A cluster sample file that meets a similarity condition is selected as the target sample file. The similarity condition includes that the feature similarity with the target file is greater than a similarity threshold, and / or the rank in the target file sequence is less than a rank threshold.
[0171] Preferably, the device further comprises a random selection unit for randomly acquiring an example number of random sample files for each category label; wherein the example data further comprises feature data of the random sample files and corresponding category labels.
[0172] The present application also provides an electronic device, Figure 8 , Figure 8 The schematic diagram of the structure of an electronic device provided in the embodiment of the present application shows the structure of the electronic device suitable for implementing the embodiment of the present application. Figure 8 As shown, the electronic devices in the embodiments of the present application may include but are not limited to fixed terminals such as mobile phones, laptops, PDAs (personal digital assistants), PADs (tablet computers), desktop computers, etc. Figure 8 The electronic device shown is merely an example and should not limit the functions and scope of use of the embodiments of the present application.
[0173] like Figure 8As shown, the electronic device may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 801, which can perform various appropriate actions and processes based on programs stored in a read-only memory (ROM) 802 or programs loaded from a storage device 808 into a random access memory (RAM) 803. When the electronic device is powered on, the RAM 803 also stores various programs and data required for the operation of the electronic device. The processing device 801, ROM 802, and RAM 803 are interconnected via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.
[0174] Typically, the following devices may be connected to the I / O interface 805: an input device 806 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 807 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 808 including, for example, a memory card, a hard disk, etc.; and a communication device 809. The communication device 809 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Figure 8 The electronic device is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed instead.
[0175] An embodiment of the present application also provides a computer program product including computer-readable instructions. When the computer-readable instructions are executed on an electronic device, the electronic device implements each step of any data processing method provided in the embodiment of the present application.
[0176] A computer-readable storage medium is also provided in an embodiment of the present application. The storage medium carries one or more computer programs. When the one or more computer programs are executed by an electronic device, the electronic device can implement the various steps of any data processing method provided in the embodiment of the present application.
[0177] It should also be noted that the device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed across multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the drawings of the device embodiments provided in this application, the connection relationship between the modules indicates that there is a communication connection between them, which can be specifically implemented as one or more communication buses or signal lines.
[0178] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus necessary general hardware, and of course can also be implemented by special hardware including application-specific integrated circuits, special CPUs, special memories, special components, etc. In general, all functions performed by computer programs can be easily implemented with corresponding hardware, and the specific hardware structures used to implement the same function can also be diverse, such as analog circuits, digital circuits or special circuits, etc. However, for the present application, software program implementation is a better implementation method in most cases. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art can be embodied in the form of a software product, which is stored in a readable storage medium, such as a computer's floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disk, etc., and includes a number of instructions to enable a computer device (which can be a personal computer, training equipment, or network equipment, etc.) to execute the methods described in each embodiment of the present application.
[0179] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented by software, all or part of the embodiments may be implemented in the form of a computer program product.
[0180] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, a computer, a training device or a data center by wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) mode to another website, a computer, a training device or a data center. The computer-readable storage medium can be any available medium that a computer can store or a data storage device such as a training device, a data center, etc. that includes one or more available media integrations. The available medium can be a magnetic medium, (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive (SSD)).
[0181] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Reference can be made to the common and similar parts between the various embodiments. For the devices disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method description.
[0182] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the components and steps of each example according to their functions. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0183] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0184] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data processing method, comprising: Obtain characteristic data of the target file; Clustering the characteristic data of the target file and the characteristic data of at least one sample file to obtain a target cluster, wherein the target cluster includes the characteristic data of the target file; Determining example data based on the target cluster, the example data including feature data of at least one target sample file in the target cluster and a category label of the target sample file; The characteristic data of the target file and the example data are input into a large model, and the large model is guided to generate a category of the target file.
2. The data processing method according to claim 1, further comprising: Acquire an original sample data set, wherein the original sample data includes original data items and category labels of the sample files, and the original data items include titles, publishing source identifiers, and content; Unifying the languages of the original data items and category labels of all the sample files into a target language, where the target language is the language with the largest number of sample files; Obtaining content summaries of each of the sample files respectively; For each of the sample files, the target data items of the sample files are spliced to obtain the original features of the sample files, wherein the target data items include the title in a unified language, the source identifier, and the content summary; Vectorized encoding is performed on the original features of each of the sample files to obtain feature data of the sample files.
3. The data processing method according to claim 2, wherein obtaining characteristic data of the target file comprises: Obtaining original data items of the target file; If the language of the original data item of the target file is not the target language, translating the original data item of the target file into the target language; Obtaining a content summary of the target file; splicing target data items of the target file to obtain original features of the target file; Vectorized encoding is performed on the original features of the target file to obtain feature data of the target file.
4. The data processing method according to claim 1, wherein clustering the characteristic data of the target file and the characteristic data of at least one sample file to obtain a target cluster comprises: Get the total number of categories; Taking the total number of the categories as the number of clusters, clustering the feature data of the target file and the feature data of at least one of the sample files to obtain a plurality of clusters; The cluster to which the characteristic data of the target file belongs is used as the target cluster.
5. The data processing method according to claim 4, wherein clustering the characteristic data of the target file and the characteristic data of at least one sample file using the total number of categories as the number of clusters comprises: After principal component analysis dimensionality reduction is performed on the feature data of the target file and the feature data of at least one of the sample files, the feature data of the target file and the feature data of at least one of the sample files are clustered using the total number of categories as the number of clusters.
6. The data processing method according to claim 1, wherein determining example data based on the target cluster comprises: Respectively obtaining feature similarities between feature data of each cluster sample file and feature data of the target file, wherein the cluster sample file is the sample file in the target cluster; Based on feature similarity, obtaining at least one target sample file from the target cluster; The sample data is obtained based on at least one of the target sample files.
7. The data processing method according to claim 6, wherein the step of obtaining at least one target sample file from the target cluster based on feature similarity comprises: Sort the sample files of each cluster according to feature similarity from high to low to obtain a target file sequence; A cluster sample file that meets a similarity condition is selected as the target sample file. The similarity condition includes that the feature similarity with the target file is greater than a similarity threshold, and / or the rank in the target file sequence is less than a rank threshold.
8. The data processing method according to claim 1, further comprising: For each category label, a random sample file with a random number of examples is obtained; The example data also includes feature data of the random sample file and corresponding category labels.
9. A data processing device comprising: A first feature acquisition unit, configured to acquire feature data of a target file; a feature clustering unit, configured to cluster the feature data of the target file and the feature data of at least one sample file to obtain a target cluster, wherein the target cluster includes the feature data of the target file; An example acquisition unit, configured to determine example data based on the target cluster, wherein the example data includes feature data of at least one target sample file in the target cluster and a category label of the target sample file; The model prediction unit is used to input the characteristic data of the target file and the example data into the large model, and guide the large model to generate the category of the target file.
10. An electronic device comprising: A memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the steps of the data processing method according to claim 1.