Data set quality evaluation method and device, equipment, medium and program product

Through hierarchical clustering and large-scale model theme extraction, the problem of insufficient semantic hierarchy and multi-dimensional coverage in large-scale data set quality evaluation is solved, and the adaptability and quality management capabilities of data sets in multi-task large-scale model training are improved.

CN120372325AActive Publication Date: 2025-07-25ZHEJIANG GEELY HLDG GRP CO LTD +1
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510873792.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-07-25
Estimated Expiration
2045-06-27

AI Technical Summary

Technical Problem

The quality evaluation of large-scale data sets in the prior art focuses only on a single dimension, ignoring semantic hierarchy, task instruction coverage and language balance, resulting in poor adaptability of multi-task big model training.

Method used

Through hierarchical clustering and large-scale theme extraction, a multi-layer label tree is constructed, a weighted frequency label tree is generated, and the node weights are calculated based on the label distribution of the text data set, and quantitative indicators of multiple dimensions are extracted for quality evaluation.

Benefits of technology

A comprehensive analysis of large-scale data sets is realized, and the adaptability and quality control capabilities of data sets in multi-task large-scale training scenarios are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372325A_ABST
    Figure CN120372325A_ABST
Patent Text Reader

Abstract

The invention provides a data set quality evaluation method and device, equipment, a medium and a program product, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a text data set of which the quality is to be evaluated; performing hierarchical clustering analysis on the text data set to obtain a clustering result of the text data set, and performing topic extraction on sample data in the text data set based on a preset large model to obtain a semantic topic of the sample data; constructing a multi-layer label tree based on the clustering result and the semantic topic, and calculating the weight of each node of the multi-layer label tree based on the label distribution of the text data set to generate a weighted frequency label tree; and extracting quantitative indexes of multiple dimensions of the text data set based on the weighted frequency tag tree so as to carry out quality evaluation on the text data set. By adopting the method and the device, the adaptability of the data set in a multi-task large model training scene can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular, to a method, device, equipment, medium, and program product for evaluating the quality of a data set. Background Art

[0002] In related technologies, the quality evaluation of large-scale data sets usually only focuses on a single dimension (such as label frequency), ignoring factors such as semantic level, task instruction coverage, and language balance, resulting in poor adaptability performance of large-scale data sets in the training and reinforcement fine-tuning of multi-task large models. For example, even two large-scale sample data sets with the same instruction frequency may be completely different at the semantic level. For another example, the candidate large-scale data set lacks data coverage in some key areas that have a direct impact on multi-task large models. These situations will ultimately affect the generalization ability of multi-task large models.

[0003] In summary, the related technologies only focus on a single dimension for the quality evaluation of large-scale data sets, making it difficult to comprehensively analyze data features, resulting in poor adaptability of large-scale data sets in the training scenario of multi-task large models. Summary of the Invention

[0004] The main purpose of the embodiments of the present application is to propose a method, device, equipment, medium, and program product for evaluating the quality of a data set, aiming to improve the adaptability of the data set in the training scenario of multi-task large models by comprehensively analyzing large-scale data sets for refined management and optimization of data quality.

[0005] To achieve the above object, a first aspect of the embodiments of the present application proposes a method for evaluating the quality of a data set, the method including: Obtain a text data set whose quality is to be evaluated; Perform hierarchical clustering analysis on the text data set to obtain the clustering result of the text data set, and extract the semantic theme of the sample data in the text data set based on a preset large model; Construct a multi-layer label tree based on the clustering result and the semantic theme, and calculate the weights of each node of the multi-layer label tree based on the label distribution of the text data set to generate a weighted frequency label tree; Extract quantitative indicators of multiple dimensions of the text data set based on the weighted frequency label tree to evaluate the quality of the text data set.

[0006] In some embodiments, the extracting the semantic theme of the sample data in the text data set based on a preset large model includes: Obtain a preset text prompt; the text prompt includes at least the processing requirements of the data and the output requirements of the data processing result; Based on the text prompt, guide the pre-trained large model to extract the topic of the sample data in the text dataset according to the processing requirements, and output the first semantic topic obtained by the extraction according to the output requirements; the semantic topic of the sample data includes the first semantic topic.

[0007] In some embodiments, after performing hierarchical clustering analysis on the text dataset to obtain the clustering result of the text dataset, the method further includes: Based on a preset text prompt, guide the pre-trained large model to generate a second semantic topic for the clustering clusters in the clustering result, and use the second semantic topic as the semantic topic of the sample data in the text dataset.

[0008] In some embodiments, the clustering result includes a tree-shaped clustering result; the constructing a multi-layer label tree based on the clustering result and the semantic topic includes: Generate the semantic node name of the target node in the tree-shaped clustering result based on the semantic topic, so as to construct the tree-shaped clustering result into a multi-layer label tree.

[0009] In some embodiments, the calculating the weights of each node of the multi-layer label tree based on the label distribution of the text dataset to generate a weighted frequency label tree includes: Determine the frequency of each node in the multi-layer label tree based on the label distribution of the text dataset; Perform weight calculation based on the frequency of each node and the depth of each node in the multi-layer label tree to obtain the weight of each node; Perform normalization processing on the weights of each node to generate a weighted frequency label tree; Wherein, the normalization processing includes horizontal normalization processing and vertical normalization processing. The horizontal normalization processing is used to perform normalization calculation on the weights of the nodes belonging to the same level in the multi-layer label tree, and the vertical normalization processing is used to perform normalization calculation on the weights of all nodes in the multi-layer label tree.

[0010] In some embodiments, the method further includes: Add labels to the sample data in the text dataset, and determine the label distribution of the text dataset based on the labels of the sample data; the labels of the sample data correspond to the nodes in the multi-layer label tree; The determining the frequency of each node in the multi-layer label tree based on the label distribution of the text dataset includes: Determine the frequency of the sample data in the text dataset based on the label distribution of the text dataset; Statistically count the frequency of each node in the multi-layer label tree based on the frequency of the sample data.

[0011] In some embodiments, extracting the quantization metrics of multiple dimensions of the text data set based on the weighted frequency label tree includes: Calculating the target dimension metrics based on the weighted frequency label tree and the sample label frequency information of the text data set to obtain the quantization metrics of multiple dimensions of the text data set; Among them, the sample label frequency information includes the frequency of the sample data in the text data set; the target dimension metrics include at least two of the domain diversity metric, the instruction diversity metric, the complexity metric, and the Chinese-English ratio metric.

[0012] In some embodiments, the text data set includes a reference text data set and a candidate text data set, and the weighted frequency label tree includes a first weighted frequency label tree of the reference text data set and a second weighted frequency label tree of the candidate text data set; The method further includes: Performing cross-entropy comparison on the first weighted frequency label tree and the second weighted frequency label tree to obtain the structural similarity between the first weighted frequency label tree and the second weighted frequency label tree; Evaluating the structural distribution difference between the candidate text data set and the reference text data set based on the structural similarity to obtain a structural distribution difference evaluation result; the structural distribution difference evaluation result is one of the results for evaluating the quality of the candidate text data set.

[0013] To achieve the above object, a second aspect of the embodiments of the present application proposes a quality evaluation device for a data set, and the device includes: An acquisition module, configured to acquire a text data set whose quality is to be evaluated; A clustering and model topic extraction module, configured to perform hierarchical clustering analysis on the text data set to obtain a clustering result of the text data set, and, based on a preset large model, extract the semantic topics of the sample data in the text data set to obtain the semantic topics of the sample data; A label tree construction module, configured to construct a multi-layer label tree based on the clustering result and the semantic topic, and calculate the weights of each node of the multi-layer label tree based on the label distribution of the text data set to generate a weighted frequency label tree; A data set quality evaluation module, configured to extract the quantization metrics of multiple dimensions of the text data set based on the weighted frequency label tree to evaluate the quality of the text data set.

[0014] To achieve the above object, a third aspect of the embodiments of the present application provides a quality evaluation device for a data set. The quality evaluation device for the data set includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the quality evaluation method for the data set described in the first aspect above is implemented.

[0015] To achieve the above object, a fourth aspect of the embodiments of the present application provides a vehicle, on which a quality evaluation device for a data set is configured. The quality evaluation device for the data set includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the quality evaluation method for the data set described in the first aspect above is implemented.

[0016] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer-readable storage medium that stores a computer program, and when the computer program is executed by a processor, the quality evaluation method for the data set described in the first aspect above is implemented.

[0017] To achieve the above object, a sixth aspect of the embodiments of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the quality evaluation method for the data set provided in the first aspect above is implemented.

[0018] The quality evaluation method, device, equipment, vehicle, computer-readable storage medium, and computer program product provided in the present application obtain a text data set whose quality is to be evaluated; perform hierarchical clustering analysis on the text data set to obtain a clustering result of the text data set, and extract semantic topics of sample data in the text data set based on a preset large model to obtain semantic topics of the sample data; construct a multi-layer label tree based on the clustering result and the semantic topic, and calculate weights of each node of the multi-layer label tree based on the label distribution of the text data set to generate a weighted frequency label tree; extract quantitative indicators of multiple dimensions of the text data set based on the weighted frequency label tree to evaluate the quality of the text data set.

[0019] Compared with the method of evaluating the quality of a large-scale data set from only a single dimension, in the embodiments of the present application, hierarchical clustering analysis is performed on the large-scale text data set to be evaluated for quality, and topic extraction is performed on the sample data therein based on a preset large model. Thus, a multi-layer label tree capable of reflecting the semantic hierarchy is constructed based on the clustering results obtained from the hierarchical clustering analysis and the semantic topics obtained from the topic extraction. Then, the weights of each node of the multi-layer label tree are calculated through the label distribution of the text data set to generate a weighted frequency label tree. Finally, quantitative indicators in multiple dimensions of the text data set are extracted based on the weighted frequency label tree for evaluating the quality of the text data set. In this way, in the embodiments of the present application, by combining hierarchical clustering of a large-scale data set with semantic extraction of a large model, a label tree structure that can reflect the semantic hierarchy is constructed, and a weighted frequency label tree is further constructed. On this basis, quantitative indicators are extracted from multiple dimensions respectively to evaluate the quality of the data set, which can comprehensively analyze the large-scale data set to perform refined management and optimization of data quality, thereby improving the adaptability of the data set in the multi-task large model training scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 is a schematic flowchart of the steps of the method for evaluating the quality of a data set provided by the embodiments of the present application in some embodiments; Figure 2 is a schematic flowchart of the steps of the method for evaluating the quality of a data set provided by the embodiments of the present application in some other embodiments; Figure 3 is Figure 1 a schematic flowchart of the refined steps of step S102 in Figure 4 is Figure 1 a schematic flowchart of the refined steps of step S103 in Figure 5 is a schematic flowchart of the multi-dimensional data set distribution analysis process of hierarchical clustering and weighted label tree involved in a complete embodiment of the method for evaluating the quality of a data set provided by the embodiments of the present application; Figure 6 is a schematic flowchart of the cooperation process of large model topic extraction and label marking involved in a complete embodiment of the method for evaluating the quality of a data set provided by the embodiments of the present application; Figure 7 is a schematic flowchart of the process of measuring data quality based on four major indicators involved in a complete embodiment of the method for evaluating the quality of a data set provided by the embodiments of the present application; Figure 8 is a schematic structural diagram of the device for evaluating the quality of a data set provided by the embodiments of the present application; Figure 9 is a schematic hardware structure diagram of the device for evaluating the quality of a data set provided by the embodiments of the present application. Detailed implementation manners

[0021] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0022] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different manner from the module division in the device or the order in the flowchart. Terms such as "first", "second", etc. in the description, claims and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0023] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.

[0024] First, the overall concept of the method for evaluating the quality of the data set provided in the embodiments of the present application will be described.

[0025] With the development of artificial intelligence technology, especially the increasingly wide application of large language models (such as GPT, GLM, BERT, etc.), and since the training effect of the model highly depends on the diversity, coverage and structural rationality of the data set, the demand for high-quality large-scale data sets is also increasing day by day. In addition, with the gradual popularization of multi-task models (MTM) and reinforcement fine-tuning (RFT), the requirements for the diversity of task types, the complexity of instruction design, and the balance of language proportion in the sample data are also getting higher and higher.

[0026] However, most of the data analysis methods in the related art still stay at the level of statistical analysis of single dimensions such as data label frequency, sample quantity, and text length, lacking a multi-dimensional and structured panoramic evaluation method. As a result, problems such as uneven data distribution, domain redundancy, and single language structure are difficult to discover in a timely manner. For example, during the quality assessment stage of large-scale data sets in the related art, usually only a single dimension (such as label frequency) is concerned, ignoring factors such as semantic level, task instruction coverage, and language balance. This will lead to poor adaptability of the data set in the training of multi-task large models. This is because even two large-scale sample data sets with the same instruction frequency may be completely different at the semantic level, or the candidate large-scale data set lacks data coverage in some key areas that have a direct impact on multi-task large models. These situations will ultimately affect the generalization ability of multi-task large models.

[0027] Therefore, there is an urgent need for a structured, multi-dimensional, and normalizable data analysis mechanism to perform multi-dimensional data set analysis with strong interpretability, clear structure, and comprehensive coverage for large-scale data sets. By comprehensively understanding the distribution characteristics of the data set, refined management and optimization of data quality can be achieved, and then scientific evaluation of the data set quality, task adaptability analysis, and distribution deviation optimization can be supported for data engineers and algorithm developers.

[0028] Based on this, the embodiments of the present application provide a method, device, equipment, vehicle, computer-readable storage medium, and computer program product for quality assessment of a data set. By comprehensively analyzing large-scale data sets for refined management and optimization of data quality, the adaptability of the data set in the training scenario of multi-task large models can be improved.

[0029] Considering that structured methods such as hierarchical clustering and label trees have good semantic parsing and hierarchical modeling capabilities, which can provide a new perspective for data analysis. Therefore, the embodiments of the present application construct an analysis system that integrates label hierarchical structure and distribution index modeling capabilities, and creatively propose a multi-dimensional data analysis mechanism that combines hierarchical clustering (Hierarchical Clustering) and large model theme extraction mechanism, constructs a weighted label tree (Label Tree), and introduces weight normalization. By integrating label frequency and hierarchical structure, it can quantify domain distribution, task instruction distribution, language ratio, and complexity indicators, and on this basis, construct a normalized distribution comparison mechanism (such as cross-entropy distribution comparison Cross-Entropy Distribution Matching), which can achieve structural alignment and evaluation between the candidate text data set and the benchmark text data set, effectively making up for the limitations of the data analysis methods in the above-mentioned related art.

[0030] It should be noted that hierarchical clustering is a clustering method that recursively divides data into hierarchical cluster groups with a tree-like structure. This method constructs a clustering hierarchy by calculating the similarity or distance between samples (such as the Ward method) and continuously merging the closest data points or clusters. A label tree is a data label system expressed in a tree structure, where each node represents a clustering theme or label category, and the child nodes are its sub-categories. The label tree can reflect the hierarchical and dependency relationships between labels and is used for multi-dimensional distribution analysis. Cross-entropy distribution comparison is used to compare the similarity of the domain distributions between the candidate text dataset and the benchmark text dataset. By constructing a weighted frequency label tree and calculating its cross-entropy, the degree of difference in the distribution structure between the two datasets is quantified.

[0031] In the embodiments of this application, the multi-dimensional data distribution analysis based on hierarchical clustering and weighted frequency label tree aims to comprehensively quantify and evaluate the distribution characteristics of large-scale text datasets from a structural perspective. By combining hierarchical clustering with large model semantic extraction, a label tree structure that can reflect semantic levels is constructed, and the weights of the label tree are calculated based on sample frequencies and node depths, thereby forming a weighted frequency tree (Weighted Frequency Tree). On this basis, quantitative indicators are extracted from multiple dimensions (such as domain diversity, instruction diversity, complexity score, and Chinese-English ratio), and at the same time, structured comparative analysis between the candidate text dataset and the benchmark text dataset is supported, so as to improve the adaptability, balance, and quality control ability of the dataset in the multi-task model training scenario.

[0032] It should be noted that a weighted frequency label tree is a variant of the label tree that combines label frequency and node hierarchy information. The weight of each node in the weighted frequency label tree is jointly determined by its sample frequency and hierarchical depth, and is used to measure the importance of the node in the overall tree structure. Domain diversity represents the distribution breadth and uniformity of different semantic domains contained in the dataset, and quantifies the distribution of each domain in the dataset through the label tree structure and information entropy calculation method. Instruction diversity represents the richness and balance of different task instruction types contained in the dataset, and is calculated by comparing the frequency distributions of various instructions in the dataset, and is often used to judge the coverage ability of the dataset. The complexity metric measures the structural complexity of the sample text or task, and obtains the complexity fluctuation level of the overall existing dataset by calculating the degree of difference between each sample and the standard complexity.

[0033] In the embodiments of the present application, through a unique data modeling and evaluation mechanism, specifically including: First, constructing a multi-layer label tree through hierarchical clustering and large language model topic extraction; Second, calculating node weights by combining sample frequency and hierarchical depth, and introducing horizontal and vertical normalization methods to generate a weighted frequency label tree; Third, defining and implementing calculation methods for multi-dimensional indicators such as domain diversity, instruction diversity, complexity, and Chinese-English ratio; Fourth, performing a comparison of the normalized distribution similarity based on cross-entropy between the candidate text data set and the benchmark text data set, constituting an integrated multi-dimensional data analysis framework that is evaluable, interpretable, and comparable.

[0034] Next, the quality evaluation method, device, equipment, vehicle, computer-readable storage medium, and computer program product of the data set provided by the embodiments of the present application will be specifically described through the following embodiments, and first, the quality evaluation method of the data set provided by the embodiments of the present application will be described in detail.

[0035] It should be noted that in each specific implementation manner of the present application, when it comes to relevant processing that needs to be based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or jumping to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.

[0036] It should be noted that the method for evaluating the quality of the dataset provided in the embodiments of the present application can be applied to a terminal, can also be applied to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be an in-vehicle terminal on a vehicle, or can be a computer device such as a smart phone, a tablet computer, a laptop computer, or a desktop computer. The terminal can also be associated with the vehicle, that is, the terminal can perform communication data interaction with the vehicle based on a network. The server side can be a back-end server terminal device of the vehicle, can be configured as an independent physical server, can also be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery network (CDN), and big data and artificial intelligence platforms. The software can be an application for implementing the method for evaluating the quality of the dataset, a computer program, and a storage medium carrying the computer program, etc. It should be understood that based on different design requirements of actual applications, in different feasible embodiments, the terminal, the server side, and the software, etc., that apply the method for evaluating the quality of the dataset provided in the embodiments of the present application can, of course, also be other forms not listed here. The method for evaluating the quality of the dataset provided in the embodiments of the present application does not specifically limit this.

[0037] In addition, the present application can also be used in many general or special computer system environments or configurations. For example: vehicles, personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer computer devices, personal computers (PCs), minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0038] For ease of understanding and explanation, in the following text, the quality assessment method for the dataset provided by the embodiments of the present application is taken as an example for terminal devices to elaborate on each specific embodiment of the present application in detail. The implementation of the quality assessment method for the dataset provided by the embodiments of the present application for any of the above-mentioned forms of the subject can refer to the process of the quality assessment method for the dataset applied by the terminal device described later.

[0039] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of the steps of the quality assessment method for the dataset provided by the embodiments of the present application in some embodiments. It should be understood that although Figure 1 and the subsequent other schematic flowcharts show the execution order of some method steps, due to different design requirements in actual applications, the quality assessment method for the dataset provided by the embodiments of the present application can of course adopt an execution order different from that shown in the figures. That is, Figure 1 the order of the shown method steps does not constitute a limitation on the execution logic order of the quality assessment method for the dataset provided by the embodiments of the present application, and any reasonable changes based on Figure 1 the order of the shown method steps should be included within the protection scope of the quality assessment method for the dataset provided by the embodiments of the present application.

[0040] As Figure 1 shown, in some embodiments, the quality assessment method for the dataset provided by the embodiments of the present application may include the following steps S101 to S104.

[0041] Step S101: Obtain a text dataset whose quality is to be evaluated.

[0042] The terminal device can continuously collect relevant sample data (such as image description data and audio text data, etc.) to automatically construct a large-scale text dataset for large model training or fine-tuning, and then use the constructed large-scale dataset as the text dataset to be evaluated for quality before actual model training or fine-tuning applications.

[0043] In some embodiments, the terminal device can also directly receive a text dataset that has been constructed and is to be evaluated for quality through a preset data interface (such as a human-computer interaction interface such as a user interface, application APP, and voice assistant).

[0044] Step S102: Perform hierarchical clustering analysis on the text dataset to obtain the clustering result of the text dataset, and, based on a preset large model, extract the semantic topics of the sample data in the text dataset to obtain the semantic topics of the sample data.

[0045] After the terminal device obtains the text data set to be evaluated for quality, it processes the text data set by combining hierarchical clustering and large model semantic extraction respectively. That is, by performing uneven clustering analysis on the text data set to obtain the clustering result of the text data set, and, based on a preset large model, extracting the theme of the sample data in the text data set, so as to obtain the semantic theme of the sample data (such as "basic mathematics", "open-ended questions", etc.).

[0046] It should be noted that the preset large model can be any language model with rich semantic expression ability, such as GPT-4, ChatGPT, etc.

[0047] Step S103: Construct a multi-layer label tree based on the clustering result and the semantic theme, and calculate the weights of each node of the multi-layer label tree based on the label distribution of the text data set to generate a weighted frequency label tree.

[0048] During the process of the terminal device processing the text data set by combining hierarchical clustering and large model semantic extraction respectively, after obtaining the clustering result of the text data set and the semantic theme of the sample data in the text data set, it can further construct a label tree structure that can reflect the semantic hierarchy based on the clustering result and the semantic theme, that is, a multi-layer label tree. And, the terminal device further calculates the weight of each node in the multi-layer label tree in combination with the label distribution of the text data set, so as to generate a weighted frequency label tree.

[0049] In some embodiments, the clustering result obtained by the terminal device through hierarchical clustering analysis of the text data set includes a tree-shaped clustering result.

[0050] The terminal device can use the clustering analysis module in the integrated multi-dimensional data analysis framework to perform hierarchical clustering analysis on the text data set. The clustering analysis module takes the original text data set as input, and based on the processing of the hierarchical clustering algorithm (such as Ward) and large model theme extraction on the text data set, outputs preliminary clustering labels and the theme representation vector of each cluster. Among them, the preliminary clustering labels can be the clustering result of the text data set, and the theme representation vector of each cluster can be the semantic theme of the sample data in the text data set.

[0051] Exemplarily, the specific steps of the clustering algorithm adopted by the clustering analysis module may include: 1. Sample representation: Convert the samples in the text data set into vector representations, commonly using TF-IDF, sentence vectors (such as BERT, SBERT Embedding). 2. Similarity calculation: Calculate the Euclidean distance or cosine similarity based on the converted vectors to form a distance matrix. 3. Clustering method selection: Adopt Ward’s linkage, and recursively merge samples or clusters by minimizing the within-class variance. 4. Construct a tree-shaped clustering structure: Generate a dendrogram, and each merging operation corresponds to a parent node in the label tree.

[0052] In this case, in the above step S102, the step of "constructing a multi-layer label tree based on the clustering result and the semantic theme" may include the following steps: Generate the semantic node name of the target node in the tree-shaped clustering result based on the semantic theme, so as to construct the tree-shaped clustering result into a multi-layer label tree.

[0053] When the terminal device constructs a multi-layer label tree based on the clustering result and the semantic theme, it can use the semantic theme as the label tree node name of the tree-shaped clustering result, so as to construct a multi-layer label tree that can reflect the semantic hierarchy.

[0054] It should be noted that the label tree node names named after the semantic theme in the multi-layer label tree are the basic framework for subsequent label marking and label weight calculation. Therefore, the accuracy of theme extraction from the sample data in the text data set by the large model determines the semantic structure quality of the label tree.

[0055] In some embodiments, the terminal device may use the label tree construction module in the integrated multi-dimensional data analysis framework to take the clustering result of the text data set as input, so as to construct a multi-layer label tree based on the clustering result through this label tree construction module, calculate the frequency and depth of each node in the multi-layer label tree, and output a label tree structure with frequency information.

[0056] In some embodiments, when the terminal device calculates the weights of each node in the multi-layer label tree based on the label distribution of the text data set to generate a weighted frequency label tree, it can determine the label frequency corresponding to each node in the multi-layer label tree based on the label distribution, so as to generate a weighted label tree (weighted frequency label tree) by combining the label frequency and the tree-shaped clustering result. In this way, the data distribution can be made hierarchically interpretable, which is better than the flat label statistics method adopted in the related art.

[0057] In some embodiments, the terminal device may further adopt the weight calculation and normalization module in the integrated multi-dimensional data analysis framework, taking the frequencies and depths of the multi-layer label tree and its nodes as inputs, so as to calculate the weight w of each node i based on the following formula 1 by the weight calculation and normalization module i , and generate and output a weighted frequency label tree by performing horizontal / vertical normalization processing on the label tree.

[0058] w i = alpha f i + beta(1 / L i ), Formula 1 where f i is the frequency of node i, and L i is the depth of node i.

[0059] Step S104: Extract quantitative indicators of multiple dimensions of the text dataset based on the weighted frequency label tree to evaluate the quality of the text dataset.

[0060] After the terminal device constructs the weighted frequency label tree, it can further calculate indicators from multiple dimensions on this basis, so as to extract quantitative indicators of multiple dimensions of the text dataset. In this way, the terminal device can independently evaluate the text dataset in terms of semantic coverage, task structure, language balance, and text complexity based on the extracted quantitative indicators, so as to form a panoramic description of the quality of the text dataset.

[0061] In some embodiments, when the terminal device extracts quantitative indicators of multiple dimensions of the text dataset based on the weighted frequency label tree, it can extract quantitative indicators from four dimensions: domain diversity, instruction diversity, complexity, and Chinese-English ratio.

[0062] Based on this, the step of "extracting quantitative indicators of multiple dimensions of the text dataset based on the weighted frequency label tree" in the above step S104 may include the following steps: Calculate the target dimension indicators based on the weighted frequency label tree and the sample label frequency information of the text dataset to obtain the quantitative indicators of multiple dimensions of the text dataset. Among them, the sample label frequency information includes the frequencies of the sample data in the text dataset; the target dimension indicators include at least two of the domain diversity indicator, the instruction diversity indicator, the complexity indicator, and the Chinese-English ratio indicator.

[0063] It should be noted that the sample label frequency information of the text dataset is the label distribution frequency of the samples in the text dataset. The terminal device can pre-label the sample data in the text dataset, and thus the label distribution frequency of the sample data in the text dataset can be calculated.

[0064] In some embodiments, the terminal device can use the label tagging module in the integrated multi-dimensional data analysis framework with the text dataset as the input, so that the label tagging module performs multi-label tagging on the sample data in the text dataset according to the domain, and / or performs single-label annotation on the sample data according to the instruction through the label tagging module, and then outputs the labeled sample data structure. Based on this sample data structure, sample label statistics can be performed to obtain the label distribution frequency of the sample data in the text dataset.

[0065] When the terminal device extracts the quantization indexes of multiple dimensions of the text dataset based on the weighted frequency label tree, by combining the weighted frequency label tree and the sample label frequency information (the label distribution frequency of the sample data in the text dataset) of the text dataset, the target dimension indexes are calculated respectively from the aspects of domain diversity, instruction diversity, complexity, and Chinese-English ratio, that is, the domain diversity index is calculated from the domain diversity dimension, the instruction diversity index is calculated from the instruction diversity dimension, the complexity index is calculated from the complexity dimension, and the Chinese-English ratio index is calculated from the Chinese-English ratio dimension. Among them, the terminal device can calculate the corresponding indexes from any two or more of the above four dimensions. In this way, the terminal device can obtain the quantization indexes of multiple dimensions of the entire text dataset, that is, the quantization indexes of at least two dimensions.

[0066] In some embodiments, the terminal device can use the index calculation module in the integrated multi-dimensional data analysis framework with the weighted frequency label tree and the sample label frequency information as the input, so as to perform the processing of information entropy, instruction distribution comparison, complexity calculation, and Chinese-English ratio statistics based on this index calculation module, and obtain the four-dimensional indexes of domain diversity D_domain, instruction diversity D_instruction, complexity D_complexity, and Chinese-English ratio D_lang.

[0067] In some embodiments, after the terminal device obtains the quantization metrics of multiple dimensions of the text data set, when evaluating the quality of the text data set based on these metrics, these metrics can be used to comprehensively quantify the distribution characteristics of the data set from different dimensions. For example, based on the domain diversity metric, it can be identified whether there is a problem of "uneven domain coverage (such as some topics being too concentrated)" in the data set; based on the instruction diversity metric, it can be identified whether there is a problem of "too single instruction type, which will affect multi-task generalization" in the data set; based on the complexity metric, it can be identified whether there is a problem of "unreasonable text structure complexity, which will affect the model fitting ability" in the data set; and / or, based on the Chinese-English ratio metric, it can be identified whether there is a problem of "imbalanced language distribution, which will affect the Chinese-English multilingual ability" in the data set. In this way, by extracting quantization metrics from the four dimensions of domain diversity, instruction diversity, complexity, and Chinese-English ratio, the terminal device can establish a system evaluation mechanism for the structural rationality, training adaptability, and generalization ability of the text data set through these metrics, thus providing an important basis for supporting data governance and data optimization.

[0068] In the embodiments of the present application, the text data set to be evaluated for quality is obtained through the terminal device, and then the text data set is processed by combining hierarchical clustering and large model semantic extraction respectively, that is, by performing staggered clustering analysis on the text data set to obtain the clustering result of the text data set, and, based on a preset large model, extracting the semantic topics of the sample data in the text data set to obtain the semantic topics of the sample data, and further constructing a label tree structure that can reflect the semantic hierarchy, that is, a multi-layer label tree, based on the clustering result and the semantic topics. And, for each node in the multi-layer label tree, the weight of each node is calculated by combining the label distribution of the text data set, so as to generate a weighted frequency label tree. Finally, based on the weighted frequency label tree, metric calculations are performed from multiple dimensions, so as to extract quantization metrics of multiple dimensions of the text data set, and based on the extracted quantization metrics, an independent evaluation is performed on the text data set in terms of semantic coverage, task structure, language balance, and text complexity, so as to constitute a panoramic description of the quality of the text data set.

[0069] Compared with the method of evaluating the quality of a large-scale data set from only a single dimension, in the embodiments of the present application, hierarchical clustering analysis is performed on the large-scale text data set to be evaluated for quality, and topic extraction is performed on the sample data therein based on a preset large model. Thus, a multi-layer label tree capable of reflecting the semantic hierarchy is constructed based on the clustering results obtained from the hierarchical clustering analysis and the semantic topics obtained from the topic extraction. Then, the weights of each node of the multi-layer label tree are calculated through the label distribution of the text data set to generate a weighted frequency label tree. Finally, quantitative indicators of multiple dimensions of the text data set are extracted based on the weighted frequency label tree for evaluating the quality of the text data set. In this way, in the embodiments of the present application, by combining hierarchical clustering of the large-scale data set with semantic extraction of the large model, a label tree structure that can reflect the semantic hierarchy is constructed, and a weighted frequency label tree is further constructed. On this basis, quantitative indicators are extracted from multiple dimensions respectively to evaluate the quality of the data set. For example, by simultaneously analyzing domain diversity, instruction coverage, complexity, and language ratio, a complete data set portrait is constructed, so as to comprehensively analyze the large-scale data set to perform refined management and optimization of data quality, thereby improving the adaptability, balance, and quality control ability of the data set in the multi-task large model training scenario.

[0070] In some embodiments, the terminal device can perform the above operations on both the candidate text data set and the reference text data set to obtain the weighted frequency label tree. That is, the above-mentioned text data set includes the reference text data set and the candidate text data set, and the above-mentioned weighted frequency label tree includes the first weighted frequency label tree of the reference text data set and the second weighted frequency label tree of the candidate text data set. Then, the terminal device performs a structured comparison analysis between the candidate text data set and the reference text data set based on the two weighted frequency label trees, so as to improve the adaptability, balance, and quality control ability of the data set in the multi-task model training scenario through the structured comparison analysis between the candidate text data set and the reference text data set.

[0071] It should be noted that the reference text data set is usually a mature existing data set, such as the training corpus of a certain large model. It can be considered that the reference text data set has an ideal data distribution. The structure of the first weighted frequency label tree represents a control sample with an ideal data distribution. The candidate text data set can be the data set whose quality is to be evaluated currently. The second weighted frequency label tree expresses the structural characteristics of the data set whose quality is to be evaluated currently.

[0072] Please refer to Figure 2 , Figure 2 which is the schematic flow chart of the steps in other embodiments of the data set quality evaluation method provided by the embodiments of the present application.

[0073] As Figure 1As shown, in some embodiments, the quality assessment method for the dataset provided by the embodiments of the present application may further include steps S201 and S202 as shown below.

[0074] Step S201: Perform cross-entropy comparison on the first weighted frequency tag tree and the second weighted frequency tag tree to obtain the structural similarity between the first weighted frequency tag tree and the second weighted frequency tag tree.

[0075] The terminal device can, through the operations described in the above embodiments, respectively process the candidate text dataset and the reference text dataset to generate the first weighted frequency tag tree of the reference text dataset and the second weighted frequency tag tree of the candidate text dataset. After that, the terminal device obtains the structural similarity between the first weighted frequency tag tree and the second weighted frequency tag tree through the process of cross-entropy comparison of the first weighted frequency tag tree and the second weighted frequency tag tree.

[0076] In some embodiments, the terminal device can use the cross-set comparison module in the integrated multi-dimensional data analysis framework, with the reference text dataset tag tree (the first weighted frequency tag tree) and the candidate text dataset tag tree (the second weighted frequency tag tree) as inputs, and then perform cross-entropy calculation and normalized output similarity pair processing on the two weighted frequency tag trees based on the cross-set comparison module to obtain the normalized cross-entropy index and output it. Among them, the normalized cross-entropy index output by the cross-set comparison module is the structural similarity between the first weighted frequency tag tree and the second weighted frequency tag tree.

[0077] Step S202: Evaluate the structural distribution difference between the candidate text dataset and the reference text dataset based on the structural similarity to obtain the structural distribution difference evaluation result; the structural distribution difference evaluation result is one of the results for evaluating the quality of the candidate text dataset.

[0078] After the terminal device obtains the structural similarity between the first weighted frequency tag tree and the second weighted frequency tag tree, it performs a structured comparison analysis between the candidate text dataset and the reference text dataset based on this structural similarity to evaluate the structural distribution difference between the two datasets, thereby obtaining the structural distribution difference evaluation result. Thus, the structural distribution difference evaluation result can also be used as one of the results obtained by evaluating the quality of the candidate text dataset.

[0079] In this embodiment, based on the construction of a weighted frequency tag tree for the text dataset by the terminal device, further through the calculation of the weighted frequency tag tree and cross-entropy, the distribution similarity analysis between the candidate text dataset and the benchmark text dataset can be realized at multiple levels, thus supporting cross-dataset structured comparison and providing a quantitative basis for data selection and optimization. That is, by evaluating the structural distribution differences between different datasets (candidate text dataset and benchmark text dataset), the selection and optimization of cross-corpus source data are supported.

[0080] In some embodiments, when the terminal device performs topic extraction based on a preset large model, a text prompt can be designed in advance, and then based on this text prompt, the large model can be guided to accurately extract the semantic topic.

[0081] Please refer to Figure 3 , Figure 3 For Figure 1 the detailed step flow diagram of step S102 in

[0082] As Figure 3 shown, in some embodiments, the step of "performing topic extraction on the sample data in the text dataset based on a preset large model" in the above step S102 may include the following step S301 and step S302.

[0083] Step S301: Obtain a preset text prompt; the text prompt includes at least the processing requirements of the data and the output requirements of the data processing result.

[0084] It should be noted that the processing requirements of the data in the text prompt can be what kind of problems the large model needs to solve for the input data, for example, summarizing the topic, etc. In addition, the output requirements of the data processing result in the text prompt can be the output form required by the large model for the result obtained by solving the problem, for example, outputting clear and concise topic words or phrases, etc. For example, the text prompt can be "You are a linguistics expert. Please summarize the common topic of the following text collection, and it is required to output clear and concise topic words or phrases".

[0085] The terminal device can generate a text prompt by combining the processing requirements of the data and the output requirements of the data processing result obtained in real time with the acquired text dataset during the process of quality assessment of the text dataset.

[0086] In some embodiments, the terminal device can also receive a text prompt artificially designed by the staff through a preset data interface, such as a human-computer interaction interface.

[0087] Step S302: Based on the text prompt, guide the pre-trained large model to extract the theme of the sample data in the text dataset according to the processing requirements, and output the first semantic theme obtained by the extraction according to the output requirements; the semantic theme of the sample data includes the first semantic theme.

[0088] After the terminal device obtains the text prompt, it inputs the text prompt and the text dataset to be subjected to theme extraction into the previously selected large model together, so as to guide the large model to perform theme extraction on the sample data in the text dataset based on the text prompt. Among them, after receiving the input text dataset and text prompt, the large model performs theme extraction on the sample data in the text dataset based on the data processing requirements in the text prompt, and after obtaining the first semantic theme of the sample data, it further outputs the first semantic theme according to the data processing result output requirements in the text prompt. In this way, the terminal device can use the first semantic theme output by the large model as the semantic theme of the sample data in the text dataset.

[0089] In some embodiments, the terminal device can also perform hierarchical clustering on the text dataset, and then perform theme extraction on each cluster in the clustering result based on a preset large model, so as to use the theme generated by the large model as the semantic theme of the sample data in the text dataset.

[0090] Based on this, after the step of "performing hierarchical clustering analysis on the text dataset to obtain the clustering result of the text dataset" in the above step S102, the dataset quality evaluation method provided by the embodiments of the present application may further include the following steps: Based on a preset text prompt, guide the pre-trained large model to generate a second semantic theme for the clusters in the clustering result, and use the second semantic theme as the semantic theme of the sample data in the text dataset.

[0091] It should be noted that the text prompt here may be the same as the above text prompt. Alternatively, the terminal device can also design the processing requirements of the data and the output requirements of the data processing results in real time based on the clustering result of the text dataset, so as to generate a text prompt.

[0092] After the terminal device performs hierarchical clustering on the text dataset to obtain the clustering result of the text dataset, it further inputs the clustering result and a preset text prompt into the previously selected large model together, so as to guide the large model to directly perform theme extraction on each cluster in the input clustering result to generate a second semantic theme for each cluster. In this way, the terminal device can use the second semantic theme generated and output by the large model as the semantic theme of the sample data in the text dataset.

[0093] It should be noted that for the step of extracting the second semantic theme for the clustering clusters in the clustering result based on the large model described herein, the terminal device can choose any one of the steps S302 described above to execute. Or, in order to improve the accuracy, the terminal device can also choose to execute both of these steps, and then select the semantic theme with higher accuracy from the first semantic theme and the second semantic theme as the final semantic theme of the sample data in the text dataset. For example, the terminal device can display the first semantic theme and the second semantic theme through the human-computer interaction interface, so as to use the semantic theme selected by the staff for manual sampling inspection as the final semantic theme of the sample data.

[0094] In some embodiments, when the terminal device extracts the theme for the clustering clusters in the clustering result based on the large model, it can merge and summarize the sample content in each clustering cluster as the input of the large model, so that the large model directly processes the summary under the guidance of the text prompt to generate the semantic theme of the clustering cluster.

[0095] In other embodiments, when the terminal device extracts the theme for the clustering clusters in the clustering result based on the large model, it can also directly select the representative sample corresponding to the clustering cluster from the text dataset as the input of the large model, so that the large model processes the representative sample under the guidance of the text prompt to generate the semantic theme of the clustering cluster.

[0096] In this embodiment, the terminal device obtains the text prompt, and then guides the large model to extract the theme to generate the semantic theme based on the text prompt, so as to combine the hierarchical clustering process for the text dataset with the large model semantic extraction process for the text dataset to construct a label tree structure that can reflect the semantic hierarchy for subsequent generation of the weighted frequency label tree and calculation of the quantization indexes of multiple dimensions of the text dataset. Thus, through this structural modeling method, the data distribution has hierarchical interpretability, which has obvious advantages compared with the flat label statistics method in the related technology.

[0097] In some embodiments, when the terminal device calculates the weights of the nodes in the multi-layer label tree to generate the weighted frequency label tree, the terminal device can calculate the weight of each node based on the frequency and depth of each node in the multi-layer label tree.

[0098] Please refer to Figure 4 , Figure 4 For Figure 1 the detailed step flow schematic diagram of step S103 in

[0099] Such as Figure 4As shown, in some embodiments, in the above step S103, the step of "calculating the weights of each node of the multi-level label tree based on the label distribution of the text dataset to generate a weighted frequency label tree" may include steps S401 to S403 as shown below.

[0100] Step S401: Determine the frequency of each node in the multi-level label tree based on the label distribution of the text dataset.

[0101] When the terminal device calculates the weight for each node in the multi-level label tree, it first determines the frequency of the node for which the weight is currently to be calculated in the multi-level label tree based on the label distribution of the text dataset.

[0102] In some embodiments, the terminal device can "match" and "map" the labels of the sample data in the text dataset with each node in the multi-level label tree - that is, determine which topic cluster (node) or multiple topic branches (node set) in the multi-level label tree the sample belongs to, that is, the attribution binding from the sample data to the multi-level label tree structure. In this way, the terminal device can, based on the frequency of the sample data, statistically obtain the frequency of each node in the multi-level label tree.

[0103] In some embodiments, the terminal device can, before calculating the weight for each node in the multi-level label tree, first label the sample data in the text dataset to obtain the label distribution of the text dataset. Then, when calculating the weight for each node in the multi-level label tree, the terminal device can, based on this label distribution, obtain the frequency of the sample data in the text dataset, and further statistically obtain the frequency of each node in the multi-level label tree.

[0104] Based on this, the method for evaluating the quality of the dataset provided in the embodiments of the present application may further include the following steps: Add labels to the sample data in the text dataset, and determine the label distribution of the text dataset based on the labels of the sample data; the labels of the sample data correspond to the nodes in the multi-level label tree.

[0105] The terminal device can perform multi-label annotation on each piece of sample data in the text dataset according to the domain, or perform single-label annotation on each piece of sample data according to an instruction (such as the instruction type). Among them, the labels added by the terminal device to the sample data in the text dataset are "matched" and "mapped" with the nodes (topic nodes) extracted based on the large model and named as the multi-level label tree, that is, the labels of the sample data correspond to the nodes in the multi-level label tree. In this way, the terminal device can obtain the overall label distribution of the text dataset by statistically counting the labels of each piece of sample data in the text dataset.

[0106] In this case, the above step S401: determining the frequency of each node in the multi-layer label tree based on the label distribution of the text dataset may include the following steps: Determine the frequency of the sample data in the text dataset based on the label distribution of the text dataset; Based on the frequency of the sample data, count the frequency of each node in the multi-layer label tree.

[0107] After the terminal device determines the label distribution of the text dataset and then determines the frequency of the nodes in the multi-layer label tree based on this label distribution, the terminal device may first determine the frequency of the sample data in the text dataset based on this label distribution. Then, based on the frequency of this sample data, count the frequency of the nodes in the multi-layer label tree corresponding to the labels of this sample data.

[0108] Step S402: Calculate the weight of each node based on the frequency of each node and the depth of each node in the multi-layer label tree to obtain the weight of each node.

[0109] When the terminal device calculates the weight for each node in the multi-layer label tree, after determining the frequency of the node for which the weight is to be calculated in the multi-layer label tree, the following formula 2 can be further used to calculate the weight by combining the frequency fi of node i and the depth (level) Li of this node i in the multi-layer label tree, so as to obtain the weight of this node i .

[0110] , formula 2 where both α and β are adjustment coefficients.

[0111] Step S403: Perform normalization processing on the weight of each node to generate a weighted frequency label tree; wherein, the normalization processing includes horizontal normalization processing and vertical normalization processing. The horizontal normalization processing is used to perform normalization calculation on the weights of the nodes belonging to the same level in the multi-layer label tree, and the vertical normalization processing is used to perform normalization calculation on the weights of all the nodes in the multi-layer label tree.

[0112] When the terminal device calculates the weight for each node in the multi-layer label tree, after calculating the weight of each node in the multi-layer label tree, further generate a weighted frequency label tree by combining the horizontal and vertical normalization methods. That is, by performing horizontal normalization processing and vertical normalization processing on the weight of each node, a weighted frequency label tree is obtained. Among them, performing horizontal normalization processing on the weight of each node means: performing normalization calculation on the weights of the nodes belonging to the same level in the multi-layer label tree, and performing vertical normalization processing on the weight of each node means: performing full-tree normalization calculation on the weights of each node in the multi-layer label tree.

[0113] In some embodiments, when the terminal device performs horizontal normalization on the weights of each node in the multi-layer label tree, the following formula 3 can be used to perform weight normalization calculation on the node set of each layer in the multi-layer label tree, so that the sum of the weights of the node set of each layer in the multi-layer label tree is 1.

[0114] , formula 3.

[0115] Among them, represents the weight value of node i after horizontal normalization, j represents the index variable of other nodes in the same layer as node i, and N l represents the set of all nodes in the l-th layer, and w j represents the initial weight of node j in the l-th layer before horizontal normalization.

[0116] In some embodiments, when the terminal device performs vertical normalization on the weights of all nodes in the multi-layer label tree, the following formula 4 can be used to perform weight normalization calculation on the node set in the entire multi-layer label tree, so that the overall weight is normalized.

[0117] , formula 4.

[0118] Among them, represents the weight value of node i after vertical normalization, K is the total number of nodes in the label tree, represents the weight of node j in the l-th layer after horizontal normalization.

[0119] In this embodiment, the terminal device calculates the weight based on the frequency of each node in the multi-layer label data and the depth of each node in the multi-layer label tree, and introduces horizontal and vertical normalization methods to generate a weighted frequency label tree. Thus, in the subsequent process, through the weighted frequency label tree and cross-entropy calculation, the distribution similarity analysis of different data sets on multiple levels can be realized, and further a quantitative basis can be provided for data selection and optimization. Moreover, by constructing the weighted frequency label tree, it is also possible to combine the previously constructed multi-layer label tree, and in the subsequent process, calculate multi-dimensional indicators such as domain diversity, instruction diversity, complexity, and Chinese-English ratio, and perform cross-entropy-based normalized distribution similarity comparison between different data sets through the weighted frequency label tree and cross-entropy calculation. The architecture is designed independently from structure modeling, index extraction to cross-set analysis, constituting an integrated multi-dimensional data analysis framework that is evaluable, interpretable, and comparable, so as to realize comprehensive analysis of large-scale data sets for refined management and optimization of data quality, thereby improving the adaptability, balance, and quality control ability of the data set in the multi-task large model training scenario.

[0120] Next, a complete embodiment of the quality evaluation method for the dataset provided in the embodiments of the present application is presented.

[0121] Please refer to Figure 5 , Figure 5 which is a schematic diagram of the multi-dimensional dataset distribution analysis process of hierarchical clustering and weighted label tree involved in a complete embodiment of the quality evaluation method for the dataset provided in the embodiments of the present application.

[0122] As Figure 5 shown, for the quality evaluation method for the dataset provided in the embodiments of the present application, after obtaining the original text data (text dataset) to be evaluated for quality through a terminal device, the original text data is input into a clustering analysis module. Then, the clustering analysis module performs hierarchical clustering and large model topic extraction on the original text data, and outputs preliminary clustering labels and topic representation vectors for each cluster. After that, the clustering results obtained by the clustering analysis module for hierarchical clustering of the original text data and the semantic topics extracted by the large model are input into a label tree construction module through the terminal device. The label tree construction module constructs a multi-layer label tree, calculates the frequency and depth of each node in the multi-layer label tree, and outputs a label tree structure with frequency information. Also, the original text data is input as the input of a label tagging module through the terminal device. Then, the label tagging module performs multi-label tagging according to the domain, and / or performs single-label annotation according to the instruction, and then outputs a labeled sample data structure. After that, the label tree structure with frequency information and the labeled sample data structure are input into a weight calculation and normalization module through the terminal device. The weight calculation and normalization module calculates the weight of each node, and generates a weighted frequency label tree by performing horizontal / vertical normalization on the label tree. Finally, the weighted frequency label tree is input into an index calculation module through the terminal device for processing such as information entropy calculation, instruction distribution comparison, complexity calculation, and Chinese-English ratio statistics, to obtain four-dimensional indicators including domain diversity D_domain, instruction diversity D_instruction, complexity D_complexity, and Chinese-English ratio D_lang. Also, the cross-entropy calculation and normalization output similarity pair processing are performed by the cross-set comparison module of the terminal device using the weighted frequency label tree of the benchmark text dataset label tree (the weighted frequency label tree of the benchmark text dataset with ideal data distribution) and the candidate text dataset label tree (the weighted frequency label tree corresponding to the original text data), to obtain a normalized cross-entropy indicator, and thus an evaluation report and similarity indicator for quality evaluation of the original text data are output.

[0123] In some embodiments, the quality evaluation method for the dataset provided by the embodiments of the present application can, through the cooperation of the terminal device based on the extraction of the large model theme and the tagging, "match" and "map" the tags marked by the tagging module for each sample (the sample data in the original text data) with the theme nodes extracted by the large model, so as to complete the attribution binding of the sample to the tag tree structure.

[0124] As Figure 6 shown, the processing of the clustering cluster samples (the sample content in the clustering cluster in the clustering result obtained by hierarchical clustering of the original text) by the terminal device to extract the tag tree theme nodes through the large model processing, and inputting the clustering cluster samples into the tagging module to obtain the sample tag information, and then combining the tag tree theme nodes and the sample tag information to construct a tag tree to achieve the binding between the sample and the nodes in the tag tree. And, the terminal device can also perform the processing of semantic consistency verification on the tag tree structure (multi-layer tag tree) after the binding of the sample and the node is completed. For example, the terminal device can display the tag tree structure through a preset human-computer interaction interface, so that the staff can perform semantic consistency verification.

[0125] In some embodiments, the quality evaluation method for the dataset provided by the embodiments of the present application can also, through the terminal device, according to the process of measuring the data quality based on the four major indicators as Figure 7 shown, first input the tagged sample data (such as the tagged sample data structure output by the tagging module), construct a weighted tag tree (weighted frequency tag tree), and then synchronously calculate the task type frequency, the domain tag frequency and the structural entropy, the complexity value of each sample, and count the number of Chinese and English samples, and compare the result obtained by calculating the task type frequency with the benchmark set to obtain the task coverage index D_instruction (instruction diversity index), fuse the node weights with the result obtained by calculating the domain tag frequency and the structural entropy to obtain the domain diversity index D_domain, compare the result obtained by calculating the complexity value of each sample with the standard complexity c-std to obtain the complexity index D_complexity, and calculate the language ratio deviation based on the result obtained by counting the number of Chinese and English samples to obtain the language structure balance index D_lang (Chinese-English ratio index). Finally, the terminal device comprehensively calculates the four-dimensional index and outputs a multi-dimensional index evaluation report for the original text data.

[0126] Please refer to Figure 8 , the embodiments of the present application also provide a quality evaluation device for the dataset, which can implement the above quality evaluation method for the dataset.

[0127] As Figure 8As shown in the figure, the quality evaluation device for the dataset provided by the embodiment of the present application includes an acquisition module 801, a clustering and model topic extraction module 802, a label tree construction module 803, and a dataset quality evaluation module 804. Among them, The acquisition module 801 is used to acquire a text dataset whose quality is to be evaluated; The clustering and model topic extraction module 802 is used to perform hierarchical clustering analysis on the text dataset to obtain the clustering result of the text dataset, and to extract the topic of the sample data in the text dataset based on a preset large model to obtain the semantic topic of the sample data; The label tree construction module 803 is used to construct a multi-layer label tree based on the clustering result and the semantic topic, and calculate the weights of each node of the multi-layer label tree based on the label distribution of the text dataset to generate a weighted frequency label tree; The dataset quality evaluation module 804 is used to extract quantitative indicators of multiple dimensions of the text dataset based on the weighted frequency label tree to evaluate the quality of the text dataset.

[0128] In some embodiments, the clustering and model topic extraction module 802 is further used to obtain a preset text prompt; the text prompt at least includes the processing requirements of the data and the output requirements of the data processing result; and, based on the text prompt, guide the pre-trained large model to extract the topic of the sample data in the text dataset according to the processing requirements, and output the first semantic topic obtained by the extraction according to the output requirements; the semantic topic of the sample data includes the first semantic topic.

[0129] In some embodiments, the clustering and model topic extraction module 802 is further used to guide the pre-trained large model to generate a second semantic topic for the clustering clusters in the clustering result based on a preset text prompt, and use the second semantic topic as the semantic topic of the sample data in the text dataset.

[0130] In some embodiments, the clustering result includes a tree-shaped clustering result; the label tree construction module 803 is further used to generate a semantic node name of the target node in the tree-shaped clustering result based on the semantic topic, so as to construct the tree-shaped clustering result into a multi-layer label tree.

[0131] In some embodiments, the label tree construction module 803 is further configured to determine the frequency of each node in the multi-level label tree based on the label distribution of the text data set; calculate weights based on the frequency of each node and the depth of each node in the multi-level label tree to obtain the weight of each node; and perform normalization processing on the weight of each node to generate a weighted frequency label tree; wherein, the normalization processing includes horizontal normalization processing and vertical normalization processing, and the horizontal normalization processing is used to perform normalization calculation on the weights of the nodes belonging to the same level in the multi-level label tree, and the vertical normalization processing is used to perform normalization calculation on the weights of all the nodes in the multi-level label tree.

[0132] In some embodiments, the label tree construction module 803 is further configured to add labels to the sample data in the text data set, and determine the label distribution of the text data set based on the labels of the sample data; the labels of the sample data correspond to the nodes in the multi-level label tree; and determine the frequency of the sample data in the text data set based on the label distribution of the text data set; and count the frequency of each node in the multi-level label tree based on the frequency of the sample data.

[0133] In some embodiments, the data set quality evaluation module 804 is further configured to calculate target dimension indicators based on the weighted frequency label tree and the sample label frequency information of the text data set to obtain quantization indicators for multiple dimensions of the text data set; wherein, the sample label frequency information includes the frequency of the sample data in the text data set; and the target dimension indicators include at least two of a domain diversity indicator, an instruction diversity indicator, a complexity indicator, and a Chinese-English ratio indicator.

[0134] In some embodiments, the text data set includes a reference text data set and a candidate text data set, and the weighted frequency label tree includes a first weighted frequency label tree of the reference text data set and a second weighted frequency label tree of the candidate text data set; the data set quality evaluation module 804 is further configured to perform cross-entropy comparison on the first weighted frequency label tree and the second weighted frequency label tree to obtain the structural similarity between the first weighted frequency label tree and the second weighted frequency label tree; and evaluate the structural distribution difference between the candidate text data set and the reference text data set based on the structural similarity to obtain a structural distribution difference evaluation result; the structural distribution difference evaluation result is one of the results for evaluating the quality of the candidate text data set.

[0135] Please refer to Figure 9 , Figure 9 which schematically shows the hardware structure of a data set quality evaluation device according to an embodiment. The data set quality evaluation device includes: The processor 901 can be implemented in ways such as a general - purpose CPU (Central Processing Unit), a microprocessor, an application - specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of the present application; The memory 902 can be implemented in forms such as a read - only memory (ROM), a static storage device, a dynamic storage device, or a random - access memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided in the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902 and are called by the processor 901 to execute the quality evaluation method of the data set in the embodiments of the present application; The input / output interface 903 is used to implement information input and output; The communication interface 904 is used to implement communication interaction between this device and other devices. It can achieve communication through wired means (such as USB, network cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.); The bus 905 transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904); Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 achieve communication connections with each other inside the device through the bus 905.

[0136] The embodiments of the present application also provide a vehicle. A data - set quality evaluation device is configured on the vehicle. The data - set quality evaluation device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above - mentioned data - set quality evaluation method is implemented.

[0137] The embodiments of the present application also provide a computer - readable storage medium. The computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, the above - mentioned data - set quality evaluation method is implemented.

[0138] The memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include a memory remotely disposed relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0139] The embodiments of the present application also provide a computer program product, including a computer program, and the steps implemented when the computer program is executed by a processor are substantially the same as the specific embodiments of the above data set quality evaluation method, and will not be described in detail herein.

[0140] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0141] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than those shown in the figures, or combine certain steps, or different steps.

[0142] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0143] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, and appropriate combinations thereof.

[0144] In the description of this application and the above-mentioned accompanying drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0145] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expressions refer to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0146] In several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0147] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0148] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0149] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The aforementioned storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0150] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings. However, this does not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall fall within the scope of the rights of the embodiments of the present application.

Claims

1. A method for evaluating the quality of a data set, characterized in that, The method includes: Obtaining a text dataset whose quality is to be evaluated; Performing hierarchical clustering analysis on the text dataset to obtain the clustering result of the text dataset, and extracting the semantic topics of the sample data in the text dataset based on a preset large model to obtain the semantic topics of the sample data; Constructing a multi-layer label tree based on the clustering result and the semantic topics, and calculating the weights of each node of the multi-layer label tree based on the label distribution of the text dataset to generate a weighted frequency label tree; Extracting quantitative indicators of multiple dimensions of the text dataset based on the weighted frequency label tree to evaluate the quality of the text dataset.

2. The method according to claim 1, characterized in that The extracting the semantic topics of the sample data in the text dataset based on a preset large model includes: Obtaining a preset text prompt; the text prompt includes at least the processing requirements of the data and the output requirements of the data processing result; Based on the text prompt, guiding the pre-trained large model to extract the semantic topics of the sample data in the text dataset according to the processing requirements, and outputting the first semantic topic extracted according to the output requirements; the semantic topics of the sample data include the first semantic topic.

3. The method according to claim 1, wherein After performing hierarchical clustering analysis on the text dataset to obtain the clustering result of the text dataset, the method further includes: Based on a preset text prompt, guiding the pre-trained large model to generate a second semantic topic for the clustering clusters in the clustering result, and using the second semantic topic as the semantic topic of the sample data in the text dataset.

4. The method according to claim 1, wherein The clustering result includes a tree-shaped clustering result; the constructing a multi-layer label tree based on the clustering result and the semantic topics includes: Generating the semantic node names of the target nodes in the tree-shaped clustering result based on the semantic topics to construct the tree-shaped clustering result into a multi-layer label tree.

5. The method according to claim 1, wherein The calculating the weights of each node of the multi-layer label tree based on the label distribution of the text dataset to generate a weighted frequency label tree includes: Determining the frequency of each node in the multi-layer label tree based on the label distribution of the text dataset; Calculating the weights of each node based on the frequency of each node and the depth of each node in the multi-layer label tree to obtain the weights of each node; Performing normalization processing on the weights of each node to generate a weighted frequency label tree; Wherein, the normalization processing includes horizontal normalization processing and vertical normalization processing. The horizontal normalization processing is used to perform normalization calculation on the weights of the nodes belonging to the same level in the multi-layer label tree, and the vertical normalization processing is used to perform normalization calculation on the weights of all the nodes in the multi-layer label tree.

6. The method according to claim 5, wherein The method further includes: Adding labels to the sample data in the text dataset, and determining the label distribution of the text dataset based on the labels of the sample data; the labels of the sample data correspond to the nodes in the multi-layer label tree; The determining the frequency of each node in the multi-layer label tree based on the label distribution of the text dataset includes: Determining the frequency of the sample data in the text dataset based on the label distribution of the text dataset; Statistically count the frequency of each node in the multi - layer label tree based on the frequency of the sample data.

7. The method according to claim 1, wherein Extract quantitative indicators of multiple dimensions of the text data set based on the weighted frequency label tree, including: Calculate the target - dimension indicators based on the weighted frequency label tree and the sample - label frequency information of the text data set to obtain quantitative indicators of multiple dimensions of the text data set; Among them, the sample - label frequency information includes the frequency of sample data in the text data set; the target - dimension indicators include at least two of the domain - diversity indicator, instruction - diversity indicator, complexity indicator, and Chinese - English ratio indicator.

8. The method according to any one of claims 1 to 7, characterized in that, The text data set includes a benchmark text data set and a candidate text data set, and the weighted frequency label tree includes a first weighted frequency label tree of the benchmark text data set and a second weighted frequency label tree of the candidate text data set; The method further includes: Perform cross - entropy comparison on the first weighted frequency label tree and the second weighted frequency label tree to obtain the structural similarity between the first weighted frequency label tree and the second weighted frequency label tree; Evaluate the structural distribution difference between the candidate text data set and the benchmark text data set based on the structural similarity to obtain a structural - distribution - difference evaluation result; the structural - distribution - difference evaluation result is one of the results for evaluating the quality of the candidate text data set.

9. A quality evaluation device for a data set, characterized in that, The device includes: An acquisition module, configured to acquire a text data set whose quality is to be evaluated; A clustering and model - topic extraction module, configured to perform hierarchical clustering analysis on the text data set to obtain a clustering result of the text data set, and extract the semantic topics of the sample data in the text data set based on a preset large - model to obtain the semantic topics of the sample data; A label - tree construction module, configured to construct a multi - layer label tree based on the clustering result and the semantic topics, and calculate the weights of each node of the multi - layer label tree based on the label distribution of the text data set to generate a weighted frequency label tree; A data - set quality evaluation module, configured to extract quantitative indicators of multiple dimensions of the text data set based on the weighted frequency label tree to evaluate the quality of the text data set.

10. A quality evaluation device for a data set, characterized in that, The quality - evaluation device of the data set includes a memory and a processor. When the processor executes the computer program, it implements the data - set quality - evaluation method according to any one of claims 1 to 8.

11. A computer-readable storage medium, characterized in that, The computer - readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the data - set quality - evaluation method according to any one of claims 1 to 8.

12. A computer program product, characterized in that, The computer - program product includes a computer program, and when the computer program is executed by a processor, it implements the data - set quality - evaluation method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • News occurrence place identification method and device, storage medium and computing equipment

    CN113268680A

  • Text analysis method and device, equipment and storage medium

    CN118467733A

  • Information collection and classification system based on big data

    CN118536008A

  • Technical literature multi-dimensional analysis method based on semantic understanding

    CN119128143A

  • Text clustering enhanced visual analysis method and device based on large language model

    CN119293254A