Data set evaluation method and device, electronic equipment and storage medium

By establishing a feature index system in a multimodal data set and evaluating the index value using spatial fill curves, the problem of how to effectively evaluate the value of new data in a multimodal data set is solved, and the completeness and quality evaluation of the data set is achieved to avoid data redundancy.

CN120144575APending Publication Date: 2025-06-13中国邮政储蓄银行股份有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510242695.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-03
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the fields of artificial intelligence and big data, how to effectively evaluate the value of new data in multimodal data sets to ensure the completeness and quality of the data set.

Method used

By entering the data set to be evaluated into the feature index module, a feature index system is established, and the index value is calculated using the spatial fill curve, the completeness of the data set and the value of the added data are evaluated.

Benefits of technology

The completeness evaluation of multimodal data sets and the evaluation of new data value are realized to ensure the quality and efficiency of the data sets and avoid data redundancy and waste of storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120144575A_ABST
    Figure CN120144575A_ABST
Patent Text Reader

Abstract

The invention discloses a data set evaluation method and device, electronic equipment and a storage medium, and the method comprises the steps: inputting a to-be-evaluated data set into a feature index module, and obtaining an index value; according to the index value, obtaining an evaluation result corresponding to the to-be-evaluated data set; establishing a feature index in the to-be-evaluated data set before judging whether the value evaluation result of the to-be-evaluated newly-added data meets the requirement or not; and if the query result is not repeated and meets a set condition, storing the to-be-evaluated new data, and updating an evaluation result corresponding to the to-be-evaluated data set. According to the method and the device, on one hand, completeness evaluation of the multi-modal data set is realized, and on the other hand, value evaluation on newly added data in the multi-modal data set is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of dataset evaluation, and particularly to a dataset evaluation method, apparatus, electronic device, and storage medium. Background Art

[0002] With the development of large models and AIGC, people have increasing requirements for the data needed to train machine learning models, demanding both larger data volumes and higher data quality. Artificial intelligence / machine learning-driven systems require a large amount of high-quality data during training and evaluation.

[0003] However, in the fields of artificial intelligence and big data, such datasets are often stored in the form of data lakes. Usually, the scale of the dataset can be expanded by continuously adding more and more data. This method is of course simple, but there has always been a lack of an effective evaluation method for determining the value of the next data to be added to the data lake. Summary of the Invention

[0004] Embodiments of this application provide a dataset evaluation method, apparatus, electronic device, and storage medium to achieve the completeness evaluation of multi-modal datasets and the value evaluation of newly added data in multi-modal datasets.

[0005] Embodiments of this application adopt the following technical solutions:

[0006] In a first aspect, embodiments of this application provide a dataset evaluation method, where the evaluation method includes:

[0007] Input the dataset to be evaluated into a feature indexing module to obtain an index value;

[0008] Obtain an evaluation result corresponding to the dataset to be evaluated according to the index value;

[0009] Before judging whether the value degree evaluation result of the newly added data to be evaluated meets the requirements, establish a feature index in the dataset to be evaluated;

[0010] If the query result is not repeated and meets the set conditions, save the newly added data to be evaluated and update the evaluation result corresponding to the dataset to be evaluated. In some embodiments, inputting the dataset to be evaluated into a feature indexing module to obtain an index value includes:

[0011] Input the dataset to be evaluated into the feature indexing module for vectorization processing to obtain multi-dimensional feature vectors of each data entry in the dataset to be evaluated;

[0012] Based on the space plane mapped by the multi-dimensional feature vectors and using a space filling curve, establish the index value.

[0013] In some embodiments, the feature vector includes an N-dimensional feature vector. Establishing the index value according to the space plane mapped by the multi-dimensional feature vector and using a space-filling curve includes:

[0014] Mapping the N-dimensional feature vector to a space plane, completing the operation of unfolding the N-dimensional sphere into an N-1 dimensional plane;

[0015] Constructing a space-filling curve according to the N-1 dimensional plane, establishing the index value, and adding the index value to a table.

[0016] In some embodiments, obtaining the evaluation result corresponding to the dataset to be evaluated according to the index value includes:

[0017] Calculating the total number of indexes according to the relationship between the index value and the maximum index number of the space-filling curve;

[0018] Determining the data coverage parameter of the dataset to be evaluated according to the total number of indexes.

[0019] In some embodiments, before determining whether the evaluation result of the value degree of the new data to be evaluated meets the requirements, establishing a feature index in the dataset to be evaluated includes:

[0020] Establishing an index value for the new data to be evaluated by using a space-filling curve;

[0021] Judging whether there is a repetition between the index value and the index values in the dataset to be evaluated;

[0022] If there is no repetition, use an evaluation module to compare the value degree of the new data to be evaluated with a preset lake-in condition;

[0023] If the value degree of the new data to be evaluated is greater than the preset lake-in condition, the value of the new data reaches the standard for data to enter the lake, and save the new data to be evaluated to the existing dataset.

[0024] In some embodiments, if the query result is not repeated and meets the set conditions, saving the new data to be evaluated and updating the evaluation result corresponding to the dataset to be evaluated includes:

[0025] Taking the evaluation result corresponding to the dataset to be evaluated as the completeness evaluation result of the dataset;

[0026] Taking the evaluation result of the value degree of the new data to be evaluated as the value evaluation result of the new dataset;

[0027] If the completeness of the dataset is insufficient, the dataset is added according to the evaluation result of the value of the newly added data entry;

[0028] If the value of the newly added data to be evaluated is insufficient, a new dataset to be added is reselected.

[0029] In some embodiments, the method further includes:

[0030] If the coverage rate of the dataset reaches A%, a data entry index value is re-established;

[0031] According to the data entry index value, the total index is increased to obtain a new total index;

[0032] According to the new total index, the coverage rate of the dataset is recalculated, where A is a constant.

[0033] In a second aspect, an embodiment of the present application further provides a dataset evaluation device, where the evaluation device includes:

[0034] An indexing module, configured to input a dataset to be evaluated into a feature indexing module to obtain an index value;

[0035] An evaluation module, configured to obtain an evaluation result corresponding to the dataset to be evaluated according to the index value;

[0036] A judgment module, configured to establish a feature index in the dataset to be evaluated before judging whether the evaluation result of the value of the newly added data to be evaluated meets the requirements;

[0037] An update module, configured to save the newly added data to be evaluated and update the evaluation result corresponding to the dataset to be evaluated if the query result is not repeated and meets the set conditions. In a third aspect, an embodiment of the present application further provides an electronic device, including: a processor; and a memory arranged to store computer-executable instructions, where the executable instructions, when executed, cause the processor to execute the above method.

[0038] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, where the computer-readable storage medium stores one or more programs, and when the one or more programs are executed by an electronic device including a plurality of application programs, the electronic device is caused to execute the above method.

[0039] The above at least one technical solution adopted in the embodiments of the present application can achieve the following beneficial effects: By inputting the dataset to be evaluated into the feature index module, a feature index system is established, and then according to the index value, the evaluation result (coverage rate) corresponding to the dataset to be evaluated is obtained. Then, before judging whether the evaluation result of the value of the newly added data to be evaluated meets the requirements, a feature index query is performed in the dataset to be evaluated. Finally, if the query result is non-repetitive and meets the set conditions, the newly added data to be evaluated is saved, and the evaluation result corresponding to the dataset to be evaluated is updated. Through the above method, the completeness evaluation of the multi-modal dataset and the value evaluation of the newly added data in the multi-modal dataset are realized. Through the above method, the completeness of the current full-scale dataset is quantitatively evaluated, and at the same time, the value of the incremental data can be automatically evaluated, and the value of the new data to be inserted into the large-scale data lake is estimated in an efficient calculation manner. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation to the present application. In the drawings:

[0041] Figure 1 is a schematic diagram of the system architecture of the dataset evaluation method in the embodiments of the present application;

[0042] Figure 2 is a schematic diagram of the process of the dataset evaluation method in the embodiments of the present application;

[0043] FIG. 3(a) is a schematic diagram of one of the feature index processes of the dataset evaluation method in the embodiments of the present application;

[0044] FIG. 3(b) is a schematic diagram of another feature index process of the dataset evaluation method in the embodiments of the present application;

[0045] Figure 4 is a schematic diagram in which the unit feature vectors are distributed on the surface of the sphere in the form of points in the embodiments of the present application;

[0046] Figure 5 is a schematic diagram of the structure of the dataset evaluation device in the embodiments of the present application;

[0047] Figure 6 is a schematic diagram of the structure of an electronic device in the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0048] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.

[0049] Multimodality refers to a mode that integrates multiple information expression methods or information sources. Multimodality emphasizes expressing and understanding things more comprehensively and richly by integrating different types of information. Therefore, in a multimodal system, information of different modalities can complement and corroborate each other, thereby improving the effect of information processing and understanding.

[0050] Dataset completeness refers to whether a dataset contains sufficiently comprehensive and complete information in a specific application or research scenario. A complete dataset should cover all relevant and important data points and features without the absence or omission of key information. It should be able to fully reflect all possible situations and states of the object or problem being studied. If a dataset lacks certain key information or situations, it may lead to insufficient or inaccurate model training, thereby affecting the final application effect. Ensuring the completeness of the dataset is of great significance for the effectiveness and reliability of the model. Dataset completeness is a crucial factor in machine learning and has a profound impact on the performance, reliability, stability, and generalization ability of the model.

[0051] Large models generally refer to models with a large number of parameters and complex structures. After being trained on a large amount of data, the models can learn rich knowledge and patterns. They possess powerful language understanding, generation capabilities, and the ability to handle various tasks, such as natural language processing, image recognition, etc. The characteristics of large models include, but are not limited to: a huge number of parameters, possibly reaching billions or more; high requirements for computing resources; the ability to process massive amounts of data; high generality and adaptability, and can be applied to multiple fields and tasks. Common large models include some pre-trained language models in the field of natural language processing. The development of large models has promoted the continuous progress and wide application of artificial intelligence technology.

[0052] AIGC: Artificial Intelligence Generated Content refers to the use of artificial intelligence technology to automatically generate various types of content, such as text, images, audio, video, etc. Through machine learning algorithms and a large amount of data training, artificial intelligence models can imitate the human creative process and thinking mode, and create content with a certain degree of logic, creativity, and uniqueness. AIGC is the application and expansion of artificial intelligence technology in the field of content creation, bringing new possibilities and efficiency to content production.

[0053] A space-filling curve is a special curve that has the property of being able to traverse or approximately traverse all points within a given space in a finite number of steps. Simply put, it is a curve that can gradually fill the entire space area in an orderly manner. Although intuitively it seems impossible for a curve to truly fill a space, through ingenious design and iteration, a space-filling curve can achieve an approximate complete coverage of the space.

[0054] The following will, in conjunction with the accompanying drawings, elaborate in detail on the technical solutions provided by each embodiment of this application.

[0055] As Figure 1 shown, it is a schematic diagram of the system architecture of the dataset evaluation method in an embodiment of this application. The overall system architecture mainly consists of three parts, namely a feature index system, a completeness evaluation module, and a value evaluation module. The feature index system is the basis for the other two modules and mainly includes two components: vectorization and a database. The functional modules in the embodiments of this application include two parts: namely, the completeness evaluation of the current dataset and the value evaluation of new data. Before determining whether the value evaluation result of new data meets the requirements, it is necessary to perform a feature index query in the dataset. After determining that it is non-repetitive and meets the set conditions, the new data is then stored in the current dataset.

[0056] The embodiments of this application provide a dataset evaluation method. As Figure 2 shown, it provides a schematic diagram of the process of the dataset evaluation method in an embodiment of this application. The method at least includes the following steps S210 to step S250:

[0057] Step S210: Input the dataset to be evaluated into the feature index module to obtain an index value.

[0058] The dataset to be evaluated includes, but is not limited to, datasets for large model training, datasets for AIGC training, etc. The number of data in the dataset to be evaluated is basically determined.

[0059] The "feature index module" has the function of vectorizing data and establishing an index and then storing it in the database. The "dataset to be evaluated" is a multi-modal dataset to meet the requirements of training samples.

[0060] Specifically, the implementation of a feature indexing system for a multi-modal dataset is mainly achieved by calculating feature vectors and storing them. Therefore, the vectorization component in the feature indexing system realizes the calculation of feature vectors, while the database component realizes the storage of feature vectors and index values (obtained by space filling curves). As shown in Fig. 3(a), multi-modal data includes, but is not limited to, text data, image data, audio data, and video data. For image data, audio data, and video data, corresponding multiple texts need to be obtained after image understanding, speech understanding, and video understanding respectively. The multiple texts are input into a natural language understanding model together with the text data to extract N-dimensional feature vectors, establish index values, and save them to the database. For data in modalities such as images, audio, and videos, corresponding deep learning models of the modalities can be used first to understand and summarize them into texts and then input them into the large model to obtain feature vectors. If a multi-modal large model is used, data in modalities such as text, image, audio, and video can be directly used as the input of the multi-modal large model, and then the multi-modal large model outputs the feature vectors of the data.

[0061] As shown in Fig. 3(b), multi-modal data includes, but is not limited to, text data, image data, audio data, and video data. A unified multi-modal large model can be used to extract N-dimensional feature vectors, establish index values, and save them to the database. For the dataset, the input data is multi-modal and supports data in forms such as text, image, audio, and video. For data in different modalities, corresponding deep learning models can be used for understanding. For example, for the text model, the large model can be directly used for understanding to output N-dimensional feature vectors. Specifically, the output of a certain layer in the large model neural network can be used as the feature vector.

[0062] Based on the above, the deep learning method is used to extract key vector features with semantics and use them as the input of the space filling curve.

[0063] Step S220, according to the index value, obtain the evaluation result corresponding to the dataset to be evaluated.

[0064] Based on the index value, obtain the completeness evaluation result and obtain the evaluation result of the dataset to be evaluated.

[0065] Step S230, before judging whether the value evaluation result of the new data to be evaluated meets the requirements, establish a feature index in the dataset to be evaluated.

[0066] The new data to be evaluated includes, but is not limited to, the data that needs to be newly added to the dataset. Similarly, the newly added data also needs to be evaluated.

[0067] The process of loading new data into the data lake of the current dataset is as follows: For new data, corresponding indexes need to be established first, and the index values of the new data are checked for duplicates through a preset index table. If duplicates are found, the new data does not need to be added. If there are no duplicates, it is determined whether to enter the lake based on whether the preset conditions are met.

[0068] Step S240, if the query result is non-duplicate and meets the preset conditions, save the new data to be evaluated and update the evaluation result corresponding to the dataset to be evaluated.

[0069] For the dataset to be evaluated, the data completeness index in the existing dataset can be determined in a quantitative manner.

[0070] For new data, an evaluation method for evaluating the value of the new data to this dataset is provided, so as to avoid data redundancy and reduce waste of storage space.

[0071] Through the above method, the evaluation of the completeness of the multi-modal dataset is realized, and at the same time, the value of the new data to the existing dataset (high or low value and whether it is necessary to be added to the dataset) is evaluated. Therefore, it has high practical benefits for constructing a high-quality multi-modal dataset.

[0072] Through the above method, it can be applied to the construction of a high-quality multi-modal dataset. Among different datasets, it can also be used to help select high-quality datasets before machine learning model training. At the same time, from a vertical perspective, a quantitative evaluation method for the value of new data in the same dataset is provided.

[0073] Different from the related technology, which is only applicable to evaluating the quality of the dataset required for supervised machine learning models and is not applicable to the dataset required for unsupervised models. Through the above method, the completeness evaluation of the multi-modal dataset and the value evaluation of the new data in the multi-modal dataset are realized.

[0074] Different from the related technology, which fails to provide a reference for evaluating the value of new data. Through the above method, the completeness of the current full-scale dataset is quantitatively evaluated, and at the same time, the value of incremental data can be automatically evaluated to estimate the value of the new data to be inserted into the large-scale data lake in an efficient calculation manner.

[0075] In an embodiment of the present application, the inputting the dataset to be evaluated into the feature index module to obtain an index value includes: inputting the dataset to be evaluated into the feature index module for vectorization processing to obtain multi-dimensional feature vectors in the dataset to be evaluated; establishing the index value according to the space plane and space filling curve mapped by the multi-dimensional feature vectors.

[0076] In an alternative embodiment, the method for obtaining the index value from the N-dimensional feature vector is as shown in the above figure, where the input N-dimensional feature vector is After that, a unitary transformation is applied to it to obtain a unit eigenvector where ||·|| 2 represents the 2-norm. Since the 2-norm of is 1, which means that the unit eigenvectors of all data

[0077] If in the case of N = 3, the unit eigenvector is distributed on the surface of the spherical space. For a sphere with dimension N>3, visualization is not possible.

[0078] On the premise that is distributed on the surface of the sphere, the feature points can be expanded using the spherical coordinate transformation formula, thus becoming points distributed in the N-1 dimensional plane. The conversion relationship between N-dimensional spherical coordinates and multi-dimensional space points is:

[0079]

[0080] Here, since the points are all on the surface of the sphere, so r = 1.

[0081] By calculating the inverse operation of the above formula, the vector can be solved, that is, the operation of expanding the N-dimensional spherical surface into an N-1 dimensional plane is completed. On this N-1 dimensional plane, an index value can be further obtained by constructing a space-filling curve. That is to say, in a high-dimensional plane, a space-filling curve is used to obtain feature vectors. If the feature vectors are similar, the distance between the feature vectors is very close.

[0082] The above method maps the multi-dimensional features of data entries to the corresponding one-dimensional representation according to the space-filling curve.

[0083] In an embodiment of the present application, the feature vector includes an N-dimensional feature vector. According to the space plane mapped by the multi-dimensional feature vector and using the space-filling curve, the establishment of the index value includes: mapping the N-dimensional feature vector to the space plane, completing the operation of expanding the N-dimensional spherical surface into an N-1 dimensional plane; constructing a space-filling curve according to the N-1 dimensional plane, establishing the index value, and adding the index value to the table.

[0084] As Figure 4As shown, it can be understood that the length N of the N-dimensional feature vector can be defined by the user. Mapping the N-dimensional feature vector to a spatial plane completes the operation of unfolding the N-dimensional sphere into an N-1 dimensional plane. A space-filling curve can be constructed on this N-1 dimensional plane.

[0085] Optionally, a space-filling curve can transform multi-dimensional data into a one-dimensional integer domain and, as much as possible, preserve the characteristics of the multi-dimensional space, such that spaces that are close in the multi-dimensional space are also as close as possible in the transformed integers. Z-curves and Hilbert curves are relatively commonly used space-filling curves, among which Z-curves are relatively easy to implement. XZ-Ordering extends the Z-curve to enable effective description of non-point spatial objects.

[0086] Finally, the system stores the calculated index value in the N-1 dimensional space into the idx column of the table data_item in the database, that is, the index value is stored in the table data_item. The definition of the table data_item is as follows:

[0087]

[0088] In an embodiment of the present application, obtaining the evaluation result corresponding to the data set to be evaluated according to the index value includes: calculating the total number of indexes according to the relationship between the index value and the maximum index number of the space-filling curve in the space; determining the data coverage parameter of the data set to be evaluated according to the total number of indexes.

[0089] In an alternative embodiment, when calculating the total number of indexes, the calculation method of the maximum index number M of each space-filling curve in the space is different. Taking XZ-Ordering as an example, in a 2-dimensional plane, a variable-length quadrant sequence is given, the minimum length of which is 0 and the maximum is g. When the length of the quadrant sequence is l, the maximum number of indexes that can be performed in the space is:

[0090] When calculating the data coverage, calculate the number of different index values in the current table data_item, denoted as C, then the coverage of the current data set is:

[0091]

[0092] The coverage of the data set can be used as a quantitative index to characterize the completeness of the data set.

[0093] In an embodiment of the present application, before determining whether the evaluation result of the value degree of the new data to be evaluated meets the requirements, a feature index is established in the dataset to be evaluated, including: establishing an index value according to the new data to be evaluated by using a space filling curve; determining whether there is a repetition between the index value and the index values in the dataset to be evaluated and meeting a preset condition; if there is no repetition, the evaluation module is used to compare the value degree of the new data to be evaluated with a preset lake entry condition; if the value degree of the new data to be evaluated is greater than the preset lake entry condition, the value of the new data reaches the standard for data lake entry, and the new data to be evaluated is saved to the existing dataset.

[0094] After the new data meets the conditions, it is stored in the data lake. The new data first establishes an index through a space filling curve, and the index value of the new data is checked for duplication through the table data_item. If there is a duplication, the process ends. If the index value does not exist in the table data_item, the evaluation module is used to determine the new data. If it passes, the new data is loaded into the data lake; otherwise, the process directly ends.

[0095] Specifically, the evaluation module needs to first calculate the value degree of the new data, including the following steps:

[0096] S1,

[0097] S2, To Are the first K unitized feature vectors with the closest distances in the table data_item, and the value of K can be customized by the system. And To To Are respectively To Representations on the N-1 dimensional plane. d is the 2-norm distance between the vector And the centroid vector .

[0098] S3, the specific calculation method is as follows: Where Is The coordinates transformed to the N-1 dimensional hyperplane, where Is The coordinates transformed to the N-1 dimensional hyperplane. The value degree v belongs to real numbers from 0 to 100. The larger v is, the greater the value of the new data.

[0099] S4, in the formula v = 100*(1 - e -k·d ) the k is also a constant greater than 0, which can be defined by the user.

[0100] After obtaining the value degree v of the new data, the evaluation module compares it with vT Compare as follows:

[0101] If v > v T , it means that the value of the newly added data has reached the standard for data to enter the lake; otherwise, the value of the newly added data is not high and it does not meet the conditions for entering the lake. Among them, v T is the condition for entering the lake defined by the system.

[0102] Through the above method, the index values on the space filling curve are used to discover the existing data samples to reduce data duplication, and the characteristics of the data samples still missing in the data lake are also identified.

[0103] In an embodiment of the present application, if the query result is not repeated and meets the preset conditions, then save the newly added data to be evaluated, and update the evaluation result corresponding to the data set to be evaluated, including: using the evaluation result corresponding to the data set to be evaluated as the evaluation result of the completeness of the data set; using the value evaluation result of the newly added data to be evaluated as the value evaluation result of the newly added data set; if the completeness of the data set is insufficient, then add the data set according to the value evaluation result of the newly added data entry; if the value of the newly added data to be evaluated is insufficient, then re-select the newly added data set.

[0104] If you want to know whether the data value of the multi-modal data set is rich enough, you can use the completeness of the data set as an evaluation index. If you want to know whether the value of the newly added data to the existing data set is high and whether it is necessary to add it to the data set, you can use the value evaluation of the newly added data set as an evaluation index. By constructing a high-quality multi-modal data set, it has high practical benefits.

[0105] In an embodiment of the present application, the method further includes: if the coverage rate of the data set reaches A%, then re-establish the index value of the data entry; according to the index value of the data entry, increase the total number of indexes to obtain a new total number of indexes; according to the new total number of indexes, re-calculate the coverage rate of the data set, and A is a constant.

[0106] When the data set coverage rate is close to A (100%), perform an operation to change the resolution, so as to optimize the total number of indexes.

[0107] Furthermore, if the coverage rate of the data set is close to 100%, the system should appropriately reduce the length of the quadrant sequence according to the situation, so as to traverse the table data_item and re-establish the index value of each data entry based on the unitized feature vector of its feat column, and then update it to the idx column of the entry.

[0108] Finally, after calculating the total number of new indexes M, recalculate the coverage rate of the data set. By gradually approaching 100% of the data set coverage rate and then changing the quadrant sequence length to re - establish the index values of data entries, and through the iterative process of increasing the total number of indexes M to reduce the coverage rate, it can effectively prevent duplicate invalid data from being added to the data set and ensure the quality of the data set is gradually improved.

[0109] The embodiment of the present application also provides a data set evaluation device 500, as Figure 5 shown, which provides a schematic structural diagram of the data set evaluation device in the embodiment of the present application. The data set evaluation device 500 at least includes: an index module 510, an evaluation module 520, a judgment module 530, and an update module 540, where:

[0110] In an embodiment of the present application, the first index module 510 is specifically configured to: input the data set to be evaluated into the feature index module to obtain index values.

[0111] The data set to be evaluated includes, but is not limited to, data sets for large model training, data sets for AIGC training, etc. The number of data in the data set to be evaluated is basically determined.

[0112] The "feature index module" has the function of vectorizing data, and after establishing indexes, it stores them in the database. The "data set to be evaluated" is a multi - modal data set, thus meeting the requirements of training samples.

[0113] Specifically, implementing a feature index system for a multi - modal data set is mainly carried out by calculating feature vectors and storing them. Therefore, the vectorization component in the feature index system realizes the calculation of feature vectors, and the database component realizes the storage of feature vectors and index values (obtained by space - filling curves). As shown in Figure 3(a), multi - modal data includes, but is not limited to, text data, image data, audio data, and video data. For image data, audio data, and video data, after image understanding, speech understanding, and video understanding respectively, the corresponding multiple texts are obtained. The multiple texts and the text data are input into a natural language understanding model together to extract N - dimensional feature vectors, establish index values, and save them to the database. For data in modalities such as images, audio, and video, corresponding modality - specific deep learning models can be used first to understand and summarize them as texts and then input them into a large model to obtain feature vectors. If a multi - modal large model is used, data in modalities such as text, image, audio, and video can be directly used as the input of the multi - modal large model, and then the multi - modal large model outputs the feature vectors of the data.

[0114] As shown in FIG. 3(b), the multimodal data includes but is not limited to text data, image data, audio data, and video data. A unified multimodal large model can be used to extract N-dimensional feature vectors, establish index values, and save them to the database. For the dataset, the input data is multimodal and supports data in forms such as text, image, audio, and video. For data of different modalities, corresponding deep learning models can be used for understanding. For example, for the text model, the large model can be directly used for understanding to output N-dimensional feature vectors. Specifically, the output of a certain layer in the large model neural network can be used as the feature vector.

[0115] Based on the above, the deep learning method is used to extract key vector features with semantics and use them as the input of the space filling curve.

[0116] In an embodiment of the present application, the evaluation module 520 is specifically configured to: obtain the evaluation result corresponding to the dataset to be evaluated according to the index value.

[0117] Based on the index value, a completeness evaluation result is obtained, and the evaluation result of the dataset to be evaluated is obtained.

[0118] In an embodiment of the present application, the judgment module 530 is specifically configured to: before judging whether the value evaluation result of the new data to be evaluated meets the requirements, establish a feature index in the dataset to be evaluated.

[0119] The new data to be evaluated includes but is not limited to the data that needs to be added to the dataset. Similarly, the new data also needs to be evaluated.

[0120] The process of loading new data into the data lake of the current dataset is as follows: The new data first establishes an index through the space filling curve, and the index value of the new data is judged for duplication through a preset index table. If it is repeated, the new data needs to be added. If it is not repeated, the new data also needs to be added.

[0121] In an embodiment of the present application, the update module 540 is specifically configured to: if the query result is not repeated and meets the set conditions, save the new data to be evaluated and update the evaluation result corresponding to the dataset to be evaluated. For the dataset to be evaluated, the data completeness index in the existing dataset can be determined in a quantitative manner.

[0122] For the new data, an evaluation method for evaluating the value of the new data to this dataset is provided, so as to avoid data redundancy and reduce the waste of storage space.

[0123] It can be understood that the above dataset evaluation device can implement each step of the dataset evaluation method provided in the foregoing embodiments. The relevant explanations regarding the dataset evaluation method are applicable to the dataset evaluation device and will not be elaborated here.

[0124] Figure 6 This is a schematic structural diagram of an electronic device according to an embodiment of the present application. Please refer to Figure 6 , at the hardware level, the electronic device includes a processor, and optionally also includes an internal bus, a network interface, and a memory. Among them, the memory may include a memory, such as a high-speed random access memory (Random-Access Memory, RAM), and may also include a non-volatile memory, such as at least one disk memory, etc. Of course, the electronic device may also include other hardware required for other services.

[0125] The processor, network interface, and memory can be interconnected through an internal bus. The internal bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of simplicity of representation, Figure 6 only a bidirectional arrow is used in

[0126] The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provide instructions and data to the processor.

[0127] The processor reads the corresponding computer program from the non-volatile memory into the memory and then runs it, forming a dataset evaluation device at the logical level. The processor executes the program stored in the memory and is specifically used to perform the following operations:

[0128] Input the dataset to be evaluated into the feature index module to obtain an index value;

[0129] According to the index value, obtain the evaluation result corresponding to the dataset to be evaluated;

[0130] Before judging whether the evaluation result of the value degree of the newly added data to be evaluated meets the requirements, establish a feature index in the dataset to be evaluated;

[0131] If the query results are not repeated and meet the set conditions, save the new data to be evaluated and update the evaluation result corresponding to the data set to be evaluated.

[0132] As described in this application Figure 2 The method executed by the data set evaluation device disclosed in the foregoing embodiments of the present application can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software. The above processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it may also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method.

[0133] The electronic device can also execute Figure 2 the method executed by the data set evaluation device in Figure 2 the embodiments shown, and implement the functions of the data set evaluation device in

[0134] The embodiments of the present application also propose a computer-readable storage medium. The computer-readable storage medium stores one or more programs. The one or more programs include instructions that, when executed by an electronic device including a plurality of application programs, can enable the electronic device to execute Figure 2 the method executed by the data set evaluation device in the embodiments shown, and specifically used to execute:

[0135] Input the data set to be evaluated into the feature index module to obtain an index value;

[0136] According to the index value, obtain the evaluation result corresponding to the dataset to be evaluated;

[0137] Before determining whether the evaluation result of the value degree of the newly added data to be evaluated meets the requirements, establish a feature index in the dataset to be evaluated;

[0138] If the query result is not repeated and meets the set conditions, save the newly added data to be evaluated and update the evaluation result corresponding to the dataset to be evaluated.

[0139] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0140] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, and the combination of processes and / or blocks in the flowchart and / or block diagram can also be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0141] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0142] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the functions specified in one Figure 1 one process or multiple processes and / or blocks Figure 1Steps of the functions specified in one or more boxes.

[0143] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0144] Memory may include non-permanent memory in computer-readable media, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.

[0145] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can store information by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0146] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but also other elements not expressly listed, or elements that are inherent to such process, method, article, or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element.

[0147] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, system, or computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0148] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.

Claims

1. A dataset evaluation method, wherein: The evaluation methods include: Input the dataset to be evaluated into the feature index module to obtain the index value; According to the index value, obtaining an evaluation result corresponding to the data set to be evaluated; Before determining whether the value evaluation result of the newly added data to be evaluated meets the requirements, establishing a feature index in the data set to be evaluated; If the query results are not repeated and meet the set conditions, the newly added data to be evaluated is saved, and the evaluation result corresponding to the data set to be evaluated is updated.

2. The method of claim 1, wherein: The step of inputting the dataset to be evaluated into the feature index module to obtain the index value includes: Inputting the data set to be evaluated into the feature index module for vectorization processing to obtain a multi-dimensional feature vector for each data item in the data set to be evaluated; The index value is established based on the spatial plane to which the multi-dimensional feature vector is mapped and by using a space filling curve.

3. The method of claim 2, wherein: The feature vector includes an N-dimensional feature vector, and the index value is established based on the spatial plane mapped to the multi-dimensional feature vector and using a space filling curve, including: Mapping the N-dimensional feature vector to a spatial plane to complete the operation of expanding the N-dimensional sphere into an N-1-dimensional plane; A space filling curve is constructed according to the N-1 dimensional plane, the index value is established and added to the table.

4. The method of claim 1, wherein: Obtaining the evaluation result corresponding to the to-be-evaluated data set according to the index value includes: Calculating the total number of indexes according to the relationship between the index value and the maximum number of spatial indexes of the space filling curve; A data coverage parameter of the data set to be evaluated is determined according to the total number of indexes.

5. The method of claim 1, wherein: Before determining whether the value evaluation result of the newly added data to be evaluated meets the requirements, establishing a feature index in the data set to be evaluated includes: Establishing an index value using a space filling curve according to the newly added data to be evaluated; Determine whether the index value is repeated with the index value in the data set to be evaluated; If there is no duplication, the evaluation module is used to calculate the value of the newly added data to be evaluated and compare it with the preset lake entry conditions; If the value of the newly added data to be evaluated is greater than the preset entry condition, the value of the newly added data meets the standard for entering the data lake, and the newly added data to be evaluated is saved in the existing data set.

6. The method of claim 1, wherein: If the query results are not repeated and meet the set conditions, the newly added data to be evaluated is saved, and the evaluation result corresponding to the data set to be evaluated is updated, including: Taking the evaluation result corresponding to the data set to be evaluated as the completeness evaluation result of the data set; The value evaluation result of the newly added data to be evaluated is used as the value evaluation result of the newly added data set; If the completeness of the data set is insufficient, adding to the data set according to the value assessment result of the newly added data item; If the value of the newly added data to be evaluated is insufficient, a new data set is selected.

7. The method according to claim 4, further comprising: If the coverage rate of the data set reaches A%, the data entry index value is re-established; According to the data entry index value, increasing the total number of indexes to obtain a new total number of indexes; The coverage of the data set is recalculated according to the total number of new indexes, and A is a constant.

8. A data set evaluation device, wherein: The evaluation device comprises: An index module is used to input the dataset to be evaluated into a feature index module to obtain an index value; An evaluation module, used to obtain an evaluation result corresponding to the to-be-evaluated data set according to the index value; A judgment module, used to establish a feature index in the data set to be evaluated before judging whether the value evaluation result of the newly added data to be evaluated meets the requirements; The updating module is used to save the newly added data to be evaluated and update the evaluation result corresponding to the data set to be evaluated if the query result is not repeated and meets the set conditions.

9. An electronic device, comprising: processor; as well as A memory arranged to store computer executable instructions, which when executed cause the processor to perform the method of any one of claims 1 to 7.

10. A computer-readable storage medium storing one or more programs, which, when executed by an electronic device including a plurality of application programs, causes the electronic device to execute any one of the methods of claims 1 to 7.