A data screening method, device and equipment
By performing feature extraction and Sharpley value calculation in the medical field, and combining prior data for data quality evaluation, the problems of low data screening efficiency and waste of resources in the prior art are solved, and fast and effective high-quality data screening is achieved.
Patent Information
- Application Number
- CN202411825581.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-12
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2044-12-12
AI Technical Summary
The prior art is difficult to quickly and efficiently screen data in the medical field, especially when processing low-quality data with noise, and existing evaluation methods consume a lot of resources and time.
By extracting the data to be evaluated, calculating representative parameters and feature weights, calculating Sharpley values with prior data, and then data quality evaluation and screening are carried out to obtain high-quality data.
This method greatly saves resources, reduces data processing load and time, improves processing efficiency, and realizes fast and effective data screening.
Smart Images

Figure CN119271860B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer technology, and particularly relates to a data screening method, apparatus and device. Background Art
[0002] With the progress of machine learning, the quality of data has received increasing attention from researchers. High-quality data will save a large amount of resources for model learning. On the contrary, using low-quality data for model training will lead to performance degradation. Therefore, the process of discarding low-quality data while selecting high-quality data is a key step for the model to learn effectively and efficiently. The main methods for selecting high-quality samples and excluding low-quality samples from a dataset usually include two steps. First, the researcher evaluates the samples. Second, the researcher ranks the samples according to their value and screens out the high-value samples. Samples with incorrect labels or noise are usually regarded as low-quality data. However, there are various types of noise, and these methods cannot distinguish them all, especially in medical scenarios. When filling out a case report, a clinician may inadvertently record incorrect information. During the process of processing a specimen, a pathologist may inadvertently over-stain a microscopic section. A medical device may produce blurred images during operation by a technician, especially when violating established regulations or protocols. The above-mentioned data is considered to be low-quality data with noise. In addition, due to the inherent complexity of the field, existing evaluation methods are rarely used in the medical field.
[0003] One of the main methods for evaluating a sample is to calculate the Shapley value of the sample. This is the expected value of the utility value traditionally generated by the utility function. The core essence of the utility function lies in quantifying the difference between the output of a model after learning a specific sample and its output before learning the sample. Many existing evaluation methods, such as TMCShapley and GTB Shapley, estimate the expected value of specific data by resampling the utility values of the data. These methods require repeated training of the model. However, these models are relatively robust. Evaluation methods based on these repeatedly trained models may ignore the details of the noise. Therefore, these methods are not general when evaluating data, especially in the medical field. Through experiments, it has been verified that the model is more sensitive to changes in the relative positions of pixels than to changes in pixel colors. It can be inferred from this that the reason for this phenomenon is that the model corrects the wrong information using relevant information. These evaluation methods perform relatively poorly when used for abnormal sample detection tasks with special noise. Unfortunately, there is a lot of special noise in medicine. The model learns from their relevant information and corrects it. In various pathophysiological processes such as inflammatory responses or tumor microenvironments, C-reactive protein is related to white blood cells. In the context of microscopic sections, if they are overstained, the shape of the cells within the section usually remains unchanged. Compared with the features extracted from normally stained microscopic sections, the features extracted from overstained microscopic sections through the model usually show very little difference. In addition, the evaluation methods of repeatedly training the model consume a large amount of resources and time. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a data screening method, device and equipment that can perform fast and effective screening operations on data with a small processing load to obtain data of different qualities.
[0005] The content of the present invention includes providing a data screening method, comprising:
[0006] Extract features from each sample data in the data to be evaluated. The data to be evaluated includes at least image data and / or text data, and the data to be evaluated is used for analysis, reference, and predictive evaluation of a specified event. The extracted features are related to the specified event;
[0007] Calculate and determine the representative parameter of the sample data based on the extracted features. The representative parameter is related to the confidence level of the features;
[0008] Calculate and determine the weight of the features based on prior data;
[0009] Calculate the Shapley value of the sample data based on the representative parameter, weight, and the correlation relationship between the features in the sample data;
[0010] Determine the quality assessment result of the sample data based on the Shapley value of each of the sample data;
[0011] Screen the data to be evaluated based on the quality assessment results of each sample data to obtain high-quality data.
[0012] In some embodiments, the extracting features from each sample data in the data to be evaluated includes:
[0013] Perform coding indexing on each sample data in the data to be evaluated;
[0014] Extract features from the sample data under each index.
[0015] In some embodiments, the method further includes:
[0016] Analyze and determine the correlation relationship between each of the features;
[0017] Determine target features that have an impact on each other based on the correlation relationship.
[0018] In some embodiments, the calculating and determining the representative parameter of the sample data based on the extracted features includes:
[0019] Compare the features in each sample with the features in other samples, and calculate the representative parameter of each sample data in combination with the following formula :
[0020]
[0021] y i represents the label of the sample data with index i, y j represents the label of the sample data with index j, v j (i) represents the degree to which the i-th sample data verifies its representativeness for the j-th sample data, represents the abnormal probability of the sample data, is the indicator function symbol.
[0022] In some embodiments, the degree of representativeness indicates that the credibility level of the j-th sample data is determined by the i-th sample data according to the distance between the two, and its formula is expressed as:
[0023]
[0024] β j is a hyperparameter, representing the balance point between normal sample data and abnormal sample data, m represents the number of features of the sample data, d i,krepresents the k-th feature of the i-th sample data, d j,k represents the k-th feature of the j-th sample data.
[0025] In some embodiments, the method further includes:
[0026] Normalize the representative parameters of all the sample data to obtain the representative parameter p corresponding to the sample data j :
[0027]
[0028] The represents the confidence level of the sample data, represents the label of the sample data corresponding to , and ⊙ represents the Hadamard matrix operator.
[0029] In some embodiments, calculating the Shapley value of the sample data based on the representative parameter, weight, and the correlation relationship between features in the sample data includes:
[0030] Based on the representative parameter, weight, and the correlation relationship between features in the sample data, calculate the Shapley value of the sample in combination with the following formula :
[0031]
[0032] The is a constant, obtained by summing the weights of all features, q is a hyperparameter representing the degree of correlation between features, and T s is the feature with a correlation relationship.
[0033] In some embodiments, screening the data to be evaluated based on the quality assessment results of each sample data to obtain high-quality data includes:
[0034] Sort each sample data based on the quality assessment results of each sample data;
[0035] Based on the target use of the sample data, screen all the sample data in the sequence to determine high-quality data, available data, and low-quality data;
[0036] The method further includes:
[0037] When the target use is model training, extract some of the available data and fuse it with the high-quality data to obtain diverse fused data;
[0038] Use the high-quality data and the fused data to train the model.
[0039] Another embodiment of the present invention also provides a data screening device, including:
[0040] A feature extraction module, configured to extract features from each sample data in the data to be evaluated, where the data to be evaluated includes at least image data and / or text data, and the data to be evaluated is used for analysis, reference, and prediction evaluation of a specified event, and the extracted features are related to the specified event;
[0041] A first calculation module, configured to calculate and determine a representative parameter of the sample data according to the extracted features, where the representative parameter is related to the confidence level of the features;
[0042] A second calculation module, configured to calculate and determine the weight of the features according to prior data;
[0043] A third calculation module, configured to calculate the Shapley value of the sample data according to the representative parameter, the weight, and the correlation relationship between the features in the sample data;
[0044] A determination module, configured to determine the quality evaluation result of the sample data according to the Shapley value of each sample data;
[0045] A screening module, configured to screen the data to be evaluated according to the quality evaluation result of each sample data to obtain high-quality data.
[0046] Another embodiment of the present invention also provides an electronic device, including
[0047] One or more processors;
[0048] A memory, configured to store one or more programs;
[0049] When the one or more programs are executed by the one or more processors, the one or more processors implement the data screening method described in any one of the above embodiments.
[0050] The beneficial effect of the present invention is that using prior data to assist in solving the Shapley value of sample data can greatly save resources, reduce the data processing load and processing time. Moreover, by introducing prior data, complex utility value calculations and repeated sampling of sample data are not required when solving the Shapley value, further reducing the processing load and improving the processing efficiency, enabling a more rapid and effective screening process when screening data.
[0051] Other features and advantages of the present application will be described in the subsequent specification, and in part will be obvious from the specification, or will be understood by implementing the present application. The objectives and other advantages of the present application can be achieved and obtained by the structures specifically pointed out in the written specification, claims, and drawings.
[0052] The technical solutions of the present application will be further described in detail below through the drawings and embodiments. Description of the Drawings
[0053] In order to more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0054] Figure 1 It is a flowchart of the data screening method in an embodiment of the present invention.
[0055] Figure 2 It is a flowchart of the data screening method in another embodiment of the present invention.
[0056] Figure 3 It is an application flowchart of the data screening method in an embodiment of the present invention.
[0057] Figure 4 It is a structural block diagram of the data screening device in an embodiment of the present invention. Detailed Embodiments
[0058] Next, the specific embodiments of the present invention will be described in detail with reference to the drawings, but it is not a limitation of the present invention.
[0059] It should be understood that various modifications can be made to the embodiments disclosed herein. Therefore, the following specification should not be regarded as restrictive, but only as an example of the embodiments. Those skilled in the art will think of other modifications within the scope of the present disclosure.
[0060] The drawings included in the specification and constituting a part of the specification show the embodiments of the present disclosure, and together with the general description of the present disclosure given above and the detailed description of the embodiments given below are used to explain the principles of the present disclosure.
[0061] These and other features of the present invention will become apparent from the following description of the preferred forms of the embodiments given by way of non-limiting examples with reference to the drawings.
[0062] It should also be understood that although the present invention has been described with reference to some specific examples, those skilled in the art can surely implement many other equivalent forms of the present invention, which have the features as described in the claims and thus are all within the protection scope defined hereby.
[0063] When combined with the accompanying drawings, the above and other aspects, features and advantages of the present disclosure will become more apparent in view of the following detailed description.
[0064] Specific embodiments of the present disclosure will be described hereinafter with reference to the accompanying drawings; however, it should be understood that the disclosed embodiments are merely examples of the present disclosure, which can be implemented in various ways. Well-known and / or repetitive functions and structures are not described in detail to avoid obscuring the present disclosure with unnecessary or redundant details. Therefore, the specific structural and functional details disclosed herein are not intended to be limiting, but merely serve as a basis and representative basis for the claims to teach those skilled in the art to use the present disclosure in substantially any suitable detailed structure in a variety of ways.
[0065] This specification may use the phrases "in one embodiment", "in another embodiment", "in yet another embodiment" or "in other embodiments", which may each refer to one or more of the same or different embodiments according to the present disclosure.
[0066] Next, embodiments of the present invention will be described in detail with reference to the accompanying drawings.
[0067] One of the main methods for evaluating a sample is to calculate the Shapley value of the sample. This is the expected value of the utility value traditionally generated by the utility function. The core essence of the utility function lies in quantifying the difference between the output of a model after learning a specific sample and its output before learning that sample. Many existing evaluation methods, such as TMCShapley and GTB Shapley, estimate the expected value of specific data by resampling the utility values of the data. However, these methods all require repeated training of the model. And when evaluating based on these repeatedly trained models, the details of the noise are often likely to be ignored. Therefore, these methods are not general when evaluating data, especially in the medical field. This kind of model is more sensitive to changes in the relative positions between pixels than to changes in pixel colors. The reason for this phenomenon is that the model corrects the wrong information using relevant information. Therefore, when these evaluation methods are used for the abnormal sample detection task with special noise, the performance is relatively poor. However, various special noises generally exist in medical data in the medical field, so the application effect of the model is not good. For example, in various pathophysiological processes such as inflammatory responses or tumor microenvironments, C-reactive protein is related to white blood cells. In the context of microscopic sections, if they are overstained, the shape of the cells within the section usually remains unchanged. Compared with the features extracted from normally stained microscopic sections, the features extracted by the model from overstained microscopic sections usually show very small differences. In addition, the evaluation methods of repeatedly training the model consume a large amount of resources and time. All these phenomena indicate that the models in the existing solutions cannot be applied to data with special noise, such as medical data.
[0068] To solve the above problems, as Figure 1 shown, an embodiment of the present invention provides a data screening method, including:
[0069] S1: Extract features from each sample data in the data to be evaluated. The data to be evaluated includes at least image data and / or text data, and the data to be evaluated is used for the analysis, reference, and prediction evaluation of a specified event. The extracted features are related to the specified event;
[0070] S2: Calculate and determine the representative parameter of the sample data based on the extracted features. The representative parameter is related to the confidence level of the features;
[0071] S3: Calculate and determine the weight of the features based on prior data;
[0072] S4: Calculate the Shapley value of the sample data based on the representative parameter, weight, and the correlation relationship between the features in the sample data;
[0073] S5: Determine the quality evaluation result of the sample data based on the Shapley values of each sample data;
[0074] S6: Screen the data to be evaluated based on the quality assessment results of each sample data to obtain high-quality data.
[0075] The specific content and field involved in the specified event are not limited. The above method can process the relevant data of any event in any field to obtain more reference-worthy high-quality data for that event. In this embodiment, the data to be evaluated is medical data, but it can also be data in other fields, which is not fixed specifically. The medical data can be, for example, but not limited to, medical image data, diagnosis and treatment data, etc. The solution in this embodiment is to first perform a quality assessment on the data to be evaluated, and then screen the data according to the assessment results. The screened data can at least eliminate low-quality data and identify high-quality data. When performing a quality assessment on the data, the solution selected in this embodiment is also to calculate the Shapley value. However, in this embodiment, by feature extraction, calculation of representative parameters, and introduction of prior data to determine the weights of the features of each sample data, and finally calculate the Shapley value based on the association relationship between the representative parameters, weights, and features. By using prior data to assist in solving the Shapley value of sample data, this embodiment can greatly save resources, reduce the data processing load and processing time. Moreover, by introducing prior data, it is not necessary to perform complex utility value calculations and repeated sampling on the sample data when solving the Shapley value, further reducing the processing load and improving the processing efficiency, so that when screening the data, the screening process can be completed more quickly and effectively.
[0076] Furthermore, based on the above content, it can be seen that in the process of solving the Shapley value, the solution in this embodiment actually redefines the utility function of the sample. Its utility value is generated by prior data, thus completely avoiding repeated training of the model. Also, since the definition of the Shapley value requires repeated sampling of the utility value, this embodiment combines the synergy function to propose a new method for calculating the Shapley value through the utility function, which is more efficient than previous methods. In addition, the method in this embodiment also does not need to manually create a validation set to calculate the utility function, so it can significantly improve the data processing efficiency and ensure data validity at the same time.
[0077] In one embodiment, as Figure 2 shown, in order to ensure accurate feature management, this embodiment will preprocess the sample data, such as performing feature extraction on each sample data in the data to be evaluated, including:
[0078] S7: Perform coding indexing on each sample data in the data to be evaluated;
[0079] S8: Perform feature extraction on the sample data under each index.
[0080] That is, first encode and index the sample data, and then each sample data can be retrieved through the index and its features can be extracted.
[0081] After the features are extracted, the system will perform a quality assessment on them. In this case, a utility function is needed to calculate the Shapley value using the utility function. Regarding the utility function, it varies for different scenarios. For example:
[0082] Taking the scenario of interaction between features as an example, in supervised machine learning, data usually consists of multiple features and a label. The importance of these features varies, and the same feature may have different values for different labels. Using to represent the features of the sample data, where the element d i represents the i-th feature of the data. When the value of the sample data is equal to the sum of the values of its constituent elements, the following hypothesis of the utility function can be obtained: . z i represents the weight corresponding to the i-th feature of d i . p represents the confidence of data j, which can be interpreted in real-world scenarios. Data from different institutions should have different confidences. For example, data collected by authoritative institutions is more credible, and intuitively, such data should have a higher confidence p. For any feature d in data i and label y, where S is the set of features, the synergy function ω(S) can be inferred as: .
[0083] When evaluating data based on instances, compared with data from inexperienced institutions, data from authoritative institutions has more weight. In addition, rare cases should receive more attention. Therefore, the confidence level p should vary among different instances within a group. Using to represent the set of instances, a group has n instances, and the instances are independent of the features. The utility function of the group can be defined as: , represents a constant value, m represents the number of features of the sample data, For any instance g in the set of instances j , the synergy function can be inferred as: .
[0084] Taking the scenario of interaction between features as an example, assume that the influence of one feature on another feature in a data is the same. That is to say, in addition to the utility value, the set of features will have a further impact on the instance during the evaluation process. At the same time, their influence degrees are equal, and its utility function can be defined as:
[0085] 。
[0086] When evaluating interaction data based on instances, assuming that the confidence is independent of features, then
[0087]
[0088] The second term of the above formula has a synergistic effect.
[0089] When evaluating interaction data under limited conditions based on features, that is, some features in the sample data have mutual influence, while some features are independent. For this scenario, use to represent the set of features that are expected to influence each other to a certain extent q. There are |T s | features in the feature set S that influence each other. The total feature set D = {d 1 , d 2 ,.., d m} and the influence set T s , label y The relationship is as follows: , and the utility function is 。
[0090] When evaluating interaction data under limited conditions based on instances, , the total feature set and |T s | features influence each other in the set S, where 。The utility function of this group can be defined as: 。
[0091] Based on the above content about the utility function in different scenarios, it can be seen that in this embodiment, the Shapley value calculation process is simplified by using the synergy degree function, improving the calculation efficiency and saving resources. Therefore, the simplified method for solving the Shapley value only requires parameters (sum of feature weights), p j (representative parameter, related to the confidence), and q (degree of association / influence between features). When calculating the Shapley value of a single example, the parameter p j plays a more important role than other parameters. In this embodiment, the parameters and q are given by known prior data.
[0092] Exemplarily, the method further includes:
[0093] S9: Analyze and determine the association relationship between each of the features;
[0094] S10: Based on the association relationship, determine the target features that have influence on each other.
[0095] That is, analyze and determine the feature set with an association relationship, and at the same time calculate and determine the degree of association q between the features in this set.
[0096] Prior data on feature weights and the degree of mutual influence between features is easy to obtain. Regarding the credibility p of the data j , data sets from different institutions show different degrees of confidence. However, it is usually challenging to determine the exact source of the data set. Therefore, this parameter is obtained relatively more quickly through actual solution. The farther a single data is from its data set, the greater the possibility of being abnormal, that is, the greater the possibility of low quality. Therefore, the parameter p j should be negatively correlated with its distance from other data points. Based on this, as Figure 3 shown, when calculating and determining the representative parameter of the feature based on the extracted feature in this embodiment, the proposed algorithm is:
[0097] S11: Compare the features in each sample with the features in other samples, and calculate the representative parameter of each sample data in combination with the following formula :
[0098]
[0099] y i represents the label of the sample data with index i, y j represents the label of the sample data with index j, v j (i) represents the degree to which the i-th sample data corroborates the representativeness of the j-th sample data, represents the abnormal probability of the sample data, is the symbol of the indicator function.
[0100] Among them, the degree of representativeness indicates that the credibility level of the j-th sample data is determined by the i-th sample data according to the distance between the two, and its formula is expressed as:
[0101]
[0102] β j is a hyperparameter, representing the balance point between normal sample data and abnormal sample data, m represents the number of features of the sample data, d i,k represents the k-th feature of the i-th sample data, d j,k represents the k-th feature of the j-th sample data.
[0103] In this embodiment, and β jIt can be obtained through prior knowledge. Given the estimated proportion of noisy data in the dataset ε and the estimated proportion of unlabeled data in the dataset δ, the calculation formulas for the two parameters can be expressed as:
[0104]
[0105] Among them, represents the percentile distance between j and the set of . Mislabeled data can also be regarded as interfering data for the class to which it is misassigned.
[0106] Furthermore, the method further includes:
[0107] S12: Normalize the representative parameters of all the sample data to obtain the representative parameter p corresponding to the sample data j :
[0108]
[0109] The represents the confidence level of the sample data, represents the label of the sample data corresponding to , and ⊙ represents the Hadamard matrix operator.
[0110] Calculating the Shapley value of the sample data based on the representative parameter, weight, and the correlation relationship between features in the sample data includes:
[0111] S13: Based on the representative parameter, weight, and the correlation relationship between features in the sample data, calculate the Shapley value of the sample in combination with the following formula :
[0112]
[0113] The is a constant obtained by summing the weights of all features. q is a hyperparameter representing the degree of correlation between features, and T s is the feature with a correlation relationship.
[0114] After determining the Shapley values of each sample data, the system can then screen the sample data. Screening the data to be evaluated based on the quality assessment results of each sample data to obtain high-quality data includes:
[0115] S14: Sort each sample data based on the quality assessment results of each sample data;
[0116] S15: Screening all sample data in the sequence based on the target use of the sample data to determine high-quality data, usable data, and low-quality data;
[0117] The method further comprises:
[0118] S16: When the target use is model training, extract part of the available data and fuse it with the high-quality data to obtain fused data with diversity;
[0119] S17: Train the model using high-quality data and fused data.
[0120] For example, when preparing training data for a medical model, or when classifying or predicting a disease based on a medical model, the method of the above embodiment can be used to screen the medical data, and then put them into different uses. Assuming that during an imaging examination in a medical scenario, some patients may become restless and uncooperative due to their condition, resulting in low-quality imaging data for the patient. If used for learning models such as deep neural networks, it will seriously affect the convergence speed and performance of the model. Therefore, before model training, or when predicting or classifying a disease, it is necessary to evaluate the collected data, discard low-quality data, and reduce its impact on model training or model application analysis. In addition, in order to improve sample diversity, especially when training a model, high-quality data and available data can be fused, so that even if more low-quality data is discarded, the diversity of the samples will not be affected.
[0121] In addition, since low-quality data is discarded, taking model training as an example, this can improve the training efficiency and quality of the model and improve the model accuracy. For example, there is no need to repeatedly train the model in large quantities. The samples to be evaluated can be directly used as training data and verification sets without manual screening. The overall data samples are more in line with the real data distribution, which is more conducive to model training.
[0122] like Figure 4 As shown, another embodiment of the present invention also provides a data screening device 100, including:
[0123] A feature extraction module is used to extract features from each sample data in the data to be evaluated, wherein the data to be evaluated includes at least image data and / or text data, and the data to be evaluated is used for analysis, reference, prediction and evaluation of a specified event, and the extracted features are related to the specified event;
[0124] A first calculation module, used to calculate and determine a representative parameter of the sample data according to the extracted features, wherein the representative parameter is related to the confidence of the feature;
[0125] A second calculation module, configured to calculate and determine the weight of the feature according to prior data;
[0126] A third calculation module, configured to calculate the Shapley value of the sample data according to the representative parameter, the weight, and the correlation relationship between features in the sample data;
[0127] A determination module, configured to determine the quality evaluation result of the sample data according to the Shapley value of each sample data;
[0128] A screening module, configured to screen the data to be evaluated according to the quality evaluation result of each sample data to obtain high-quality data.
[0129] In some embodiments, the extracting features from each sample data in the data to be evaluated includes:
[0130] Performing coding indexing on each sample data in the data to be evaluated;
[0131] Extracting features from the sample data under each index.
[0132] In some embodiments, the method further includes:
[0133] Analyzing and determining the correlation relationship between each feature;
[0134] Determining target features that have an impact on each other based on the correlation relationship.
[0135] In some embodiments, the calculating and determining the representative parameter of the sample data based on the extracted features includes:
[0136] Comparing the features in each sample with the features in other samples, and calculating the representative parameter of each sample data in combination with the following formula :
[0137]
[0138] y i represents the label of the sample data with index i, y j represents the label of the sample data with index j, v j (i) represents the degree to which the i-th sample data corroborates the representativeness of the j-th sample data, represents the abnormal probability of the sample data, is the indicator function symbol.
[0139] In some embodiments, the degree of representativeness indicates that the credibility level of the j-th sample data is determined by the i-th sample data according to the distance between the two, and its formula is expressed as:
[0140]
[0141] β j is a hyperparameter, representing the balance point between normal sample data and abnormal sample data, m represents the number of features of the sample data, d i,k represents the k-th feature of the i-th sample data, d j,k represents the k-th feature of the j-th sample data.
[0142] In some embodiments, the method further includes:
[0143] Normalize the representative parameters of all the sample data to obtain the representative parameter p corresponding to the sample data j :
[0144]
[0145] The represents the confidence level of the sample data, represents and the label of the corresponding sample data, ⊙ represents the Hadamard matrix operator.
[0146] In some embodiments, calculating the Shapley value of the sample data based on the representative parameter, weight, and the correlation relationship between features in the sample data includes:
[0147] Based on the representative parameter, weight, and the correlation relationship between features in the sample data, calculate the Shapley value of the sample in combination with the following formula :
[0148]
[0149] The is a constant, obtained by summing the weights of all features, q is a hyperparameter, representing the degree of correlation between features, T s is a feature with a correlation relationship.
[0150] In some embodiments, screening the data to be evaluated based on the quality evaluation results of each sample data to obtain high-quality data includes:
[0151] Sort each sample data based on the quality evaluation results of each sample data;
[0152] Based on the target use of the sample data, screen all the sample data in the sequence to determine high-quality data, available data, and low-quality data;
[0153] The device further includes:
[0154] A fusion module, configured to extract some of the available data and fuse it with high-quality data to obtain diverse fused data when the target use is model training;
[0155] A training module, configured to train the model using the high-quality data and the fused data.
[0156] Another embodiment of the present invention further provides an electronic device, including:
[0157] One or more processors;
[0158] A memory, configured to store one or more programs;
[0159] When the one or more programs are executed by the one or more processors, the one or more processors implement the data screening method described in any of the above embodiments.
[0160] Furthermore, an embodiment of the present invention further provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, the data screening method described above is implemented. It should be understood that each of the solutions in this embodiment has the corresponding technical effects in the above method embodiment, and will not be elaborated here.
[0161] Furthermore, an embodiment of the present invention further provides a computer program product, the computer program product is tangibly stored on a computer-readable medium and includes computer-readable instructions, and the computer-executable instructions, when executed, cause at least one processor to execute a network attack defense method such as the above-described embodiment.
[0162] It should be noted that the computer storage medium of the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable medium can, for example, but is not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access storage medium (RAM), a read-only storage medium (ROM), an erasable programmable read-only storage medium (EPROM or flash memory), an optical fiber, a portable compact disk read-only storage medium (CD-ROM), an optical storage medium, a magnetic storage medium, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which the computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable medium can send, propagate, or transmit a program configured to be used by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wireless, antenna, optical cable, RF, etc., or any suitable combination of the above.
[0163] In addition, those skilled in the art should understand that the embodiments of the present invention can be provided as a method, a system, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage and optical storage, etc.) that contain computer-usable program code.
[0164] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the process Figure 1one or more processes and / or blocks Figure 1 a system of functions specified in one or more blocks
[0165] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufacture including an instruction system that implements the functions specified in one Figure 1 one or more processes and / or blocks Figure 1 one or more blocks
[0166] Those of ordinary skill in the art should understand that the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of protection of this application is limited to these examples; under the concept of this application, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of one or more embodiments of this application as described above, which are not provided in detail for the sake of brevity.
Claims
1. A data screening method, characterized in that: include: Extracting features from each sample data in the data to be evaluated, wherein the data to be evaluated includes at least image data and / or text data, and the data to be evaluated is used for analysis, reference, and prediction and evaluation of a specified event, and the extracted features are related to the specified event; Determining representative parameters of the sample data based on the extracted features, wherein the representative parameters are related to the confidence of the features; Determining the weight of the feature based on prior data calculation; Calculate the Shapley value of the sample data based on the representative parameter, the weight and the correlation between the features in the sample data; Determine a quality assessment result of the sample data based on the Shapley value of each of the sample data; Based on the quality assessment results of each of the sample data, the data to be assessed are screened to obtain high-quality data; The step of calculating and determining representative parameters of the sample data based on the extracted features includes: The features in each sample are compared with the features in other samples, and the representative parameters of each sample data are calculated by the following formula: : y i Represents the label of the sample data with index i, y j Represents the label of the sample data with index j, v j (i) indicates the degree to which the i-th sample data confirms its representativeness to the j-th sample data, represents the abnormal probability of sample data, is the indicator function symbol.
2. The data screening method according to claim 1, characterized in that: The feature extraction of each sample data in the data to be evaluated includes: Encoding and indexing each sample data in the data to be evaluated; Perform feature extraction on the sample data under each index.
3. The data screening method according to claim 1, characterized in that: The method further comprises: Analyze and determine the correlation between each of the features; Based on the association relationship, target features that have an impact on each other are determined.
4. The data screening method according to claim 1, characterized in that: The representative degree means that the credibility level of the j-th sample data is determined by the i-th sample data according to the distance between the two, and its formula is expressed as: β j is a hyperparameter, indicating the balance point between normal sample data and abnormal sample data, m indicates the number of features of sample data, and d i,k represents the kth feature of the i-th sample data, d j,k Represents the kth feature of the jth sample data.
5. The data screening method according to claim 1, characterized in that: The method further comprises: The representative parameters of all the sample data are normalized to obtain the representative parameters p corresponding to the sample data. j : Said Indicates the trustworthiness level of the sample data, Indicates n The credibility level of the sample data, Representation and The label of the corresponding sample data, ⊙ represents the Hadamard matrix calculator.
6. The data screening method according to claim 5, characterized in that: The calculating the Shapley value of the sample data based on the representative parameter, the weight and the correlation between the features in the sample data includes: Based on the representative parameters, weights and the correlation between the features in the sample data, the Shapley value of the sample is calculated by combining the following formula: : Said is a constant obtained by summing the weights of all features, q is a hyperparameter indicating the degree of association between features, T s Features with associated relationships.
7. The data screening method according to claim 1, characterized in that: The step of screening the data to be evaluated based on the quality evaluation results of each sample data to obtain high-quality data includes: sorting each sample data based on the quality assessment result of each sample data; Screening all sample data in the sequence based on the target use of the sample data to determine high-quality data, usable data, and low-quality data; The method further comprises: When the target use is model training, extract part of the available data and fuse it with high-quality data to obtain fused data with diversity; The model is trained using high-quality data and fused data.
8. A data screening device, characterized in that: include: A feature extraction module is used to extract features from each sample data in the data to be evaluated, wherein the data to be evaluated includes at least image data and / or text data, and the data to be evaluated is used for analysis, reference, prediction and evaluation of a specified event, and the extracted features are related to the specified event; A first calculation module, used to calculate and determine a representative parameter of the sample data according to the extracted features, wherein the representative parameter is related to the confidence of the feature; A second calculation module, used to calculate and determine the weight of the feature according to prior data; A third calculation module, used for calculating the Shapley value of the sample according to the representative parameter, the weight and the correlation relationship between the features in the sample data; A determination module, configured to determine a quality assessment result of the sample data according to the Shapley value of each sample data; A screening module, used to screen the data to be evaluated according to the quality evaluation results of each sample data to obtain high-quality data; The process of calculating and determining representative parameters of the sample data based on the extracted features includes: The features in each sample are compared with the features in other samples, and the representative parameters of each sample data are calculated by the following formula: : y i Represents the label of the sample data with index i, y j Represents the label of the sample data with index j, v j (i) indicates the degree to which the i-th sample data confirms its representativeness to the j-th sample data, represents the abnormal probability of sample data, is the indicator function symbol.
9. An electronic device, characterized in that: include one or more processors; a memory configured to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the data screening method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Malicious encrypted traffic detection method, system and related device
CN113965390A
Data processing method and related device
CN116541684A