A data screening method, device, computer program product, equipment and medium
By adjusting the weight, error compensation and data distinction on the target data set on the calculation card, and comparing the reference feature set, the accurate screening of the data set is achieved, solving the problem of deterioration in data screening accuracy, and improving the accuracy and flexibility of data screening.
Patent Information
- Application Number
- CN202411733248.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-11-29
AI Technical Summary
As the data scale increases, how to accurately screen data has become an urgent problem. It is difficult for the existing technology to effectively handle the impact of various data, resulting in the deterioration of the accuracy of data screening.
By applying a data filtering method on the calculation card, the target feature set is obtained by obtaining the target data set that is processed transmitted by the operation interface, determining its filtering method, and performing weight adjustment, error compensation and data distinction to obtain the target feature set. Then, compare with the reference feature set, filter the target data set according to the comparison results, obtain the filtering results and transmit them to the operation interface.
Through weight adjustment, error compensation and data distinction, the calculation card can accurately extract the target feature set of the target data set and flexibly filter it, which improves the accuracy and flexibility of data screening and meets the operation interface's processing needs for the data set.
Smart Images

Figure CN119226584B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing technology, and more specifically, to a data screening method, device, computer program product, equipment and medium. Background Art
[0002] As the number of data sources increases, the scale of data also increases. Although it is easier to obtain data, it also brings new problems. For example, various types of data will affect the acquisition of required data, resulting in poor accuracy of data screening.
[0003] In summary, how to accurately screen data is a problem that currently needs to be solved urgently by those skilled in the art. Summary of the invention
[0004] The purpose of the present invention is to provide a data screening method, which can solve the technical problem of how to accurately screen data to a certain extent. The present invention also provides a data screening device, a computer program product, an electronic device and a computer non-volatile readable storage medium.
[0005] In order to achieve the above object, the present invention provides the following technical solutions:
[0006] In a first aspect, a data screening method is provided, which is applied to a computing card, comprising:
[0007] Obtain the target data set to be processed transmitted by the operation interface;
[0008] Determining a screening method for the target data set;
[0009] According to the screening method, weight adjustment, error compensation and data differentiation are performed on the target data set to obtain a target feature set corresponding to the target data set;
[0010] Acquire a reference feature set for screening the target data set;
[0011] Comparing the target feature set with the reference feature set to obtain a comparison result;
[0012] The target data set is screened according to the comparison result to obtain a screening result, and the screening result is transmitted to the operation interface.
[0013] On the other hand, weight adjustment, error compensation and data differentiation are performed on the target data set to obtain a target feature set corresponding to the target data set, including:
[0014] Performing weight operation, linear transformation and nonlinear transformation on the target data set to obtain a target feature set corresponding to the target data set;
[0015] Among them, weight operation is used to adjust the weight of data according to the data weight; linear transformation is used to compensate for data errors; and nonlinear transformation is used to distinguish data.
[0016] On the other hand, the screening method includes a method of screening by importance; performing weight operation, linear transformation and nonlinear transformation on the target data set to obtain a target feature set corresponding to the target data set, including:
[0017] Performing a weight operation on the target data set to obtain a first weight operation result;
[0018] Determining first compensation data corresponding to the target data set;
[0019] Performing a linear transformation on the first compensation data and the first weight operation result to obtain a first linear transformation result;
[0020] A nonlinear transformation is performed on the first linear transformation result to obtain a target feature set corresponding to the target data set; the nonlinear transformation is used to amplify non-negative values in the first linear transformation result and reduce negative values in the first linear transformation result.
[0021] On the other hand, comparing the target feature set with the reference feature set to obtain a comparison result includes:
[0022] For each data in the target data set, determining a target feature of the data in the target feature set, and determining a reference feature of the data in the reference feature set;
[0023] generating a first difference between the reference feature and the target feature, and generating a first error rate of data according to the first difference;
[0024] Detecting whether the first error rate is less than or equal to a first standard value;
[0025] In response to the first error rate being less than or equal to the first standard value, obtaining a comparison result indicating that the data is important data;
[0026] In response to the first error rate being greater than the first standard value, a comparison result indicating that the data is non-significant data is obtained.
[0027] On the other hand, after filtering the target data set according to the comparison result and obtaining the filtering result, the method further includes:
[0028] The important data in the target data set is encrypted, and the non-important data in the target data set is desensitized.
[0029] On the other hand, the screening method includes a method of screening security data; performing weight operation, linear transformation and nonlinear transformation on the target data set to obtain a target feature set corresponding to the target data set, including:
[0030] Performing a weight operation on the target data set to obtain a second weight operation result;
[0031] Determining second compensation data corresponding to the target data set;
[0032] Performing a linear transformation on the second compensation data and the second weight operation result to obtain a second linear transformation result;
[0033] Performing a nonlinear transformation on the second linear transformation result to obtain a first nonlinear transformation result; the nonlinear transformation is used to map the second linear transformation result to a positive number less than or equal to 1;
[0034] A weight operation is performed on the first nonlinear transformation result to obtain a target feature set corresponding to the target data set.
[0035] On the other hand, comparing the target feature set with the reference feature set to obtain a comparison result includes:
[0036] For each group of data in the target data set, determining a target feature group of the data group in the target feature set, and determining a reference feature group of the data group in the reference feature set;
[0037] generating a second difference between the reference feature group and the target feature group, and generating a second error rate of the data group according to the second difference;
[0038] Detecting whether the second error rate is less than or equal to a second standard value;
[0039] In response to the second error rate being less than or equal to the second standard value, obtaining a comparison result indicating the safety of the data group;
[0040] In response to the second error rate being greater than the second standard value, a comparison result characterizing the risk of the data set is obtained.
[0041] On the other hand, after filtering the target data set according to the comparison result and obtaining the filtering result, the method further includes:
[0042] Taking the target feature group of the dangerous data group in the target data set as the data group to be purified;
[0043] Perform weight calculation, linear transformation and nonlinear transformation on the data group to be purified to obtain a purification feature set corresponding to the data group to be purified;
[0044] Comparing the cleansing feature set with the reference feature set, and detecting whether the data groups to be cleansed are all safe according to the comparison result;
[0045] In response to the data groups to be purified being all safe, the data groups to be purified are used as safe data groups;
[0046] In response to the existence of a dangerous data group in the data group to be purified, the purified feature set of the dangerous data group in the data group to be purified is used as a new data group to be purified, and the step of performing weight calculation, linear transformation and nonlinear transformation on the data group to be purified is returned to obtain the purified feature set corresponding to the data group to be purified.
[0047] On the other hand, the screening method includes a method of setting data screening; performing weight operation, linear transformation and nonlinear transformation on the target data set to obtain a target feature set corresponding to the target data set, including:
[0048] Performing a weight operation on the target data set to obtain a third weight operation result;
[0049] Determining third compensation data corresponding to the target data set;
[0050] Performing a linear transformation on the third compensation data and the third weight operation result to obtain a third linear transformation result;
[0051] A nonlinear transformation is performed on the third linear transformation result to obtain a target feature set corresponding to the target data set; the nonlinear transformation is used to map the third linear transformation result to a positive number less than or equal to 1.
[0052] On the other hand, comparing the target feature set with the reference feature set to obtain a comparison result includes:
[0053] Determine the length of the word vector;
[0054] In the target feature set, a feature phrase having a length equal to the length value is read out;
[0055] For each of the feature phrases, detect whether the feature phrase contains the set features in the reference feature set. If so, obtain a comparison result indicating that the corresponding data group of the feature phrase in the target data set has the set data. If not, obtain a comparison result indicating that the corresponding data group of the feature phrase in the target data set does not have the set data.
[0056] On the other hand, after filtering the target data set according to the comparison result and obtaining the filtering result, the method further includes:
[0057] The data groups containing setting data in the target data set are filtered to obtain a processed data set.
[0058] On the other hand, obtaining a reference feature set for screening the target data set includes:
[0059] Get the actual data set that carries the set data;
[0060] Performing a weight operation on the actual data set to obtain a fourth weight operation result;
[0061] Determining fourth compensation data corresponding to the actual data set;
[0062] Performing a linear transformation on the fourth compensation data and the fourth weight operation result to obtain a fourth linear transformation result;
[0063] Performing a nonlinear transformation on the fourth linear transformation result to obtain an actual feature set corresponding to the actual data set; the nonlinear transformation is used to map the fourth linear transformation result to a positive number less than or equal to 1;
[0064] A set feature set is obtained, and features existing in both the actual feature set and the set feature set are combined into a reference feature set.
[0065] In a second aspect, a data screening device is provided, comprising a computing card, a power module and a clock module connected to the computing card, and an operation interface connected to the computing card; wherein the computing card, the power module and the clock module are mounted on a board;
[0066] The operation interface is used to operate the target data set to be processed and the screening results;
[0067] The computing card is used to determine a screening method for the target data set; perform weight adjustment, error compensation and data differentiation on the target data set according to the screening method to obtain a target feature set corresponding to the target data set; obtain a reference feature set for screening the target data set; compare the target feature set with the reference feature set to obtain a comparison result; and screen the target data set according to the comparison result to obtain the screening result.
[0068] According to a third aspect, a computer program product is provided, comprising a computer program / instruction, which, when executed by a processor, implements the steps of any of the above-mentioned data screening methods.
[0069] In a fourth aspect, an electronic device is provided, including:
[0070] Memory for storing computer programs;
[0071] A processor is used to implement the steps of any of the above-mentioned data screening methods when executing the computer program.
[0072] In a fifth aspect, a computer non-volatile readable storage medium is provided, wherein the computer non-volatile readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of any of the above data screening methods are implemented.
[0073] A data screening method provided by the present invention is applied to a computing card, and obtains a target data set to be processed transmitted by an operation interface; determines a screening method for the target data set; performs weight adjustment, error compensation and data differentiation on the target data set according to the screening method to obtain a target feature set corresponding to the target data set; obtains a reference feature set for screening the target data set; compares the target feature set with the reference feature set to obtain a comparison result; screens the target data set according to the comparison result to obtain a screening result, and transmits the screening result to the operation interface. The beneficial effects of the present invention are as follows: since the data after weight adjustment is easier to be accurately distinguished, the data after error compensation is more accurate, and the relationship between the data after data distinction and the screening method is clearer, in this way, after the computing card processes the target data set with the help of weight adjustment, error compensation and data distinction, it can accurately extract the target feature set corresponding to the target data set, and can distinguish the features in the target feature set that are related to and unrelated to screening, so as to facilitate the subsequent application of the target feature set to accurately screen the data; in addition, by comparing this target feature set with the reference feature set, the target data set can be screened accordingly. In other words, only the reference feature set needs to be adjusted according to the screening method to adjust the screening results of the data, which has good flexibility and realizes the use of the computing card to provide accurate and flexible data screening results for the operation interface, thereby facilitating the operation interface to process the data set. The data screening device, computer program product, electronic device and computer non-volatile readable storage medium provided by the present invention also solve the corresponding technical problems. BRIEF DESCRIPTION OF THE DRAWINGS
[0074] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0075] Figure 1 A flow chart of a data screening method provided by an embodiment of the present invention;
[0076] Figure 2 Schematic diagram of data feature extraction for importance screening;
[0077] Figure 3 It is a structural diagram of the linear transformation module;
[0078] Figure 4 Schematic diagram of data feature extraction for security screening;
[0079] Figure 5 A schematic diagram of applying sliding window filtering to screen illegal data;
[0080] Figure 6 A schematic diagram for setting data screening;
[0081] Figure 7 A schematic diagram of a data screening device provided by the present invention;
[0082] Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention;
[0083] Fig. 9 Another structural schematic diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0084] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0085] See also Figure 1 , Figure 1 A flow chart of a data screening method provided by an embodiment of the present invention.
[0086] A data screening method provided by an embodiment of the present invention is applied to a computing card and may include the following steps:
[0087] Step S101: Acquire a target data set to be processed transmitted by an operation interface.
[0088] Step S102: Determine a screening method for the target data set.
[0089] In actual applications, the computing card can first obtain the target data set to be processed transmitted by the operation interface, and determine the filtering method of the target data set. The number and type of data in the target data set can be flexibly determined according to the application scenario. For example, in the image processing process, the data in the target data set can be image data, in the text processing process, the data in the target data set can be text data, in the language processing process, the data in the target data set can be language data, etc. The filtering method of the target data set refers to the method of filtering the target data set, which can be determined according to the application requirements of the data. For example, the filtering method can be level filtering of the target data set or specific data filtering.
[0090] It should be noted that in the process of receiving the target data set, you can first check whether the target data set has changed. If it has not changed, then apply it to ensure that the target data set received is consistent with the sending end. In this process, you can first receive the target data set and the original verification value corresponding to the target data set. This verification value can be a hash value, etc., generate a true verification value of the received target data set, and check whether the true verification value is consistent with the original verification value. If they are inconsistent, the target data set is abandoned. If they are consistent, the target data set can be stored for application.
[0091] Step S103: According to the screening method, weight adjustment, error compensation and data differentiation are performed on the target data set to obtain a target feature set corresponding to the target data set.
[0092] In practical applications, the data itself has characteristics that can be used to filter the data. Therefore, after obtaining the target data set, the target feature set of the target data set can be extracted for filtering the target data set. In this process, considering that each filtering method requires different features, the target data set can be weighted, error compensated, and data differentiated according to the filtering method to obtain the target feature set corresponding to the target data set.
[0093] In a specific application scenario, in the process of adjusting the weight of the target data set, compensating for errors and distinguishing data to obtain the target feature set corresponding to the target data set, the target data set can be subjected to weight operation, linear transformation and nonlinear transformation to obtain the target feature set corresponding to the target data set; wherein, the weight operation is used to adjust the weight of the data according to the data weight; the linear transformation is used to compensate for the errors of the data; and the nonlinear transformation is used to distinguish the data. It should be noted that the weight operation refers to the operation of the data according to the data weight, such as multiplication operation. The data weight can be adjusted according to the application scenario. Assuming that the data is x and the data weight is a, the weight operation result can be ax. Assuming that the data and the data weight are both arrays, the weight operation result is the weighted sum of the data and the data weight, etc.; the linear transformation refers to the error compensation of the data, such as the linear transformation result can be the sum of the data required for the linear transformation and the compensation data; the nonlinear transformation is used to distinguish the data according to the screening method. For example, in the process of level screening of the data, the nonlinear transformation can scale the data to distinguish the data of different levels. In the process of specific data screening of the data, the nonlinear transformation can map the data to distinguish the difference between the data and the reference data. In the process of performing weight operations, linear transformations, and nonlinear transformations on the target data set, weight operations, linear transformations, and nonlinear transformations can be performed on the target data matrix or target data vector corresponding to the target data set, so as to quickly and effectively extract the target feature set with the help of the matrix or vector, and the vector of the data can be a vector obtained after One-hot encoding and other operations. In addition, weight operations, linear transformations, and nonlinear transformations can be implemented through software or hardware, etc., and are not specifically limited here.
[0094] Step S104: obtaining a reference feature set for screening the target data set.
[0095] Step S105: Compare the target feature set and the reference feature set to obtain a comparison result.
[0096] Step S106: Filter the target data set according to the comparison result to obtain a filtering result, and transmit the filtering result to the operation interface.
[0097] In practical applications, considering that data screening needs to refer to screening criteria, and the screening criteria can be quantified into corresponding features after being determined, after obtaining the target feature set of the target data set, the reference feature set for screening the target data set can be obtained, and the type of the reference feature set can be flexibly determined according to the data screening requirements; the target feature set and the reference feature set are compared to obtain a comparison result, for example, an error rate between the target feature set and the reference feature set can be generated, and the comparison result is determined according to the relationship between the error rate and the set value, and the error rate between a single target feature and a single reference feature can be the square value of the difference between the target feature and the reference feature, and the error rate between multiple features, that is, the target feature group and the reference feature group can be the sum of the error rates of all single features in the group, etc.; the target data set is screened according to the comparison result to obtain a screening result, and the screening result is transmitted to the operation interface, so that the operation interface can further process the target data set according to the screening result, so that the processed target data set can meet user needs.
[0098] A data screening method provided by the present invention is applied to a computing card, and obtains a target data set to be processed transmitted by an operation interface; determines a screening method for the target data set; performs weight adjustment, error compensation and data differentiation on the target data set according to the screening method to obtain a target feature set corresponding to the target data set; obtains a reference feature set for screening the target data set; compares the target feature set with the reference feature set to obtain a comparison result; screens the target data set according to the comparison result to obtain a screening result, and transmits the screening result to the operation interface. The beneficial effects of the present invention are as follows: since the data after weight adjustment is easier to be accurately distinguished, the data after error compensation is more accurate, and the relationship between the data after data distinction and the screening method is clearer, in this way, after the computing card processes the target data set with the help of weight adjustment, error compensation and data distinction, the target feature set corresponding to the target data set can be accurately extracted, and the features in the target feature set that are related to and unrelated to screening can be distinguished, so as to facilitate the subsequent application of the target feature set to accurately screen the data; in addition, by comparing this target feature set with the reference feature set, the target data set can be screened accordingly. In other words, only the reference feature set needs to be adjusted according to the screening method to adjust the screening results of the data, which has good flexibility and realizes the use of the computing card to provide accurate and flexible data screening results for the operation interface, thereby facilitating the operation interface to process the data set.
[0099] exist Figure 1On the basis of the illustrated embodiment, taking into account the different contributions of the data in the target data set, that is, the different importance of the data, the target data set can be screened for importance, that is, the target data set includes data to be screened for importance, and the screening method at this time includes a method for screening for importance, a reference feature set, that is, a feature set used to screen the target data for importance, and the reference feature set can be generated based on defined important data, etc.; accordingly, in the process of performing weight operations, linear transformations, and nonlinear transformations on the target data set to obtain a target feature set corresponding to the target data set, a weight operation can be performed on the target data set to obtain a first weight operation result; determine first compensation data corresponding to the target data set; perform a linear transformation on the first compensation data and the first weight operation result to obtain a first linear transformation result; perform a nonlinear transformation on the first linear transformation result to obtain a target feature set corresponding to the target data set, the nonlinear transformation is used to amplify non-negative values in the first linear transformation result, and reduce negative values in the first linear transformation result, and the nonlinear transformation can be a piecewise function, such as .
[0100] To facilitate understanding of this process, assume that there are n sets of data in the target data set, represented by a matrix: ; The weight coefficient is expressed as a matrix , element a in A1 ij (1≤i≤n, 1≤j≤n) corresponds to the weight of the input data, which can be defined by the user; the first compensation data is represented by a matrix as , can also be defined by the user, etc.; after linear transformation, vector , C1 is then input into the nonlinear transformation, and after being calculated by the f1() function, the final output is , so the target feature set can be expressed as Y1=f1(A1*X+P1), such as Figure 2 As shown, , after expansion, it is y 1n =f1(a n1× x1+a n2 ×x2+…+a nn × n +p 1n ). It should be noted that the linear transformation realizes the accumulation and summation of data, and can be composed of n groups of linear transformation modules, each group of linear transformation modules can be composed of n multipliers and 1 accumulator, and its structure can be as follows Figure 3 As shown, that is, c 11 =a 11 ×x1+a 12 ×x2+…+a 1n × n +p 11 , c 12 ~c1n The calculation process is similar.
[0101] From this implementation process, it can be seen that in the process of screening the importance of data, the present invention can perform a weight operation on the target data set to obtain a first weight operation result, determine the first compensation data corresponding to the target data set, and perform a linear transformation on the first compensation data and the first weight operation result to obtain a first linear transformation result; perform a nonlinear transformation on the first linear transformation result to obtain a target feature set corresponding to the target data set, and the nonlinear transformation is used to amplify the non-negative values in the first linear transformation result and reduce the negative values in the first linear transformation result. In this way, the probability that the target feature output as non-negative is important data is relatively large, and the data with negative input will be ignored, that is, the unimportant data will be ignored, so that the important data and unimportant data in the target data set can be screened more accurately, and the accuracy of detecting important data can be improved.
[0102] In a specific application scenario, in the process of comparing the target feature set and the reference feature set to obtain a comparison result, for each data in the target data set, the target feature of the data can be determined in the target feature set, and the reference feature of the data can be determined in the reference feature set; a first difference between the reference feature and the target feature is generated, and a first error rate of the data is generated according to the first difference; it is detected whether the first error rate is less than or equal to a first standard value; in response to the first error rate being less than or equal to the first standard value, a comparison result characterizing that the data belongs to important data is obtained; in response to the first error rate being greater than the first standard value, a comparison result characterizing that the data belongs to non-important data is obtained, that is, the first standard value is used to screen whether the data belongs to important data, and its value can be flexibly determined according to actual needs, for example, it can be set to a value between (0,1], and the smaller the value, the higher the recognition accuracy of important data.
[0103] To facilitate understanding of this process, assume that the user defines a reference feature set of n important data as E 1i (1≤i≤n), calculate the characteristics of each column of data in the target data set in turn, that is, calculate the data Features , (1≤j≤m), then the first error rate of the data can be expressed as , (1≤i≤n, 1≤j≤m).
[0104] In specific application scenarios, after the target data set is screened according to the comparison results and the screening results are obtained, the important data in the target data set can be encrypted, such as using AES (Advanced Encryption Standard), 3DES (Triple Data Encryption Algorithm), SM4 and other symmetric encryption algorithms to encrypt the important data; and the non-important data in the target data set can be desensitized, such as replacing characters in non-important text data with "*" or spaces, etc., to avoid the risk of data diffusion and leakage in the target data set.
[0105] In a specific application scenario, in the process of desensitizing the non-important data in the target data set, the non-important data in the target data set can be first used as the data to be desensitized, and the characters in the data to be desensitized can be randomly selected to obtain the first selected characters; the first selected characters in the data to be desensitized are replaced with characters * to obtain the first desensitized data; the characters in the first desensitized data are randomly selected to obtain the second selected characters; the second selected characters in the first desensitized data are replaced with spaces to obtain the target desensitized data corresponding to the data to be desensitized. In this way, the first desensitization of the desensitized data is first achieved by randomly selecting characters in the data to be desensitized and replacing them with characters *, and then the second desensitization of the desensitized characters is achieved by randomly selecting characters in the first desensitized data and replacing them with spaces. Since the characters are randomly selected and not deterministic, the first desensitization and the second desensitization can make any characters in the data to be desensitized desensitized, which improves the randomness and uncertainty of the desensitization process, making it difficult to restore the desensitization results to obtain the original data, and ensuring data security.
[0106] To facilitate understanding of this process, assume that the target data set is composed of passwords, mailboxes, IP (Internet Protocol) addresses, protocols, etc. Important data is represented by level 1, and non-important data is represented by level 2. The reference data can be shown in Table 1; the reference feature set can be represented by a matrix as , the first column f i1 (1≤i≤N) indicates the type of data, such as password, email address, port number, etc. The second column f i2 Indicates the characteristic description of the data, the third column f i3 Indicates the importance level of the data; then according to the present embodiment, each data D in the target data set Data is read in turn. ij (1≤i≤n, 1≤j≤m) to traverse the feature table and convert D ijPerform feature matching with N records in the feature table one by one: Encrypt the data that matches successfully and is at level "1" to generate data ciphertext, for example, the data "368*#Za" identified as a password is encrypted to form the ciphertext "13FA454B", and desensitize the data that matches successfully and is at level "2", for example, the data "192.168.1.15" identified as an IP address is desensitized to form the data "192.*.*.*", and the unmatched data is not processed, thereby encrypting the important data in the target data set and desensitizing the non-important data, ensuring data security. Assume that the target data set is 5 server log records, each of which includes 5 fields, namely source IP address, source port number, destination IP address, destination port number, and protocol, as shown in Table 2. After structural conversion, a basic data set Data with 5 rows and 5 columns is formed:
[0107] ;
[0108] The data set obtained after processing according to the scheme of this embodiment is:
[0109] .
[0110] Table 1 Reference data example table
[0111]
[0112] Table 2 Target dataset example table
[0113]
[0114] exist Figure 1On the basis of the illustrated embodiment, considering that data is susceptible to poisoning, for example, poisoned data, malicious data, etc. are inserted into the data, which may affect the safe use of the data, in order to avoid this situation, the target data set may be screened for security data, and the criteria for defining data security and danger may be flexibly adjusted as needed, that is, the target data set includes data to be screened for security data, and the screening method at this time includes a method for screening for security data, a reference feature set, that is, a feature set used to perform security screening on the target data, and the reference feature set may be generated based on defined security data, etc. Accordingly, in the process of performing weight operations, linear transformations, and nonlinear transformations on the target data set to obtain a target feature set corresponding to the target data set, a weight operation may be performed on the target data set to obtain a second weight operation result; second compensation data corresponding to the target data set may be determined; a linear transformation may be performed on the second compensation data and the second weight operation result to obtain a second linear transformation result; a nonlinear transformation may be performed on the second linear transformation result to obtain a first nonlinear transformation result; a nonlinear transformation is used to map the second linear transformation result to a positive number less than or equal to 1, and a nonlinear transformation function may be flexibly adjusted as needed, for example, a nonlinear transformation function may be , x∈(-∞,+∞), or it can be , , x∈(-∞,+∞), etc.; perform weight operation on the first nonlinear transformation result to obtain the target feature set corresponding to the target data set. The extraction process of the target feature set can be as follows: Figure 4 As shown, A2 and B represent data weights, P2 represents the second compensation data, and the target feature set can be expressed as H=f2(A2*X+P2), Y=B×H, and X represents the target data set in matrix form, that is:
[0115] ;
[0116] .
[0117] From this implementation process, it can be seen that in the process of screening security data, the present invention can perform a weight operation on the target data set to obtain a second weight operation result, and determine the second compensation data corresponding to the target data set; perform a linear transformation on the second compensation data and the second weight operation result to obtain a second linear transformation result; perform a nonlinear transformation on the second linear transformation result to obtain a first nonlinear transformation result, and the nonlinear transformation is used to map the second linear transformation result to a positive number less than or equal to 1. In this way, the error between the target feature set and the reference feature set can be reduced, and the screening accuracy of security data and dangerous data can be improved; and then the first nonlinear transformation result needs to be weighted again to further optimize the features of the security data and dangerous data in the first nonlinear transformation result according to the data weight, so as to obtain a target feature set that is easier to screen security data, thereby improving the efficiency and accuracy of subsequent security data screening.
[0118] In a specific application scenario, in the process of comparing the target feature set and the reference feature set to obtain a comparison result, for each group of data in the target data set, the target feature group of the data group can be determined in the target feature set, and the reference feature group of the data group can be determined in the reference feature set; a second difference between the reference feature group and the target feature group is generated, and a second error rate of the data group is generated according to the second difference; it is detected whether the second error rate is less than or equal to a second standard value, and the second standard value can be set to a value between (0,1], and the smaller the second standard value, the higher the detection accuracy; in response to the second error rate being less than or equal to the second standard value, a comparison result characterizing the safety of the data group is obtained; in response to the second error rate being greater than the second standard value, a comparison result characterizing the danger of the data group is obtained.
[0119] To facilitate understanding of this process, assume that the user defines a reference feature set of n safe standard data as E 2i (1≤i≤n), calculate the characteristics of each column of the target data set in turn, that is, calculate each data group Features , (1≤j≤m), then the second error rate of the data group can be expressed as , (1≤j≤m).
[0120] In a specific application scenario, after the target data set is screened according to the comparison result and the screening result is obtained, the dangerous data group in the target data set can be discarded, and the dangerous data group in the target data set can also be processed to obtain a safe data group, that is, the target feature group of the dangerous data group in the target data set can be used as the data group to be purified; weight operation, linear transformation and nonlinear transformation are performed on the data group to be purified to obtain a purified feature set corresponding to the data group to be purified, and the generation principle of the purified feature set is the same as the generation principle of the target feature set for screening safe data in this embodiment; the purified feature set is compared with the reference feature set, and this comparison process is the same as the comparison principle between the target feature set and the reference feature set for screening safe data in this embodiment, and whether the data groups to be purified are all safe is detected according to the comparison result; in response to the data groups to be purified being all safe, the data groups to be purified are used as safe data groups; in response to the existence of dangerous data groups in the data groups to be purified, the purified feature set of the dangerous data groups in the data groups to be purified is used as a new data group to be purified, and the step of performing weight operation, linear transformation and nonlinear transformation on the data group to be purified is returned to obtain the purified feature set corresponding to the data group to be purified.
[0121] exist Figure 1 On the basis of the illustrated embodiments, considering that in some scenarios, it is necessary to filter specific data in the user, such as hoping that the data set does not carry setting data such as illegal content, the type of setting data can be determined according to the application scenario, such as the setting data can be data that violates driving rules, communication specifications or production rules, etc. In order to meet this demand, the target data set can be data to be screened for setting data. At this time, the screening method includes the method to be screened for setting data, the reference feature set is also the feature set used to screen the target data for setting data, the reference feature set can be generated according to the defined setting data, etc. Correspondingly, in the process of performing weight operation, linear transformation and nonlinear transformation on the target data set to obtain the target feature set corresponding to the target data set, the target data set can be weighted to obtain a third weight operation result; determine the third compensation data corresponding to the target data set, perform linear transformation on the third compensation data and the third weight operation result, and obtain a third linear transformation result; perform nonlinear transformation on the third linear transformation result to obtain the target feature set corresponding to the target data set; and the nonlinear transformation is used to map the third linear transformation result to a positive number less than or equal to 1.
[0122] From this implementation process, it can be seen that in the process of screening the set data, the present invention can perform a weight operation on the target data set to obtain a third weight operation result, and determine the third compensation data corresponding to the target data set; perform a linear transformation on the third compensation data and the third weight operation result to obtain a third linear transformation result; perform a nonlinear transformation on the third linear transformation result to obtain a target feature set corresponding to the target data set, and the nonlinear transformation is used to map the third linear transformation result to a positive number less than or equal to 1. In this way, the error between the features of the set data and the reference feature set can be reduced, thereby improving the screening accuracy of the set data.
[0123] In a specific application scenario, in the process of comparing the target feature set and the reference feature set to obtain the comparison result, the length value of the word vector can be determined. For example, the length of the word vector in the reference feature set can be counted, and the length obtained by counting can be used as the length value; in the target feature set, a feature phrase with a length equal to the length value is read out; for each feature phrase, whether the feature phrase contains the set feature in the reference feature set is detected. If so, a comparison result is obtained that the corresponding data group of the feature phrase in the target data set has the set data; if not, a comparison result is obtained that the corresponding data group of the feature phrase in the target data set does not have the set data. And after filtering the target data set according to the comparison result and obtaining the filtering result, the data group with the set data in the target data set can also be filtered to obtain a processed data set, so that the processed data set does not carry the set data.
[0124] It should be noted that the sliding window filter function can be used to take the length value as the window length to select the value of the target feature set, and compare it with the set features in the reference feature set, and finally decide whether to filter the corresponding data group based on the comparison result. Take the process of screening illegal sentences as an example, Figure 5 As shown in the figure, the window length value is len (1≤len≤n), and the window filter function reads the feature vectors of len consecutive word vectors in the target feature set Y(n×L) in turn to obtain the feature vector set S={S1,S2,…} of all sentences. The feature vectors S of all sentences are compared with the vector Q(m×L) of the illegal features in the reference feature set. If a certain S j Contains vectors in Q(m×L), indicating the difference with S j The corresponding statement contains illegal information, and the statement in the target data set needs to be filtered out and cannot be used for output; if S j Does not contain vectors in Q(m×L), indicating that it is different from S j The corresponding statement does not contain any violation information and does not need to be filtered. The statement in the target data set can be directly output.
[0125] In a specific application scenario, in the process of obtaining a reference feature set for filtering a target data set, an actual data set carrying set data can be obtained. This actual data set can be determined according to the data application scenario, such as a data set carrying set data in a certain business scenario provided to a user; a weight operation is performed on the actual data set to obtain a fourth weight operation result; fourth compensation data corresponding to the actual data set is determined; a linear transformation is performed on the fourth compensation data and the fourth weight operation result to obtain a fourth linear transformation result; a nonlinear transformation is performed on the fourth linear transformation result to obtain an actual feature set corresponding to the actual data set; the nonlinear transformation is used to map the fourth linear transformation result to a positive number less than or equal to 1; a set feature set is obtained, the set feature set includes features of the set data, and features that exist in both the actual feature set and the set feature set are combined into a reference feature set, that is, features in the actual feature set that are the same as those in the set feature set are used as reference feature sets. Assuming that the actual data set is represented by a vector as , the corresponding actual feature set is It means that D=f2(A4*W+P4), , A4 represents the corresponding data weight, P4 represents the fourth compensation data, and the feature set is set with E 3i (1≤i≤n) indicates that the third difference between the features in the actual feature set and the features in the set feature set can be generated, and the third error rate of the data can be generated according to the third difference. For example, the third error rate is , (1≤i≤n), detect whether the third error rate is less than or equal to the third standard value, the third standard value can be set to a value between (0,1], the smaller the third standard value, the higher the detection accuracy; in response to the third error rate being less than or equal to the third standard value, the features in the actual feature set are used as the features in the reference feature set, and in response to the third error rate being greater than the third standard value, the features in the actual feature set are prohibited from being used as the features in the reference feature set. The entire data processing process is as follows Figure 6 As shown, the target dataset is represented by word vector as , (1≤i≤n), the corresponding data weight is represented by A3, and the third compensation data is represented by P3. Then the target feature set at this time can be expressed as Y=f2(A3*X+P3), .
[0126] It should be noted that the actual data set may be a data set carrying set data corresponding to the target data set. Taking the process of asking and answering illegal questions as an example, the actual data set may be questions carrying illegal content, and the target data set may be answers carrying illegal content. For example, the actual data set may be “how to safely run a red light”, and the target data set may be its corresponding answer, etc.
[0127] In order to facilitate understanding of the scheme of the present invention, the process of applying the data screening method of the present invention to perform text processing on an AI (Artificial Intelligence) model is now described. The processing method of image data or language data is similar. Tools can be used to extract relevant information of the image or language, and then converted into a text format for processing. The scheme of the present invention can screen the training data, and can also screen the output data of the AI model, and can perform multiple screenings, etc. For example, the text processing process can include the following steps:
[0128] Authentication is performed on users who use generative AI models. Only authenticated users can use the AI model for text processing to prevent illegal users from damaging the AI model by inserting malicious data. Authentication methods include users entering the correct username and password, holding a legal digital certificate, etc.
[0129] After the user's identity is authenticated, the basic data set and the original hash value of the basic data set are obtained;
[0130] Perform hash operations on the basic data set to generate actual hash values;
[0131] Check whether the actual hash value is consistent with the original hash value. If not, end the process. If consistent, proceed to the next step.
[0132] Screening the important data in the basic data set, encrypting the screened important data, and screening the unimportant data in the basic data set, desensitizing the screened unimportant data to obtain a first data set;
[0133] Screening and purifying dangerous data in the first data set to obtain a second data set without data poisoning risk;
[0134] Using the second data set to train the AI model to obtain a trained AI model;
[0135] Obtaining a search statement input by a user, and generating a violation feature set for screening violation data based on the search statement;
[0136] Obtain the initial search results of the trained AI model for the search statement;
[0137] Screening and filtering the illegal contents in the initial search results according to the illegal feature set, and outputting the search results without illegal contents;
[0138] During the entire process, the user's behavior in using the AI model is audited, and the user's login time, operations performed, etc. are recorded. Logs are generated and stored externally so that administrators can determine whether there are any violations, abnormal operations, etc. through log information, so as to take timely remedial measures to ensure the security of the AI model.
[0139] Based on the data screening method provided in the above embodiment, the present invention also provides a data screening device, including a computing card, a power module and a clock module connected to the computing card, and an operation interface connected to the computing card; and the computing card, the power module and the clock module are mounted on a board; the operation interface is used to operate the target data set to be processed and the screening result; the computing card is used to determine the screening method of the target data set; according to the screening method, the target data set is weighted, error compensated and data distinguished to obtain a target feature set corresponding to the target data set; a reference feature set is obtained for screening the target data set; the target feature set is compared with the reference feature set to obtain a comparison result; the target data set is screened according to the comparison result to obtain a screening result.
[0140] In practical applications, the data screening device provided by the present invention can be built based on a main processor, and the structure can be as follows: Figure 7 As shown, the computing card can be an FPGA (Field-Programmable Gate Array), the board can be a PCI-E (peripheral component interconnect express, a high-speed serial computer expansion bus standard) board, and the main processor can be a server, workstation or other equipment with network communication functions. The FPGA-based PCI-E board includes an FPGA, a random number generator, a power supply / clock module, and a program interface. The FPGA includes a protocol conversion core, a data cache, a main control module, a read-only memory, an encryption algorithm engine, and a hash algorithm engine. The functions of each device are as follows:
[0141] The host computer software is also the user's operating interface for data processing. It can process the original data to form a target data set and send it to the data buffer area of the PCI-E board through the data transmission interface; the PCI-E board transmits the screening results to the host computer software through the data transmission interface; the keyword library can be used to store reference feature sets, etc.
[0142] The protocol conversion core can be a PCI-E protocol IP core, which realizes bus protocol conversion and converts the physical PCI-E bus into a bus signal on the board end.
[0143] The data buffer area can be a volatile memory such as RAM (Random Access Memory) or FIFO (Firstin First out), which is used to cache basic data sets, target data sets, keywords, final output content and other data, and send them to related modules for data processing;
[0144] The main control module can adopt state machine control logic as the control unit of FPGA, which is used to receive the basic data set in the data buffer area, call the encryption algorithm engine, hash algorithm engine, read-only memory and other modules to process the basic data set, generate a target data set, and send the target data set to the GPU (Graphics Processing Unit, graphics processing unit) module through the data buffer area; receive the generated content of the GPU module through the data buffer area, and filter the generated content according to the keyword library of the host computer software, and finally return the content to the host computer software through the data buffer area for the user to read; execute the data screening method of the present invention;
[0145] The read-only memory can be a non-volatile memory such as ROM (Read Only Memory). The data will not be lost after the board is powered off. It is used to store required data, such as storing feature tables.
[0146] The encryption algorithm engine is used to implement symmetric encryption algorithm functions such as AES, 3DES, and SM4;
[0147] The hashing engine is used to implement hashing algorithm functions such as SHA-256 and SM3;
[0148] The random number generator is used to generate physical true random numbers as the key of the encryption algorithm. The key and the encryption algorithm together realize the encryption function of the data;
[0149] The clock module can use a crystal oscillator as the clock module to provide clock frequency for each FPGA module;
[0150] The power module provides operating voltage for each module of the FPGA board, such as 3.3V, 2.5V, 1.2V, etc.
[0151] The program interface may be a JATG (Joint Test Action Group) / AS interface, etc., which is a program debugging / downloading interface for debugging and downloading FPGA programs;
[0152] The GPU module can authenticate users, generate reference feature sets, filter data, audit user operations, etc. It can also interact with the FPGA's data cache, hash algorithm engine, and read-only memory to complete the data security processing process. In addition, it can generate logs that can be read by users through the host computer software.
[0153] The present invention also provides an electronic device and a computer non-volatile readable storage medium, both of which have the corresponding effects of a data screening method provided in an embodiment of the present invention. Figure 8 , Figure 8 A schematic diagram of the structure of an electronic device provided by an embodiment of the present invention.
[0154] An electronic device provided by an embodiment of the present invention includes a memory 201 and a processor 202. The memory 201 stores a computer program. When the processor 202 executes the computer program, the steps of the data screening method described in any of the above embodiments are implemented.
[0155] See also Fig. 9 Another electronic device provided in an embodiment of the present invention may also include: an input port 203 connected to the processor 202, used to transmit commands input from the outside to the processor 202; a display unit 204 connected to the processor 202, used to display the processing results of the processor 202 to the outside; a communication module 205 connected to the processor 202, used to realize communication between the electronic device and the outside. The display unit 204 can be a display panel, a laser scanning display, etc.; the communication method adopted by the communication module 205 includes but is not limited to mobile high-definition link technology (Mobile High-Definition Link, MHL), Universal Serial Bus (Universal Serial Bus, USB), High-Definition Multimedia Interface (High-Definition Multimedia Interface, HDMI), wireless connection: wireless fidelity technology (WIreless Fidelity, WiFi), Bluetooth communication technology, low-power Bluetooth communication technology, and communication technology based on IEEE802.11s.
[0156] A computer-readable storage medium provided in an embodiment of the present invention may be a computer non-volatile readable storage medium, etc. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data screening method described in any of the above embodiments are implemented.
[0157] The computer-readable storage medium involved in the present invention includes random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, hard disk, removable disk, CD-ROM (Compact Disc Read-Only Memory), or any other form of storage medium known in the technical field.
[0158] A computer program product provided by an embodiment of the present invention includes a computer program / instruction, which, when executed by a processor, implements the steps of the data screening method described in any of the above embodiments.
[0159] For the description of the relevant parts of a data screening device, a computer program product, an electronic device, and a computer non-volatile readable storage medium provided by the embodiment of the present invention, please refer to the detailed description of the corresponding parts in a data screening method provided by the embodiment of the present invention, which will not be repeated here. In addition, the parts of the above technical solutions provided by the embodiment of the present invention that are consistent with the implementation principles of the corresponding technical solutions in the prior art are not described in detail to avoid excessive elaboration.
[0160] It should also be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0161] The above description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A data screening method, characterized in that: Applied to computing cards, including: Obtain the target data set to be processed transmitted by the operation interface; Determining a screening method for the target data set; According to the screening method, weight adjustment, error compensation and data differentiation are performed on the target data set to obtain a target feature set corresponding to the target data set; Acquire a reference feature set for screening the target data set; Comparing the target feature set with the reference feature set to obtain a comparison result; Filtering the target data set according to the comparison result to obtain a filtering result, and transmitting the filtering result to the operation interface; The target data set is weighted, error compensated, and data differentiated to obtain a target feature set corresponding to the target data set, including: Performing weight operation, linear transformation and nonlinear transformation on the target data set to obtain a target feature set corresponding to the target data set; Among them, weight operation is used to adjust the weight of data according to the data weight; linear transformation is used to compensate for data errors; and nonlinear transformation is used to distinguish data.
2. The data screening method according to claim 1, characterized in that: The screening method includes a method of screening by importance; Performing weight operation, linear transformation and nonlinear transformation on the target data set to obtain a target feature set corresponding to the target data set, including: Performing a weight operation on the target data set to obtain a first weight operation result; Determining first compensation data corresponding to the target data set; Performing a linear transformation on the first compensation data and the first weight operation result to obtain a first linear transformation result; A nonlinear transformation is performed on the first linear transformation result to obtain a target feature set corresponding to the target data set; the nonlinear transformation is used to amplify non-negative values in the first linear transformation result and reduce negative values in the first linear transformation result.
3. The data screening method according to claim 2, characterized in that: Comparing the target feature set with the reference feature set to obtain a comparison result includes: For each data in the target data set, determining a target feature of the data in the target feature set, and determining a reference feature of the data in the reference feature set; generating a first difference between the reference feature and the target feature, and generating a first error rate of data according to the first difference; Detecting whether the first error rate is less than or equal to a first standard value; In response to the first error rate being less than or equal to the first standard value, obtaining a comparison result indicating that the data is important data; In response to the first error rate being greater than the first standard value, a comparison result indicating that the data is non-significant data is obtained.
4. The data screening method according to claim 3, characterized in that: The target data set is screened according to the comparison result, and after the screening result is obtained, the method further includes: The important data in the target data set is encrypted, and the non-important data in the target data set is desensitized.
5. The data screening method according to claim 1, characterized in that: The screening method includes a method for performing security data screening; Performing weight operation, linear transformation and nonlinear transformation on the target data set to obtain a target feature set corresponding to the target data set, including: Performing a weight operation on the target data set to obtain a second weight operation result; Determining second compensation data corresponding to the target data set; Performing a linear transformation on the second compensation data and the second weight operation result to obtain a second linear transformation result; Performing a nonlinear transformation on the second linear transformation result to obtain a first nonlinear transformation result; the nonlinear transformation is used to map the second linear transformation result to a positive number less than or equal to 1; A weight operation is performed on the first nonlinear transformation result to obtain a target feature set corresponding to the target data set.
6. The data screening method according to claim 5, characterized in that: Comparing the target feature set with the reference feature set to obtain a comparison result includes: For each group of data in the target data set, determining a target feature group of the data group in the target feature set, and determining a reference feature group of the data group in the reference feature set; generating a second difference between the reference feature group and the target feature group, and generating a second error rate of the data group according to the second difference; Detecting whether the second error rate is less than or equal to a second standard value; In response to the second error rate being less than or equal to the second standard value, obtaining a comparison result indicating the safety of the data group; In response to the second error rate being greater than the second standard value, a comparison result characterizing the risk of the data set is obtained.
7. The data screening method according to claim 6, characterized in that: The target data set is screened according to the comparison result, and after the screening result is obtained, the method further includes: Taking the target feature group of the dangerous data group in the target data set as the data group to be purified; Perform weight calculation, linear transformation and nonlinear transformation on the data group to be purified to obtain a purification feature set corresponding to the data group to be purified; Comparing the cleansing feature set with the reference feature set, and detecting whether the data groups to be cleansed are all safe according to the comparison result; In response to the data groups to be purified being all safe, the data groups to be purified are used as safe data groups; In response to the existence of a dangerous data group in the data group to be purified, the purified feature set of the dangerous data group in the data group to be purified is used as a new data group to be purified, and the step of performing weight calculation, linear transformation and nonlinear transformation on the data group to be purified is returned to obtain the purified feature set corresponding to the data group to be purified.
8. The data screening method according to claim 1, characterized in that: The screening method includes a method for setting data screening; Performing weight operation, linear transformation and nonlinear transformation on the target data set to obtain a target feature set corresponding to the target data set, including: Performing a weight operation on the target data set to obtain a third weight operation result; Determining third compensation data corresponding to the target data set; Performing a linear transformation on the third compensation data and the third weight operation result to obtain a third linear transformation result; A nonlinear transformation is performed on the third linear transformation result to obtain a target feature set corresponding to the target data set; the nonlinear transformation is used to map the third linear transformation result to a positive number less than or equal to 1.
9. The data screening method according to claim 8, characterized in that: Comparing the target feature set with the reference feature set to obtain a comparison result includes: Determine the length of the word vector; In the target feature set, a feature phrase having a length equal to the length value is read out; For each of the feature phrases, detect whether the feature phrase contains the set features in the reference feature set. If so, obtain a comparison result indicating that the corresponding data group of the feature phrase in the target data set has the set data. If not, obtain a comparison result indicating that the corresponding data group of the feature phrase in the target data set does not have the set data.
10. The data screening method according to claim 9, characterized in that: The target data set is screened according to the comparison result, and after the screening result is obtained, the method further includes: The data groups containing setting data in the target data set are filtered to obtain a processed data set.
11. The data screening method according to any one of claims 8 to 10, characterized in that: Obtaining a reference feature set for screening the target data set includes: Get the actual data set that carries the set data; Performing a weight operation on the actual data set to obtain a fourth weight operation result; Determining fourth compensation data corresponding to the actual data set; Performing a linear transformation on the fourth compensation data and the fourth weight operation result to obtain a fourth linear transformation result; Performing a nonlinear transformation on the fourth linear transformation result to obtain an actual feature set corresponding to the actual data set; the nonlinear transformation is used to map the fourth linear transformation result to a positive number less than or equal to 1; A set feature set is obtained, and features existing in both the actual feature set and the set feature set are combined into a reference feature set.
12. A data screening device, characterized in that: It includes a computing card, a power module and a clock module connected to the computing card, and an operation interface connected to the computing card; and the computing card, the power module and the clock module are mounted on a board; The operation interface is used to operate the target data set to be processed and the screening results; The computing card is used to determine a screening method for the target data set; According to the screening method, weight adjustment, error compensation and data differentiation are performed on the target data set to obtain a target feature set corresponding to the target data set; and a reference feature set for screening the target data set is obtained; Comparing the target feature set with the reference feature set to obtain a comparison result; Filtering the target data set according to the comparison result to obtain the filtering result; The target data set is weighted, error compensated, and data differentiated to obtain a target feature set corresponding to the target data set, including: Performing weight operation, linear transformation and nonlinear transformation on the target data set to obtain a target feature set corresponding to the target data set; Among them, weight operation is used to adjust the weight of data according to the data weight; linear transformation is used to compensate for data errors; and nonlinear transformation is used to distinguish data.
13. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the steps of the data screening method according to any one of claims 1 to 11 are implemented.
14. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the steps of the data screening method according to any one of claims 1 to 11 when executing the computer program.
15. A computer non-volatile readable storage medium, characterized in that: The computer non-volatile readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the data screening method according to any one of claims 1 to 11 are implemented.
Citation Information
Patent Citations
Data processing method, storage medium and computer terminal
CN114943273A
Radar reconnaissance big data processing method and platform
CN117688304A