Data processing method, device and storage medium
Automatic data cleaning through feature extraction models and similarity algorithms solves the problem of low efficiency of manual labeling in existing technologies and achieves efficient and accurate data cleaning.
Patent Information
- Application Number
- CN202310678932.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-08
- Publication Date
- 2025-09-16
- Estimated Expiration
- 2043-06-08
AI Technical Summary
Existing data cleaning methods are based on manual labeling, which is inefficient and cannot meet the needs of fast and efficient data analysis.
By obtaining the preset scene tracking information and the information to be trained, and using the feature extraction model and similarity algorithm, data cleaning is automatically performed to eliminate data irrelevant to the preset scene.
It improves the accuracy and efficiency of data cleaning, reduces manual intervention, and meets the data analysis needs under the business model.
Smart Images

Figure CN116821109B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to a data processing method, device, and storage medium. Background Art
[0002] Data tracking is a commonly used data collection method that mainly collects data based on specific user behaviors or time characteristics, and can provide fast, efficient, and rich data support.
[0003] Currently, when tracking data is collected and reported, it may be incomplete or contain errors due to various reasons, making it unsuitable for subsequent data analysis and business operations. Therefore, tracking data usually needs to be cleaned before it can be used.
[0004] Existing data cleaning methods are usually based on manual labeling, which is time-consuming, labor-intensive and inefficient. Summary of the Invention
[0005] The present application provides a data processing method, device and storage medium for solving the technical problem of low efficiency of data cleaning of buried point data by general technology.
[0006] To achieve the above objectives, this application adopts the following technical solutions:
[0007] In a first aspect, a data processing method is provided, comprising:
[0008] Obtaining preset scene buried point information and information to be trained; the preset scene buried point information includes preset scene buried point data; the information to be trained includes data to be trained;
[0009] Performing model training based on the preset scene buried point information to obtain a feature extraction model and a first information feature corresponding to the preset scene buried point information;
[0010] Inputting the information to be trained into the feature extraction model to obtain a second information feature corresponding to the information to be trained;
[0011] Target training data corresponding to the to-be-trained data is determined based on the similarity between the first information feature and the second information feature.
[0012] Optionally, when the preset scene buried point information further includes: a first clustering result for representing a data category of the preset scene buried point data, the first information feature includes a first clustering feature of the first clustering result; when the information to be trained further includes: a second clustering result for representing a data category of the data to be trained, the second information feature includes a second clustering feature of the second clustering result; obtaining the preset scene buried point information and the information to be trained includes:
[0013] Clustering the preset scene buried point data according to the first clustering algorithm to obtain a first clustering result;
[0014] The training data is clustered according to the second clustering algorithm to obtain a second clustering result.
[0015] Optionally, determining target training data corresponding to the to-be-trained data based on the similarity between the first information feature and the second information feature includes:
[0016] When the similarity between the first information feature and the second information feature is greater than or equal to a first preset similarity, determining the to-be-trained data as target training data;
[0017] When the similarity between the first information feature and the second information feature is less than or equal to the second preset similarity, the data to be trained is removed; the second preset similarity is less than the first preset similarity;
[0018] When the similarity between the first information feature and the second information feature is less than a first preset similarity and greater than a second preset similarity, interpolation processing is performed on the training data to obtain target training data.
[0019] Optionally, the data processing method further includes:
[0020] Based on the target training data and the preset scene buried point information, the feature extraction model is trained to obtain the target model; the target model is used to extract the preset scene buried point data from the data to be processed.
[0021] In a second aspect, a data processing device is provided, comprising: an acquisition unit and a processing unit;
[0022] An acquisition unit is used to acquire preset scene buried point information and information to be trained; the preset scene buried point information includes preset scene buried point data; the information to be trained includes the data to be trained;
[0023] A processing unit, configured to perform model training based on the preset scene buried point information to obtain a feature extraction model and a first information feature corresponding to the preset scene buried point information;
[0024] The processing unit is further configured to input the information to be trained into the feature extraction model to obtain a second information feature corresponding to the information to be trained;
[0025] The processing unit is further configured to determine target training data corresponding to the to-be-trained data based on the similarity between the first information feature and the second information feature.
[0026] Optionally, when the preset scene buried point information further includes: a first clustering result for representing a data category of the preset scene buried point data, the first information feature includes a first clustering feature of the first clustering result; when the to-be-trained information further includes: a second clustering result for representing a data category of the to-be-trained data, the processing unit is specifically configured to:
[0027] Clustering the preset scene buried point data according to the first clustering algorithm to obtain a first clustering result;
[0028] The training data is clustered according to the second clustering algorithm to obtain a second clustering result.
[0029] Optionally, a processing unit is used to:
[0030] When the similarity between the first information feature and the second information feature is greater than or equal to a first preset similarity, determining the to-be-trained data as target training data;
[0031] When the similarity between the first information feature and the second information feature is less than or equal to the second preset similarity, the data to be trained is removed; the second preset similarity is less than the first preset similarity;
[0032] When the similarity between the first information feature and the second information feature is less than a first preset similarity and greater than a second preset similarity, interpolation processing is performed on the training data to obtain target training data.
[0033] Optionally, the processing unit is also used to perform model training on the feature extraction model based on the target training data and the preset scene buried point information to obtain a target model; the target model is used to extract the preset scene buried point data from the data to be processed.
[0034] In a third aspect, a data processing device is provided, comprising a memory and a processor; the memory is used to store computer-executable instructions, and the processor is connected to the memory via a bus; when the data processing device is running, the processor executes the computer-executable instructions stored in the memory, so that the data processing device executes the data processing method described in the first aspect.
[0035] The data processing device may be a network device, or a portion of a network device, such as a chip system in the network device. The chip system is used to support the network device in implementing the functions involved in the first aspect and any possible implementation thereof, such as acquiring, determining, and sending the data and / or information involved in the above-mentioned data processing method. The chip system includes a chip and may also include other discrete devices or circuit structures.
[0036] In a fourth aspect, a computer-readable storage medium is provided, the computer-readable storage medium including computer execution instructions, which, when executed on a computer, enable the computer to execute the data processing method described in the first aspect.
[0037] In a fifth aspect, a computer program product is also provided, which includes computer instructions. When the computer instructions are executed on a data processing device, the data processing device executes the data processing method as described in the first aspect above.
[0038] It should be noted that the above-mentioned computer instructions may be stored in whole or in part on a computer-readable storage medium. The computer-readable storage medium may be packaged together with the processor of the data processing device or separately from the processor of the data processing device, and this embodiment of the application is not limited to this.
[0039] The description of the second, third, fourth and fifth aspects of this application can refer to the detailed description of the first aspect.
[0040] In the embodiments of this application, the names of the aforementioned data processing devices do not limit the devices or functional modules themselves. In actual implementation, these devices or functional modules may appear with other names. For example, the receiving unit may also be called a receiving module, a receiver, etc. As long as the functions of each device or functional module are similar to those of this application, they are within the scope of the claims of this application and their equivalents.
[0041] The technical solution provided by this application brings at least the following beneficial effects:
[0042] Based on any of the above aspects, an embodiment of the present application provides a data processing method, including: obtaining preset scene point information and information to be trained. The preset scene point information includes preset scene point data; the information to be trained includes the data to be trained. Then, model training can be performed based on the preset scene point information to obtain a feature extraction model and a first information feature corresponding to the preset scene point information. Then, the information to be trained can be input into the feature extraction model to obtain a second information feature corresponding to the information to be trained, and based on the similarity between the first information feature and the second information feature, the target training data corresponding to the data to be trained is determined.
[0043] From the above, it can be seen that since this application first extracts the first information feature corresponding to the preset scene burial point information, and then determines the target training data corresponding to the data to be trained based on the similarity between the second information feature corresponding to the information to be trained and the first information feature (that is, data cleaning is performed based on the similarity), the information to be trained is brought as close to the preset scene as possible, thereby eliminating data in the information to be trained that is not related to the preset scene, thereby improving the accuracy of data cleaning.
[0044] The beneficial effects of the first, second, third, fourth and fifth aspects of this application can all be referred to the analysis of the above beneficial effects, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 A schematic diagram of the structure of a data processing system provided in an embodiment of the present application;
[0046] Figure 2 A schematic diagram of the hardware structure of a data processing device provided in an embodiment of the present application Figure 1 ;
[0047] Figure 3 A schematic diagram of the hardware structure of a data processing device provided in an embodiment of the present application Figure 2 ;
[0048] Figure 4 A schematic diagram of a data processing method provided in an embodiment of the present application Figure 1 ;
[0049] Figure 5 A schematic diagram of a data processing method provided in an embodiment of the present application Figure 2 ;
[0050] Figure 6 A schematic diagram of the structure of a data processing device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0051] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0052] It should be noted that in the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0053] In order to facilitate a clear description of the technical solutions of the embodiments of the present application, in the embodiments of the present application, words such as "first" and "second" are used to distinguish between identical or similar items with basically the same functions and effects. Those skilled in the art can understand that words such as "first" and "second" do not limit the quantity and execution order.
[0054] Data tracking is a commonly used data collection method that mainly collects data based on specific user behaviors or time characteristics, and can provide fast, efficient, and rich data support.
[0055] The collected feature data can be used to analyze the usage of websites or applications (APPs), user behavior habits, etc., and is the basis for establishing data products such as user portraits and user behavior paths.
[0056] However, when the amount of tracking data is large enough, dirty data (generally referring to data that is untrue, incomplete, incorrect, duplicate, or meaningless) may be generated due to product anomalies, network transmission failures, improper configuration of data collection components, etc. Dirty data is often difficult to correct and utilize, so it is usually necessary to clean the tracking data before it can be used.
[0057] The goal of data cleaning is to eliminate dirty data as much as possible and repair the impact of dirty data. Existing data cleaning methods are usually based on manual labeling, which is time-consuming, labor-intensive, and inefficient.
[0058] In addition, existing technologies can also use unsupervised model methods to perform data cleaning. However, unsupervised model methods do not take into account the needs of the business model and cannot meet the data analysis requirements under the required business model.
[0059] Existing technologies can also perform data cleaning through data verification mode methods. However, the data cleaning mode through data verification mode methods is relatively fixed, lacks flexibility, and the model generalization ability is not strong enough.
[0060] In response to the above problems, an embodiment of the present application provides a data processing method, comprising: obtaining preset scene point information and information to be trained. The preset scene point information includes preset scene point data; the information to be trained includes the data to be trained. Then, model training can be performed based on the preset scene point information to obtain a feature extraction model and a first information feature corresponding to the preset scene point information. Then, the information to be trained can be input into the feature extraction model to obtain a second information feature corresponding to the information to be trained, and based on the similarity between the first information feature and the second information feature, the target training data corresponding to the data to be trained is determined.
[0061] From the above, it can be seen that since this application first extracts the first information feature corresponding to the preset scene burial point information, and then determines the target training data corresponding to the data to be trained based on the similarity between the second information feature corresponding to the information to be trained and the first information feature (that is, data cleaning is performed based on the similarity), the information to be trained is brought as close to the preset scene as possible, thereby eliminating data in the information to be trained that is not related to the preset scene, thereby improving the accuracy of data cleaning.
[0062] The data processing method is applicable to a data processing system. Figure 1 FIG. 1 shows a structure of the data processing system. Figure 1 As shown, the data processing system includes: an electronic device 101, a tracking point server 102 and a data providing device 103.
[0063] Among them, the electronic device 101 is communicated with the tracking server 102 and the data providing device 103 respectively.
[0064] In practical applications, the electronic device 101 can be connected to any number of tracking servers and data providing devices. Figure 1 An example of an electronic device 101 connected to a tracking server 102 and a data providing device 103 is used for explanation.
[0065] Optionally, the physical devices of the electronic device 101 and the data providing device 103 may be servers, terminals, or other types of electronic devices, which is not limited in the embodiments of the present application.
[0066] In an embodiment of the present application, the data providing device 103 can provide the information to be trained to the electronic device 101, and the burying point server can provide the preset scene burying point information to the electronic device 101. The electronic device can clean the dirty data in the training information based on the information to be trained provided by the data providing device 103 and the preset scene burying point information provided by the burying point server, thereby obtaining the target training data.
[0067] Optionally, the terminal may be a device that provides voice and / or data connectivity to a user, a handheld device with wireless connection capabilities, or other processing devices connected to a wireless modem. A wireless terminal may communicate with one or more core networks via a radio access network (RAN). A wireless terminal may be a mobile terminal, such as a mobile phone (or "cellular" phone) and a computer with a mobile terminal, or a portable, pocket-sized, handheld, computer-built-in, or vehicle-mounted mobile device that exchanges voice and / or data with a radio access network, such as a mobile phone, tablet computer, laptop computer, netbook, or personal digital assistant (PDA).
[0068] Optionally, the above-mentioned server can be a server in a server cluster (consisting of multiple servers), or a chip in the server, or a system on a chip in the server, or can be implemented through a virtual machine (VM) deployed on a physical machine. This embodiment of the present application does not limit this.
[0069] The basic hardware structure of electronic device 101 includes Figure 2 or Figure 3 The data processing device shown in FIG. Figure 2 and Figure 3 Taking the data processing device shown in FIG. 1 as an example, the hardware structure of the electronic device 101 is introduced.
[0070] like Figure 2 FIG2 is a schematic diagram of a hardware structure of a data processing device provided in an embodiment of the present application. The data processing device includes a processor 21, a memory 22, a communication interface 23, and a bus 24. The processor 21, the memory 22, and the communication interface 23 can be connected via a bus 24.
[0071] Processor 21 is the control center of the data processing device and can be a single processor or a collective term for multiple processing elements. For example, processor 21 can be a general-purpose central processing unit (CPU) or other general-purpose processor. The general-purpose processor can be a microprocessor or any conventional processor.
[0072] As an embodiment, the processor 21 may include one or more CPUs, such as Figure 2 CPU 0 and CPU 1 are shown in Figure 1.
[0073] The memory 22 may be a read-only memory (ROM) or other type of static network access device that can store static information and instructions, a random access memory (RAM) or other type of dynamic network access device that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a disk storage medium or other magnetic network access device, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto.
[0074] In one possible implementation, the memory 22 may exist independently of the processor 21 and may be connected to the processor 21 via a bus 24 for storing instructions or program codes. When the processor 21 calls and executes the instructions or program codes stored in the memory 22, the data processing method provided in the following embodiments of the present application can be implemented.
[0075] In the embodiment of the present application, for the electronic device 101, the software programs stored in the memory 22 are different, so the functions implemented by the electronic device 101 are different. The functions performed by each device will be described in conjunction with the following flowchart.
[0076] In another possible implementation, the memory 22 may also be integrated with the processor 21 .
[0077] The communication interface 23 is used to connect the data processing device to other devices via a communication network, which may be Ethernet, a wireless access network, a wireless local area network (WLAN), etc. The communication interface 23 may include a receiving unit for receiving data and a sending unit for sending data.
[0078] The bus 24 may be an ISA bus, a PCI bus, or an EISA bus. The bus may be divided into an address bus, a data bus, a control bus, and the like. For ease of presentation, Figure 2 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.
[0079] Figure 3 FIG. 2 shows another hardware structure of the data processing device in the embodiment of the present application. Figure 3 As shown, the data processing device may include a processor 31 and a communication interface 32. The processor 31 is coupled to the communication interface 32.
[0080] The functions of the processor 31 may refer to the description of the processor 21. In addition, the processor 31 also has a storage function and can play the role of the memory 22.
[0081] The communication interface 32 is used to provide data to the processor 31. The communication interface 32 can be an internal interface of the data processing device, or an external interface of the data processing device (equivalent to the communication interface 23).
[0082] It should be pointed out that Figure 2 (or Figure 3 ) does not constitute a limitation on the data processing device, except Figure 2 (or Figure 3 ) In addition to the components shown in the figure, the data processing device may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.
[0083] The data processing method provided in the embodiments of the present application is described in detail below with reference to the accompanying drawings.
[0084] The data processing method provided in the embodiment of the present application is applied to Figure 1 The electronic device 101 in the data processing system shown is Figure 4 As shown, the data processing method provided in the embodiment of the present application includes:
[0085] S401: The electronic device obtains preset scene tracking information and information to be trained.
[0086] The preset scene tracking information includes preset scene tracking data, and the to-be-trained information includes to-be-trained data.
[0087] Optionally, the electronic device can obtain preset scene burial point information (also called preset scene burial point data set) under a preset scene from the burial point server, and obtain information to be trained (also called training sample set to be cleaned) from the data providing device.
[0088] Optionally, the data providing device may be a user device, a cloud server, or a local server.
[0089] Optionally, the preset scenario may be a scenario for data security warning, a scenario for building a user profile, a scenario for user behavior prediction, etc.
[0090] The information to be trained may include data to be trained in various scenarios.
[0091] S402. The electronic device performs model training based on the preset scene buried point information to obtain a feature extraction model and a first information feature corresponding to the preset scene buried point information.
[0092] Specifically, after obtaining the preset scene buried point information and the information to be trained, in order to accurately filter out training data close to the preset scene from the information to be trained, the electronic device can perform model training based on the preset scene buried point information to obtain the feature extraction model and the first information feature corresponding to the preset scene buried point information.
[0093] Optionally, the electronic device can use a machine learning algorithm to train the feature extraction model. Machine learning algorithms include but are not limited to: convolutional neural networks, recurrent neural networks, Transformers, etc.
[0094] Optionally, after the electronic device trains the feature extraction model, the preset scene point information can be input into the feature extraction model to obtain the first information feature.
[0095] Optionally, the electronic device can also use but is not limited to the following methods to extract the first information feature: Local Binary Pattern (LBP) feature extraction algorithm, Directed Gradient Histogram feature extraction algorithm, Gaussian-Laplacian (LOG) feature extraction algorithm, Scale Invariant Feature Transform (SIFT) feature extraction operator, Speeded Up Robust Features (SURF) feature extraction algorithm, etc.
[0096] S403: The electronic device inputs the information to be trained into a feature extraction model to obtain a second information feature corresponding to the information to be trained.
[0097] S404: The electronic device determines target training data corresponding to the to-be-trained data based on the similarity between the first information feature and the second information feature.
[0098] Optionally, the electronic device may determine the similarity between the first information feature and the second information feature using, but not limited to, the following methods: Euclidean distance algorithm, Mahalanobis distance algorithm, Minkowski distance algorithm, Hamming distance algorithm, Pearson correlation coefficient algorithm, cosine similarity algorithm, etc.
[0099] Since the first information feature may include multiple first information features corresponding to multiple preset scene burial point information, and the second information feature may include multiple second information features corresponding to multiple information to be trained, the electronic device can determine the similarity between each first information feature and each second information feature, that is, multiple similarities.
[0100] When a certain similarity is greater than a preset similarity, it indicates that the information to be trained corresponding to the similarity is close to the training data of the preset scenario. In this case, the electronic device may determine the information to be trained corresponding to the similarity as target training data.
[0101] Correspondingly, when a certain similarity is less than a preset similarity, it indicates that the information to be trained corresponding to the similarity is not close to the training data of the preset scene. In this case, the electronic device may eliminate the information to be trained corresponding to the similarity.
[0102] As can be seen from the above, an embodiment of the present application provides a data processing method, including: obtaining preset scene point information and information to be trained. The preset scene point information includes preset scene point data; the information to be trained includes the data to be trained. Then, model training can be performed based on the preset scene point information to obtain a feature extraction model and a first information feature corresponding to the preset scene point information. Then, the information to be trained can be input into the feature extraction model to obtain a second information feature corresponding to the information to be trained, and based on the similarity between the first information feature and the second information feature, the target training data corresponding to the data to be trained is determined.
[0103] From the above, it can be seen that since this application first extracts the first information feature corresponding to the preset scene burial point information, and then determines the target training data corresponding to the data to be trained based on the similarity between the second information feature corresponding to the information to be trained and the first information feature (that is, data cleaning is performed based on the similarity), the information to be trained is brought as close to the preset scene as possible, thereby eliminating data in the information to be trained that is not related to the preset scene, thereby improving the accuracy of data cleaning.
[0104] In some embodiments, when the preset scene buried point information further includes: a first clustering result for representing the data category of the preset scene buried point data, the first information feature includes the first clustering feature of the first clustering result. When the information to be trained further includes: a second clustering result for representing the data category of the data to be trained, the second information feature includes the second clustering feature of the second clustering result. In this case, Figure 5 As shown, the data processing method provided in the embodiment of the present application specifically includes:
[0105] S501: The electronic device obtains preset scene buried point data and data to be trained.
[0106] The specific implementation method of the electronic device obtaining the preset scene buried point data and the data to be trained can refer to the specific description of S401 and will not be repeated here.
[0107] S502. The electronic device clusters the preset scene buried point data according to a first clustering algorithm to obtain a first clustering result.
[0108] Specifically, since the amount of preset scene buried point data is large, if similarity calculation is performed on each preset scene buried point data, the amount of calculation is large. In this case, the electronic device can cluster the preset scene buried point data according to the first clustering algorithm to obtain a first clustering result.
[0109] The first clustering algorithm may be, but is not limited to, the following methods: K-means clustering algorithm, K-medoids clustering algorithm, random selection-based clustering algorithm, and the like.
[0110] S503: The electronic device clusters the training data according to a second clustering algorithm to obtain a second clustering result.
[0111] Specifically, since the amount of data to be trained is large, if similarity calculation is performed on each pre-trained data, the amount of calculation is large. In this case, the electronic device can cluster the training data according to the second clustering algorithm to obtain a second clustering result.
[0112] The second clustering algorithm may be, but is not limited to, the following methods: K-means clustering algorithm, K-medoids clustering algorithm, clustering algorithm based on random selection, and the like.
[0113] Optionally, the first clustering algorithm and the second clustering algorithm may be the same clustering algorithm or different clustering algorithms.
[0114] S504: The electronic device performs model training according to the first clustering result to obtain a feature extraction model and a first clustering feature corresponding to the first clustering result.
[0115] Specifically, the electronic device may perform training based on the cluster classification data of the preset scene data (i.e., the first clustering result), obtaining feature data of each cluster classification of the preset scene data (i.e., the first clustering feature) and a training model including a feature extraction layer (i.e., the feature extraction model). S505: The electronic device inputs the second clustering result into the feature extraction model to obtain a second clustering feature corresponding to the information to be trained.
[0116] Specifically, the electronic device can train the cluster classification of the training sample data to be cleaned (i.e., the second clustering result) based on the training model including the feature extraction layer, and extract feature data (i.e., the second clustering feature) of each training sample data in the cluster classification of the training sample data to be cleaned.
[0117] S505: When the similarity between the first information feature and the second information feature is greater than or equal to a first preset similarity, the electronic device determines the to-be-trained data as target training data.
[0118] Specifically, when the first information feature includes a first cluster feature and the second information feature includes a second cluster feature, the electronic device can calculate the clustering distance (i.e., similarity) of the feature data of each training sample data in the cluster classification of the training sample data to be cleaned (i.e., each second cluster feature) and the feature data of each cluster classification of the preset scene data (i.e., each first cluster feature).
[0119] When the distance is less than the first preset distance (i.e., the similarity between the first information feature and the second information feature is greater than or equal to the first preset similarity), it indicates that the training sample data and the preset scene data are highly similar. In this case, the training sample data can be used as the target training data for model training.
[0120] S506: When the similarity between the first information feature and the second information feature is less than or equal to a second preset similarity, the electronic device removes the data to be trained.
[0121] The second preset similarity is smaller than the first preset similarity.
[0122] If the distance is greater than a second preset distance (i.e., the similarity between the first information feature and the second information feature is less than or equal to the second preset similarity), the training sample data and the preset scene data are less similar. In this case, the training sample data is dirty data, and the electronic device can clean the training sample data.
[0123] S507: When the similarity between the first information feature and the second information feature is less than the first preset similarity and greater than the second preset similarity, the electronic device performs interpolation processing on the training data to obtain target training data.
[0124] When the distance is between the first preset distance and the second preset distance (i.e., the similarity between the first information feature and the second information feature is less than the first preset similarity and greater than the second preset similarity), it means that the training sample data and the preset scene data are not very similar, but are somewhat similar. In this case, it means that the training sample data still has analytical value, and the electronic device can choose to interpolate the missing values in the training sample data.
[0125] Optionally, the interpolation method may be, but is not limited to, the following methods: special value, mean, median, mode, random interpolation, multiple interpolation, hot plateau interpolation, Lagrange difference method, Newton interpolation method, adversarial generative network, neural network, etc.
[0126] S508. The electronic device performs model training on the feature extraction model based on the target training data and the preset scene embedding information to obtain a target model.
[0127] Among them, the target model is used to extract preset scene embedded data from the data to be processed.
[0128] Specifically, after determining the target training data, the electronic device can input the cleaned training samples (ie, target training data) into the above-mentioned feature extraction model for training optimization to enhance its generalization ability.
[0129] The above mainly introduces the solution provided by the embodiment of the present application from the perspective of method. In order to realize the above functions, it includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should easily appreciate that, in combination with the units and algorithm steps of each example described in the embodiments disclosed herein, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in a hardware or computer software driven hardware manner depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0130] The embodiment of the present application can divide the data processing device into functional modules according to the above method example. For example, each functional module can be divided according to each function, or two or more functions can be integrated into one processing module. The above integrated modules can be implemented in the form of hardware or in the form of software functional modules. Optionally, the division of modules in the embodiment of the present application is schematic and is only a logical functional division. In actual implementation, there may be other division methods.
[0131] like Figure 6The figure is a schematic diagram of the structure of a data processing device provided in an embodiment of the present application. The data processing device can be used to execute Figure 4-Figure 5 The data processing method shown. Figure 6 The data processing device shown includes: an acquisition unit 601 and a processing unit 602;
[0132] The acquisition unit 601 is used to acquire preset scene buried point information and information to be trained; the preset scene buried point information includes preset scene buried point data; the information to be trained includes the data to be trained;
[0133] Processing unit 602, configured to perform model training based on the preset scene buried point information to obtain a feature extraction model and first information features corresponding to the preset scene buried point information;
[0134] The processing unit 602 is further configured to input the information to be trained into the feature extraction model to obtain a second information feature corresponding to the information to be trained;
[0135] The processing unit 602 is further configured to determine target training data corresponding to the to-be-trained data based on the similarity between the first information feature and the second information feature.
[0136] Optionally, when the preset scene buried point information further includes: a first clustering result for representing a data category of the preset scene buried point data, the first information feature includes a first clustering feature of the first clustering result; when the to-be-trained information further includes: a second clustering result for representing a data category of the to-be-trained data, the processing unit 602 is specifically configured to:
[0137] Clustering the preset scene buried point data according to the first clustering algorithm to obtain a first clustering result;
[0138] The training data is clustered according to the second clustering algorithm to obtain a second clustering result.
[0139] Optionally, the processing unit 602 is specifically configured to:
[0140] When the similarity between the first information feature and the second information feature is greater than or equal to a first preset similarity, determining the to-be-trained data as target training data;
[0141] When the similarity between the first information feature and the second information feature is less than or equal to the second preset similarity, the data to be trained is removed; the second preset similarity is less than the first preset similarity;
[0142] When the similarity between the first information feature and the second information feature is less than a first preset similarity and greater than a second preset similarity, interpolation processing is performed on the training data to obtain target training data.
[0143] Optionally, the processing unit 602 is further used to perform model training on the feature extraction model based on the target training data and the preset scene buried point information to obtain a target model; the target model is used to extract the preset scene buried point data from the data to be processed.
[0144] An embodiment of the present application further provides a computer-readable storage medium, which includes computer-executable instructions. When the computer-executable instructions are executed on a computer, the computer executes the data processing method provided in the above embodiment.
[0145] An embodiment of the present application also provides a computer program, which can be directly loaded into a memory and contains software code. After being loaded and executed by a computer, the computer program can implement the data processing method provided in the above embodiment.
[0146] Those skilled in the art will appreciate that, in one or more of the examples above, the functions described herein can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. Computer-readable media include computer-readable storage media and communication media, wherein communication media include any medium that facilitates the transmission of a computer program from one place to another. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0147] Through the description of the above implementation methods, technical personnel in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional modules is used as an example. In actual applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above.
[0148] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules or units is merely a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms. The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place, or they may be distributed in multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0149] In addition, each functional unit in each embodiment of the present application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit. If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or in other words, the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for making a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disk.
[0150] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A data processing method, characterized in that: include: Obtain preset scene tracking information and information to be trained; The preset scene burying point information includes preset scene burying point data; The information to be trained includes data to be trained; Performing model training based on the preset scene buried point information to obtain a feature extraction model and a first information feature corresponding to the preset scene buried point information; Inputting the information to be trained into the feature extraction model to obtain a second information feature corresponding to the information to be trained; Target training data corresponding to the to-be-trained data is determined based on the similarity between the first information feature and the second information feature.
2. The data processing method according to claim 1, wherein: When the preset scene buried point information further includes: a first clustering result for representing the data category of the preset scene buried point data, the first information feature includes a first clustering feature of the first clustering result; when the information to be trained further includes: a second clustering result for representing the data category of the data to be trained, the second information feature includes a second clustering feature of the second clustering result; the obtaining of the preset scene buried point information and the information to be trained includes: Clustering the preset scene buried point data according to a first clustering algorithm to obtain the first clustering result; The data to be trained is clustered according to a second clustering algorithm to obtain the second clustering result.
3. The data processing method according to claim 1, wherein: The determining, based on the similarity between the first information feature and the second information feature, target training data corresponding to the to-be-trained data includes: When the similarity between the first information feature and the second information feature is greater than or equal to a first preset similarity, determining the to-be-trained data as the target training data; When the similarity between the first information feature and the second information feature is less than or equal to a second preset similarity, removing the to-be-trained data; and the second preset similarity is less than the first preset similarity; When the similarity between the first information feature and the second information feature is less than the first preset similarity and greater than the second preset similarity, interpolation processing is performed on the data to be trained to obtain the target training data.
4. The data processing method according to any one of claims 1 to 3, characterized in that: Also includes: Performing model training on the feature extraction model based on the target training data and the preset scene embedding information to obtain a target model; The target model is used to extract preset scene buried point data from the data to be processed.
5. A data processing device, characterized in that: include: Acquisition unit and processing unit; The acquisition unit is used to acquire the preset scene buried point information and the information to be trained; The preset scene burying point information includes preset scene burying point data; The information to be trained includes data to be trained; The processing unit is configured to perform model training based on the preset scene buried point information to obtain a feature extraction model and a first information feature corresponding to the preset scene buried point information; The processing unit is further configured to input the information to be trained into the feature extraction model to obtain a second information feature corresponding to the information to be trained; The processing unit is further configured to determine target training data corresponding to the to-be-trained data based on the similarity between the first information feature and the second information feature.
6. The data processing device according to claim 5, characterized in that When the preset scene buried point information further includes: a first clustering result for representing the data category of the preset scene buried point data, the first information feature includes a first clustering feature of the first clustering result; when the to-be-trained information further includes: a second clustering result for representing the data category of the to-be-trained data, the processing unit is specifically configured to: Clustering the preset scene buried point data according to a first clustering algorithm to obtain the first clustering result; The data to be trained is clustered according to a second clustering algorithm to obtain the second clustering result.
7. The data processing device according to claim 5, characterized in that The processing unit is specifically configured to: When the similarity between the first information feature and the second information feature is greater than or equal to a first preset similarity, determining the to-be-trained data as the target training data; When the similarity between the first information feature and the second information feature is less than or equal to a second preset similarity, removing the to-be-trained data; and the second preset similarity is less than the first preset similarity; When the similarity between the first information feature and the second information feature is less than the first preset similarity and greater than the second preset similarity, interpolation processing is performed on the data to be trained to obtain the target training data.
8. The data processing device according to any one of claims 5 to 7, characterized in that: The processing unit is also used to perform model training on the feature extraction model based on the target training data and the preset scene buried point information to obtain a target model; the target model is used to extract the preset scene buried point data from the data to be processed.
9. A data processing device, characterized in that: It comprises a memory and a processor; the memory is used to store computer-executable instructions, and the processor is connected to the memory via a bus; when the data processing device is running, the processor executes the computer-executable instructions stored in the memory, so that the data processing device performs the data processing method according to any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium includes computer-executable instructions. When the computer-executable instructions are executed on a computer, the computer is enabled to execute the data processing method according to any one of claims 1 to 4.
Citation Information
Patent Citations
Data processing model training method and device, data processing method and device and electronic equipment
CN111401558A
Data screening method and device, storage medium and electronic equipment
CN111797288A