Computer data automatic labeling and classifying system based on deep learning
By adopting deep learning technology and adaptive multi-task learning algorithms in the data annotation and classification system, the problem of data preprocessing in the existing technology that cannot be dynamically optimized is solved, and efficient and accurate data processing and labeling is achieved, which is suitable for complex multi-task environments and large-scale data scenarios.
Patent Information
- Application Number
- CN202411847429.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-16
- Publication Date
- 2025-05-06
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The existing data annotation and classification technology has problems such as inadequate data preprocessing and inability to dynamically optimize task priority and resource allocation when dealing with complex multi-task environments and large-scale data, resulting in competition and conflict between tasks and reducing the overall performance of classification.
It adopts a computer data automatic annotation and classification system based on deep learning, including data preprocessing module, deep learning classification module and annotation module. The data preprocessing module improves data quality through cleaning, denoising and standardizing processing. The deep learning classification module uses deep neural networks and adaptive multi-task learning algorithms to classify, and the labeling module supports multi-category label allocation.
It improves the efficiency and accuracy of data processing, can adapt to a variety of data types and complex task scenarios, dynamically optimizes task priority and resource allocation, reduces competition and conflicts between tasks, and significantly improves the automation level and application effect of data annotation and classification.
Smart Images

Figure CN119939382A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer data processing, and in particular to a computer data automatic labeling and classification system based on deep learning. Background Art
[0002] With the development of big data technology and artificial intelligence, the automatic labeling and classification of computer data has become an important task in the field of data analysis and processing. The generation and accumulation of massive data provide valuable resources for data mining, machine learning and intelligent decision-making. However, how to classify and label data efficiently and accurately is the key to improving data utilization and mining potential value. The application of automated processing technology is particularly urgent in various data types such as images, texts, sensors, etc., especially in large-scale data scenarios. Manual labeling is inefficient and prone to errors, and traditional rule-driven methods are also difficult to adapt to the diversity and dynamic changes of complex data.
[0003] Existing data labeling and classification technologies have many shortcomings in practical applications. On the one hand, the data preprocessing process lacks systematicity, and the data processing often fails to fully clean the noise and standardize the format, resulting in uneven data quality, affecting the accuracy of subsequent classification and labeling results. On the other hand, traditional classification models are usually unable to efficiently handle complex multi-task environments and fail to dynamically optimize task priorities and resource allocation, resulting in competition and conflict between tasks, reducing the overall classification performance. In addition, the existing automatic labeling technology is relatively simple and lacks a flexible multi-label allocation mechanism. It is difficult to meet the diverse needs of data labeling in complex scenarios and cannot guarantee the comprehensiveness and accuracy of data labeling.
[0004] The purpose of the present invention is to overcome the shortcomings of the prior art and provide a computer data automatic labeling and classification system based on deep learning, which can not only adapt to a variety of data types and complex task scenarios, but also improve the efficiency and accuracy of data processing, and provide a flexible, intelligent and scalable solution for large-scale data analysis and processing. Summary of the invention
[0005] The present invention provides a computer data automatic labeling and classification system based on deep learning.
[0006] A computer data automatic labeling and classification system based on deep learning, including a data preprocessing module, a deep learning classification module and a labeling module, wherein;
[0007] The data preprocessing module receives computer data and performs preprocessing, including cleaning, denoising and standardization;
[0008] The deep learning classification module classifies the preprocessed computer data based on a deep neural network (DNN) model and optimizes the weights of the classification tasks through an adaptive multi-task learning algorithm;
[0009] The labeling module automatically labels computer data based on the output results of the deep learning classification module in combination with preset rules, and supports a multi-category label assignment mechanism.
[0010] Optionally, the data preprocessing module includes:
[0011] Data reception and import: receiving computer data transmitted by external units or devices, and importing computer data;
[0012] Data cleaning: use mean interpolation to fill missing values and remove duplicate data;
[0013] De-noising: Use median filtering to remove noise from computer data;
[0014] Data standardization: Standardize the input computer data and convert it into a standard normal distribution with zero mean and unit variance.
[0015] Optionally, the deep learning classification module includes:
[0016] Data conversion: convert the format of the preprocessed computer data, and determine and assign classification tasks based on the format-converted computer data. Each classification task receives the corresponding computer data as input for processing;
[0017] Data classification: classify input data through a deep neural network (DNN) model and generate classification results;
[0018] Adaptive multi-task learning optimization: Based on the adaptive multi-task learning algorithm, the weight distribution of classification tasks is dynamically adjusted to optimize the priority of classification tasks.
[0019] Optionally, the data conversion includes:
[0020] Format conversion and data structure adjustment: Format conversion of received pre-processed computer data, including dimension adjustment and batch processing;
[0021] Determine and assign classification tasks: Determine and assign classification tasks based on the converted computer data by calculating similarity;
[0022] Task assignment results: The format-converted computer data and the corresponding classification task assignment results are output simultaneously.
[0023] Optionally, the format conversion and data structure adjustment include:
[0024] Dimensionality adjustment: Dimensionality adjustment is performed on the preprocessed computer data, converting the computer data from its original high-dimensional form (image, text, numerical value) to a standardized form (one-dimensional vector);
[0025] Batch processing: Divide the format-converted computer data into multiple batches to improve data processing efficiency.
[0026] Optionally, the determining and assigning classification tasks includes:
[0027] Calculate similarity: For each format-converted computer data X', calculate the similarity between the data and each preset classification task feature T i The similarity between them determines the most relevant classification task;
[0028] Assign classification tasks: According to the similarity calculation results, automatically assign the classification task with the highest similarity to each format-converted computer data X', that is, select the task corresponding to the maximum similarity value Assign X' to the task
[0029] Optionally, the data classification includes:
[0030] Input data preparation: input the format-converted computer data X' into the deep neural network (DNN) model;
[0031] Forward propagation: The input data is forward propagated through the deep neural network (DNN) model. Assume that the deep neural network contains L layers, and the output of each layer is calculated by weighted sum, bias term, and activation function;
[0032] Calculate the classification output: In the last layer (output layer), the output Y of the deep neural network (DNN) model ′ The classification results are calculated through the Softmax function, which is used to convert the output into the probability distribution of each category;
[0033] Generate classification results: According to the Softmax output results, select the category with the highest probability as the final classification result y pred .
[0034] Optionally, the adaptive multi-task learning optimization includes:
[0035] Task weight initialization: At initialization, for each classification task T i Assigning initial weights Used to indicate the importance of each classification task;
[0036] Calculate task loss: For each classification task Ti , calculate the task loss L according to the difference between its corresponding classification result and the true label i ;
[0037] Calculate task loss gradient: Calculate the gradient of the loss of each task with respect to the weight
[0038] Dynamically adjust task weights: based on the loss gradient of each task Use adaptive algorithms to dynamically adjust task weights;
[0039] Task priority optimization: After each round of iteration, the priority of tasks is automatically optimized according to the update of task losses and weights.
[0040] Optionally, the annotation module includes:
[0041] Receive classification results: Receive the output results from the deep learning classification module, which include the classification probability or predicted category corresponding to each input data;
[0042] Matching preset labeling rules: According to the preset labeling rules, the classification results are matched and analyzed. The labeling rules include single-category labeling and multi-category label assignment rules. When the classification results meet the multi-category label assignment conditions, labels are assigned according to the classification probabilities of each category.
[0043] Automatic label assignment: Assign one or more labels to each data point based on the classification probability value;
[0044] Output the annotation results: bind the annotated computer data to the corresponding label set and output the final annotation results.
[0045] Optionally, the automatic label allocation includes:
[0046] Single-category label assignment: select the category C with the highest classification probability max As a single label;
[0047] Multi-category label assignment: When the classification probability exceeds the preset label assignment threshold θ, the category label is assigned to the input data.
[0048] Beneficial effects of the present invention:
[0049] The present invention, through the data preprocessing module, completes the reception, cleaning, denoising and standardization of computer data, ensuring the quality and consistency of input data. Data cleaning removes missing values and redundant data, denoising filters out noise in the data, and standardization converts the data into a standard normal distribution, effectively improving the stability and availability of the data. It can efficiently process data from different sources, provide high-quality input data for subsequent deep learning classification and automatic labeling, and ensure the reliability and efficiency of the data processing process.
[0050] The present invention realizes efficient classification of computer data based on a deep neural network model through a deep learning classification module, can automatically extract complex data features and complete nonlinear mapping. At the same time, data conversion ensures that the data adapts to the input requirements of the deep learning model through format conversion and classification task allocation, and accurately allocates data to the most relevant classification tasks according to similarity. The adaptive multi-task learning algorithm dynamically adjusts the weight distribution of classification tasks, optimizes task priorities, and ensures that high-priority tasks obtain more resources, thereby improving the overall classification accuracy and system performance, and meeting the efficient classification requirements in a multi-task environment.
[0051] The present invention realizes automatic label allocation of data through the labeling module, and has the flexibility of single-category and multi-category label allocation. Single-category label allocation ensures the uniqueness and reliability of label allocation by selecting the category label with the highest classification probability. Multi-category label allocation is based on a threshold mechanism. By statistically analyzing the mean and standard deviation of the classification probability, the label allocation threshold is adaptively set, and multiple relevant category labels can be allocated to the data, meeting the labeling requirements of complex scenarios, effectively reducing manual intervention, and improving the automation level and result quality of data labeling. It is suitable for large-scale and diversified data processing scenarios, and significantly improves the labeling efficiency and adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings in the following description are only for the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0053] Figure 1 A schematic diagram of system function modules according to an embodiment of the present invention;
[0054] Figure 2 Schematic diagram of a deep learning classification module according to an embodiment of the present invention. DETAILED DESCRIPTION
[0055] The present invention is described in detail below in conjunction with the accompanying drawings and specific embodiments. At the same time, it is explained here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments, and those skilled in the art may also adopt other alternatives to implement some known technologies; and the accompanying drawings are only for more specific description of the embodiments, and are not intended to specifically limit the present invention.
[0056] It should be noted that the references to "one embodiment", "an embodiment", "an exemplary embodiment", "some embodiments" and the like in the specification indicate that the embodiments described may include specific features, structures or characteristics, but not every embodiment may include the specific features, structures or characteristics. In addition, when a specific feature, structure or characteristic is described in conjunction with an embodiment, it should be within the knowledge of a person skilled in the art to implement such feature, structure or characteristic in conjunction with other embodiments (whether or not explicitly described).
[0057] In general, a term can be understood, at least in part, from its use in context. For example, depending, at least in part, on the context, the term "one or more" as used herein can be used to describe any feature, structure, or characteristic in the singular sense, or can be used to describe a combination of features, structures, or characteristics in the plural sense. Additionally, the term "based on" can be understood as not necessarily intended to convey an exclusive set of factors, but can instead, depending, at least in part, on the context, allow for the presence of other factors that are not necessarily explicitly described.
[0058] like Figure 1-Figure 2 As shown, the computer data automatic labeling and classification system based on deep learning includes a data preprocessing module, a deep learning classification module and a labeling module, wherein;
[0059] The data preprocessing module receives computer data and performs preprocessing, including cleaning, denoising and standardization;
[0060] The deep learning classification module classifies the preprocessed computer data based on the deep neural network (DNN) model and optimizes the weights of the classification tasks through an adaptive multi-task learning algorithm;
[0061] The labeling module automatically labels computer data based on the output results of the deep learning classification module and preset rules, and supports a multi-category label assignment mechanism;
[0062] Through the above content, computer data can be processed efficiently and automatic labeling and classification can be realized. The deep neural network model and adaptive multi-task learning algorithm are adopted. It not only has efficient data processing capabilities, but also can adapt to a variety of data types and task requirements, providing a flexible, accurate and scalable solution, which significantly improves the automation level and application effect of data labeling and classification.
[0063] The data preprocessing module includes:
[0064] Data reception and import: receiving computer data transmitted by external units or devices, and importing computer data;
[0065] Data cleaning: Use mean interpolation to fill missing values, remove duplicate data, and ensure the integrity and consistency of input data, expressed as:
[0066]
[0067] Among them, x filled is the missing value after filling, {x i} is the valid data of the feature, and n is the number of valid data;
[0068] Denoising: Use median filtering to remove noise from computer data, expressed as:
[0069] x filtered =median({x ′ i});
[0070] Among them, {x ′ i} is the surrounding neighborhood data, x filtered is the data value after denoising;
[0071] Data standardization: Standardize the input computer data and convert it into a standard normal distribution with zero mean and unit variance, expressed as:
[0072]
[0073] Among them, x norm is the standardized data value, x is the original data, μ is the mean of the original data, and σ is the standard deviation of the original data;
[0074] Through the above content, it is possible to ensure the consistency and high quality of computer data received from different sources. The cleaning step removes missing values and duplicate data, the denoising process effectively filters out the noise in the original data, and the standardization step converts the data into a standard normal distribution, thereby improving the stability and availability of the data.
[0075] The deep learning classification module includes:
[0076] Data conversion: Convert the format of the preprocessed computer data to ensure that it can be accepted by the deep neural network, and determine and assign classification tasks based on the format-converted computer data. Each classification task receives the corresponding computer data as input for processing;
[0077] Data classification: classify input data through a deep neural network (DNN) model and generate classification results;
[0078] Adaptive multi-task learning optimization: Based on the adaptive multi-task learning algorithm, the weight distribution of classification tasks is dynamically adjusted to optimize the priority of classification tasks;
[0079] Through the above content, the computer data processing capability and classification accuracy are improved. Data conversion ensures that the preprocessed data can adapt to the input requirements of the deep neural network and allocates appropriate input to each classification task, so that the data remains consistent when flowing between different tasks. Data classification utilizes the powerful feature learning ability of the deep neural network model to efficiently and accurately assign data to the correct category, and the adaptive multi-task learning optimization submodule dynamically adjusts the task weights to ensure that each classification task can obtain the optimal learning resources in a multi-task environment, thereby improving the overall classification effect.
[0080] Data conversion includes:
[0081] Format conversion and data structure adjustment: Format conversion of received pre-processed computer data, including dimension adjustment and batch processing, to make it suitable for the input requirements of deep neural networks;
[0082] Determine and assign classification tasks: Determine and assign classification tasks based on the converted computer data by calculating similarity;
[0083] Task assignment results: Output the format-converted computer data and the corresponding classification task assignment results simultaneously;
[0084] Through the above content, the efficiency and accuracy of data processing are improved. Format conversion and data structure adjustment ensure that the input data can fully adapt to the requirements of deep neural networks, so that the data can smoothly enter the subsequent deep learning processing flow. By calculating the similarity between data and task characteristics, the most suitable classification task is automatically determined and assigned, thereby ensuring that each task can process data that matches its characteristics, improving classification accuracy, and ensuring the consistency and efficiency of the entire classification process.
[0085] Format conversion and data structure adjustment include:
[0086] Dimensionality adjustment: Dimensionality adjustment is performed on the preprocessed computer data, converting the computer data from the original high-dimensional form (image, text, numerical value) to a standardized form (one-dimensional vector), expressed as:
[0087] X ′ =f(X);
[0088] Where X is the preprocessed computer data, f(X) is the function that converts the preprocessed computer data into a format suitable for neural network processing, and X ′ is the converted data;
[0089] Batch processing: Divide the format-converted computer data into multiple batches to improve data processing efficiency, expressed as:
[0090] B i ={x1,x2,...,x m},i∈{1,2,...,k};
[0091] Among them, B i is the i-th batch, x1,x2,...,x m are the 1st, 2nd, ..., mth data samples in the batch, k is the total number of batches, and m is the number of samples in each batch;
[0092] The format function f(X) specifically includes:
[0093] Dimension adjustment of image data: Image data is a two-dimensional matrix (height, width). Deep neural networks require input data to be a one-dimensional vector. For a color image of size H×W×C (H is height, W is width, and C is the number of color channels), it is converted into a one-dimensional vector using the following formula, expressed as:
[0094] X ′ = Flatten(X);
[0095] Among them, Flatten is to flatten the multi-dimensional tensor of input data into a one-dimensional vector;
[0096] Dimension adjustment of text data: Text data is converted into vector representation through word embedding. The original text is a word sequence of length L. Assuming that each word is represented as a d-dimensional word vector, the dimension adjustment of text data converts each word into its corresponding embedding vector through the following formula, expressed as:
[0097] X ′ =[e1,e2,…,e L ];
[0098] Among them, e1, e2, …, e L is the d-dimensional word embedding vector of the 1st, 2nd, ..., Lth word, and the final X ′ is a matrix of dimension L×d, representing the embedding representation of the entire text;
[0099] Dimensionality adjustment of numerical data: For numerical data (such as tabular data or sensor data), normalization is used to adjust it to a fixed-length vector, converting each data point xi Normalized to the range [0,1], expressed as:
[0100]
[0101] Among them, X is the original data set, x ′ i Normalized data to ensure that all data are on the same scale;
[0102] Through the above content, the efficiency and accuracy of data processing are improved. Dimension adjustment ensures that different types of data (such as images, text or numerical data) can be converted into the input form required by the neural network, so that the data can be effectively processed by the network. Through batch processing, the data is divided into multiple batches, which can accelerate the training process, reduce the consumption of computing resources, and improve the convergence speed of the model.
[0103] Identify and assign classification tasks including:
[0104] Calculate similarity: For each format-converted computer data X', calculate the similarity between the data and each preset classification task feature T i The similarity between them determines the most relevant classification task, which is expressed as:
[0105]
[0106] Among them, S(X',T i ) is the data X' and the classification task T i The similarity between them, X' is the data after format conversion, T i is the feature vector of the i-th preset classification task, ‖X'‖ and ‖T i ‖ are X' and T respectively i The norm of
[0107] Assign classification tasks: According to the similarity calculation results, automatically assign the classification task with the highest similarity to each format-converted computer data X', that is, select the task corresponding to the maximum similarity value Assign X' to the task
[0108] Through the above content, the efficiency and accuracy of the classification process are improved. The cosine similarity algorithm can accurately evaluate the relationship between data and tasks, ensuring that each data can be assigned the most matching classification task. It not only reduces manual intervention and improves the system's automation level, but also optimizes the allocation of classification tasks so that each task can focus on processing the most relevant data. Through this intelligent allocation method, the accuracy of classification tasks is improved and the processing speed is accelerated. At the same time, in a multi-task learning environment, the execution of different tasks can be effectively managed and coordinated.
[0109] Data categories include:
[0110] Input data preparation: input the format-converted computer data X' into the deep neural network (DNN) model;
[0111] Forward propagation: The input data is forward propagated through the deep neural network (DNN) model. Assume that the deep neural network contains L layers. The output of each layer is calculated by weighted sum, bias term and activation function. The output of the lth layer is H l It is expressed as:
[0112] H l =σ(W l H l-1 +b l );
[0113] Among them, W l is the weight matrix of the lth layer, H l-1 is the output of the l-1th layer, b l is the bias term, σ(·) is the ReLU activation function;
[0114] Calculate the classification output: In the last layer (output layer), the output Y of the deep neural network (DNN) model ′ The classification results are calculated by the Softmax function, which is used to convert the output into the probability distribution of each category, expressed as:
[0115]
[0116] Among them, P(y i |X ′ ) is the input data X ′ Belongs to category y i The probability of z i is the score of the i-th category, C is the number of categories, z j is the score of the jth category;
[0117] Generate classification results: According to the Softmax output results, select the category with the highest probability as the final classification result y pred , expressed as:
[0118] y pred =argmax(P(y i |X ′ ));
[0119] Among them, y pred is the final classification result, that is, the category to which the data predicted by the network belongs;
[0120] Through the above content, the input data can be accurately assigned to the most appropriate category. Through the learning of multi-layer networks, DNN can extract complex features from the data and perform nonlinear mapping, effectively improving the accuracy and robustness of classification. The Softmax function ensures the probabilistic nature of the output results, so that the prediction results of each category can be quantified as a probability value, thereby providing higher interpretability for the classification task. Finally, the system ensures the accuracy of the classification by selecting the category with the maximum probability as the prediction result. In addition, the deep learning model has the ability to learn automatically without human intervention, and can maintain high efficiency and high precision when processing large-scale data, thereby greatly improving the automation and execution efficiency of the classification task.
[0121] Adaptive multi-task learning optimization includes:
[0122] Task weight initialization: At initialization, for each classification task T i Assigning initial weights Used to indicate the importance of each classification task;
[0123] Calculate task loss: For each classification task T i , according to the difference between its corresponding classification result and the true label, calculate the task loss L i , for the i-th task, the loss function L i It is expressed as:
[0124]
[0125] in, is the loss function, y j is the true label of the jth data point, is the predicted value of the i-th task, and N is the number of samples;
[0126] Calculate task loss gradient: In order to optimize the weight distribution of tasks, calculate the gradient of the loss of each task relative to the weight The gradient represents the degree of influence of each task on the parameter update;
[0127] Dynamically adjust task weights: based on the loss gradient of each task Use an adaptive algorithm to dynamically adjust task weights, expressed as:
[0128]
[0129] in, is the weight of the i-th task at the t+1-th iteration, is the weight of the i-th task at the t-th iteration, η is the learning rate, is the gradient of the loss function with respect to the weight;
[0130] Task priority optimization: After each round of iteration, the priority of tasks is automatically optimized according to the update of task loss and weight. Tasks with small task loss and large weight will be given higher priority, thus affecting the parameter update during model training;
[0131] Through the above content, it is ensured that the system can intelligently optimize resource allocation according to the learning progress and difficulty of the task, improve the overall classification performance, and by calculating the gradient of the task loss and updating the weights, the system can automatically identify tasks that have a greater impact on the overall performance and give them a higher priority, while reducing attention to secondary tasks to avoid waste of resources. This not only improves the accuracy of classification tasks, but also can effectively coordinate competition and conflicts between tasks in a multi-task environment, accelerate the convergence of the model, and improve the robustness and processing efficiency of the system in complex task scenarios.
[0132] The annotation module includes:
[0133] Receive classification results: Receive the output results from the deep learning classification module, which include the classification probability or predicted category corresponding to each input data;
[0134] Matching preset labeling rules: According to the preset labeling rules, the classification results are matched and analyzed. The labeling rules include single-category labeling and multi-category label assignment rules. When the classification results meet the multi-category label assignment conditions, labels are assigned according to the classification probabilities of each category.
[0135] Automatic label assignment: Assign one or more labels to each data point based on the classification probability value;
[0136] Output the annotation results: bind the annotated computer data with the corresponding label set and output the final annotation results;
[0137] Through the above content, automatic label assignment of computer data is realized, with the flexibility of single-category and multi-category label assignment, and the ability to automatically select the most appropriate label based on the classification probability, assign a single label or multiple labels to the data, meet the data labeling needs in different scenarios, improve the efficiency and accuracy of data labeling, reduce manual intervention, and realize efficient and intelligent labeling of batch data, which is suitable for complex tasks and large-scale data processing environments.
[0138] Automatic label assignment includes:
[0139] Single-category label assignment: select the category C with the highest classification probability max As a single tag, it is represented as:
[0140] C max =argmax(P(y i |X ′));
[0141] Among them, P(y i |X ′ ) represents the input data X ′ Belongs to category y i probability;
[0142] Multi-category label assignment: When the classification probability exceeds the preset label assignment threshold θ, the category label is assigned to the input data, expressed as:
[0143] Assign(y i )={y i |P(y i |X ′ )≥θ};
[0144] Among them, θ is the label assignment threshold, Assign(y i ) represents the assigned category label set;
[0145] The setting of the label allocation threshold θ specifically includes:
[0146] Statistical classification probability distribution: Statistic all classification probabilities output by the deep learning classification module to obtain the probability distribution of all data points in different categories. Suppose the probability set of the classification result is P = {P(y1), P(y2), ..., P(y n )}, where P(y1), P(y2), ..., P(y n ) represents the probability of the 1st, 2nd, ..., nth category, where n is the total number of categories;
[0147] Calculate the mean and standard deviation of the probability distribution: Calculate the mean μ and standard deviation σ of all classification probabilities, expressed as:
[0148]
[0149] Among them, P(y i ) is the probability of the i-th category;
[0150] Set the label assignment threshold: Based on the statistical results, the label assignment threshold θ is set to the classification probability mean μ plus a weight factor k times the standard deviation σ, expressed as:
[0151] θ=μ+k·σ;
[0152] Among them, k is an adjustable weight factor. When k is small (k=0.5), the threshold is low and the label allocation is more relaxed, which is suitable for scenarios with high requirements for label coverage. When k is large (k=1.5), the threshold is high and the label allocation is more strict, which is suitable for scenarios with high requirements for label accuracy.
[0153] Through the above content, the flexible allocation of single-category and multi-category labels is realized, which significantly improves the accuracy and adaptability of data labeling. Single-category allocation ensures the uniqueness of data labels and the reliability of classification results by selecting the category label with the highest probability, while multi-category label allocation is based on the threshold judgment mechanism, which can assign multiple relevant category labels to data to meet the labeling requirements of complex data scenarios. The adaptive adjustment of the threshold further improves the accuracy and flexibility of the system in different tasks, making the labeling process accurate and efficient, reducing manual intervention, and improving the automation level and result quality of large-scale data labeling. It is suitable for multi-label classification tasks and diversified data processing scenarios.
[0154] The present invention covers any substitution, modification, equivalent method and scheme made on the essence and scope of the present invention. In order to make the public have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, but those skilled in the art can fully understand the present invention without the description of these details. In addition, in order to avoid unnecessary confusion about the essence of the present invention, well-known methods, processes, procedures, components and circuits are not described in detail.
[0155] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A computer data automatic labeling and classification system based on deep learning, characterized by: It includes data preprocessing module, deep learning classification module and labeling module, among which; The data preprocessing module receives computer data and performs preprocessing, including cleaning, denoising and standardization; The deep learning classification module classifies the preprocessed computer data based on a deep neural network model and optimizes the weights of the classification tasks through an adaptive multi-task learning algorithm; The labeling module automatically labels computer data based on the output results of the deep learning classification module in combination with preset rules, and supports a multi-category label assignment mechanism.
2. The computer data automatic labeling and classification system based on deep learning according to claim 1 is characterized in that: The data preprocessing module comprises: Data reception and import: receiving computer data transmitted by external units or devices, and importing computer data; Data cleaning: use mean interpolation to fill missing values and remove duplicate data; De-noising: Use median filtering to remove noise from computer data; Data standardization: Standardize the input computer data and convert it into a standard normal distribution with zero mean and unit variance.
3. The computer data automatic labeling and classification system based on deep learning according to claim 1 is characterized in that: The deep learning classification module includes: Data conversion: convert the format of the preprocessed computer data, and determine and assign classification tasks based on the format-converted computer data. Each classification task receives the corresponding computer data as input for processing; Data classification: classify input data through a deep neural network model and generate classification results; Adaptive multi-task learning optimization: Based on the adaptive multi-task learning algorithm, the weight distribution of classification tasks is dynamically adjusted to optimize the priority of classification tasks.
4. The computer data automatic labeling and classification system based on deep learning according to claim 3 is characterized in that: The data conversion includes: Format conversion and data structure adjustment: Format conversion of received pre-processed computer data, including dimension adjustment and batch processing; Determine and assign classification tasks: Determine and assign classification tasks based on the converted computer data by calculating similarity; Task assignment results: The format-converted computer data and the corresponding classification task assignment results are output simultaneously.
5. The computer data automatic labeling and classification system based on deep learning according to claim 4 is characterized in that: The format conversion and data structure adjustment include: Dimensionality adjustment: Dimensionality adjustment is performed on the preprocessed computer data, converting the computer data from its original high-dimensional form into a standardized form; Batch processing: Divide the format-converted computer data into multiple batches to improve data processing efficiency.
6. The computer data automatic labeling and classification system based on deep learning according to claim 5 is characterized in that: Determining and assigning classification tasks includes: Calculate similarity: For each format-converted computer data X', calculate the similarity between the data and each preset classification task feature T i The similarity between them determines the most relevant classification task; Assign classification tasks: According to the similarity calculation results, automatically assign the classification task with the highest similarity to each format-converted computer data X', that is, select the task corresponding to the maximum similarity value Assign X' to the task 7. The computer data automatic labeling and classification system based on deep learning according to claim 6 is characterized in that: The data categories include: Input data preparation: input the format-converted computer data X' into the deep neural network model; Forward propagation: The input data is forward propagated through the deep neural network model. Assume that the deep neural network contains L layers, and the output of each layer is calculated by weighted sum, bias term, and activation function; Calculate the classification output: In the last layer, the output Y of the deep neural network model ′ The classification results are calculated through the Softmax function, which is used to convert the output into the probability distribution of each category; Generate classification results: According to the Softmax output results, select the category with the highest probability as the final classification result y pred .
8. The computer data automatic labeling and classification system based on deep learning according to claim 7 is characterized in that: The adaptive multi-task learning optimization includes: Task weight initialization: At initialization, for each classification task T i Assigning initial weights Used to indicate the importance of each classification task; Calculate task loss: For each classification task T i , calculate the task loss L according to the difference between its corresponding classification result and the true label i ; Calculate task loss gradient: Calculate the gradient of the loss of each task with respect to the weight Dynamically adjust task weights: based on the loss gradient of each task Use adaptive algorithms to dynamically adjust task weights; Task priority optimization: After each round of iteration, the priority of tasks is automatically optimized according to the update of task losses and weights.
9. The computer data automatic labeling and classification system based on deep learning according to claim 1, characterized in that: The marking module comprises: Receive classification results: Receive the output results from the deep learning classification module, which include the classification probability or predicted category corresponding to each input data; Matching preset labeling rules: According to the preset labeling rules, the classification results are matched and analyzed. The labeling rules include single-category labeling and multi-category label assignment rules. When the classification results meet the multi-category label assignment conditions, labels are assigned according to the classification probabilities of each category. Automatic label assignment: Assign one or more labels to each data point based on the classification probability value; Output the annotation results: bind the annotated computer data to the corresponding label set and output the final annotation results.
10. The computer data automatic labeling and classification system based on deep learning according to claim 9, characterized in that: The automatic label allocation includes: Single-category label assignment: select the category C with the highest classification probability max As a single label; Multi-category label assignment: When the classification probability exceeds the preset label assignment threshold θ, the category label is assigned to the input data.
Citation Information
Cited By
Federal learning-oriented label perception distribution dynamic weighted backdoor defense method
CN122475940A