Data acquisition cleaning and labeling processing system and method based on AI technology
By introducing deep learning, variational method, information gain, quantum computing and other AI technologies into the data acquisition, cleaning and labeling processing system, the problems of low efficiency, low accuracy and poor flexibility in the data acquisition, cleaning and labeling processing in the existing technology are solved, and efficient and accurate data processing is achieved.
Patent Information
- Application Number
- CN202510195352.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-21
- Publication Date
- 2025-06-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing AI technologies have problems such as inefficiency, low accuracy and poor flexibility in data acquisition, cleaning and labeling. Especially when facing multi-source heterogeneous data, it is difficult to adaptively process the data sources that change, resulting in low data acquisition efficiency.
The data acquisition, cleaning and labeling processing system based on AI technology is adopted, including a data acquisition module, a data preprocessing module, a data cleaning module, a data labeling module, a data storage and query optimization module, and a feedback and optimization mechanism module. Through deep learning, variational method, information gain, quantum computing, matrix decomposition and reinforcement learning, adaptive data acquisition, precise cleaning and efficient labeling are achieved.
It significantly improves the efficiency and accuracy of data acquisition, cleaning and labeling, reduces manual intervention and labeling errors, improves the flexibility and accuracy of data processing, and solves the problem of storage and query bottlenecks in traditional methods.
Smart Images

Figure CN120144931A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of AI data processing, and particularly to a data acquisition, cleaning and annotation processing system and method based on AI technology. Background Art
[0002] With the advent of the big data era, artificial intelligence (AI) technology has been widely applied to data processing tasks in various industries. Especially in the fields of data acquisition, cleaning and annotation processing, the introduction of AI technology has made data processing more efficient and intelligent. Data acquisition, as the starting link of data processing, usually involves collecting a large amount of raw data from various data sources, which may come from multiple channels such as sensors, social networks, enterprise databases, etc. Subsequently, the collected raw data often contains noise, redundancy or missing values, so data cleaning is required to improve data quality. And data annotation is a key step to ensure that data can be correctly used for machine learning or other analysis tasks. Especially in supervised learning, the quality of annotated data directly determines the performance of the model.
[0003] In the existing AI technology for data acquisition, cleaning and annotation processing, in the data acquisition link, automated tools or crawler technology are used to collect data from multiple data sources. With the development of the Internet of Things and big data technology, the types of data sources are numerous and the data volume is huge. During the acquisition process, it is necessary to deal with the data format differences from different data sources. The collected data will enter the data cleaning stage, which includes steps such as removing noise, repairing missing data, and standardization processing. Common methods include rule-based cleaning and cleaning algorithms based on statistics or machine learning. In the data annotation stage, AI technology assigns labels to data points through automatic annotation, semi-automatic annotation or manual review.
[0004] However, although the existing data acquisition technology can automatically obtain data from different data sources, in the face of multi-source heterogeneous data, the non-uniform data format and the diversity of data sources are still a major challenge. The existing technology often relies on fixed rules for formatting processing and cannot adaptively process changing data sources, resulting in low data acquisition efficiency. In the data cleaning link, traditional cleaning methods often rely on manually formulated rules or simple statistical methods for noise removal and missing value filling, lacking flexibility and adaptability. These methods are prone to low efficiency when dealing with large-scale data. Especially in the face of complex data patterns, the performance of traditional algorithms is far inferior to that of intelligent methods based on deep learning. Therefore, the present invention provides a data acquisition, cleaning and annotation processing system and method based on AI technology to solve the deficiencies existing in the prior art. Summary of the Invention
[0005] In view of the deficiencies of the prior art, the present invention provides a data collection, cleaning, and annotation processing system and method based on AI technology, which solves the problems of low efficiency, low accuracy, and poor flexibility in data collection, cleaning, and annotation processing in existing AI technologies.
[0006] To achieve the above objectives, the present invention is realized through the following technical solutions: A data collection, cleaning, and annotation processing system based on AI technology, comprising:
[0007] A data collection module for automatically collecting raw data from multiple data sources;
[0008] A data preprocessing module for formatting, de-duplicating, and preliminarily cleaning the collected raw data;
[0009] A data cleaning module for optimizing the data cleaning process based on the variational method and high-order differential equations to remove noise and repair missing values;
[0010] A data annotation module for optimizing the data annotation process through information gain and calculation;
[0011] A data storage and query optimization module for optimizing the data storage structure through matrix decomposition; Therefore, the present invention provides an efficient cold forging forming die and forming process to solve the deficiencies existing in the prior art
[0012] A feedback and optimization mechanism module for automatically adjusting the data processing process according to the feedback of the cleaning and annotation results.
[0013] Preferably, the data collection module includes an automated data collection algorithm, which dynamically identifies and adapts to multiple data sources through deep learning technology, automatically collects data from structured and unstructured data sources, and converts the data into a standard format for subsequent processing.
[0014] Preferably, the data cleaning module includes a variational method optimization objective function, and the variational method optimization objective function is defined as:
[0015]
[0016] Wherein, is the loss function of data cleaning, x i is the raw data point, c i is the data point after cleaning, and n is the total number of data points.
[0017] Preferably, the data annotation module optimizes label assignment through information gain, and the information gain formula is:
[0018] IG(D,Y) = H(Y) - H(Y∣D);
[0019] Among them, IG(D, Y) represents the information gain, H(Y) is the label entropy, and H(Y|D) is the conditional entropy.
[0020] Preferably, the data storage and query optimization module decomposes the data set into a latent feature matrix W and a label matrix H through matrix factorization, which is defined as:
[0021] D α ≈WH;
[0022] Among them, D α is the original data matrix, W is the feature matrix, and H is the label matrix.
[0023] Preferably, the feedback and optimization mechanism module includes:
[0024] Feedback unit: used to feedback the optimization strategy according to the results after data cleaning and annotation;
[0025] Self-learning unit: used to dynamically adjust the cleaning and annotation process through reinforcement learning algorithms;
[0026] Adjustment unit: used to adjust the parameters of the data cleaning and annotation module.
[0027] Preferably, the data cleaning module includes:
[0028] Variational optimization unit: used to construct and solve the optimization problem of data cleaning;
[0029] Higher-order differential equation unit: used to model the acceleration in the data cleaning process;
[0030] Data repair unit: used to repair missing values according to the results optimized by the higher-order differential equation.
[0031] Preferably, the data acquisition module includes:
[0032] Data source identification unit: used to automatically identify and classify different types of data sources;
[0033] Intelligent acquisition unit: used to automatically acquire data from the identified data sources;
[0034] Data preprocessing unit: used to perform preliminary cleaning and standardization on the acquired data.
[0035] Preferably, the data annotation module includes:
[0036] Information gain unit: used to calculate the contribution of each data point to the overall annotation task;
[0037] Calculation unit: used to accelerate the label assignment process;
[0038] Annotation adjustment unit: used to optimize label assignment based on information gain and quantum computing results.
[0039] The present invention also provides a data collection, cleaning and annotation processing method based on AI technology, including the following steps:
[0040] Automatically collect raw data from multiple data sources;
[0041] Format, deduplicate and preliminarily clean the collected raw data;
[0042] Use variational methods and higher-order differential equations to clean the data, remove noise and repair missing values;
[0043] Optimize data annotation through information gain and quantum computing, and automatically assign labels to each data point;
[0044] Optimize the storage of the cleaned and annotated data, optimize the data storage structure through matrix factorization, and improve the data access efficiency through query optimization algorithms;
[0045] Automatically adjust the data processing process according to the feedback of the cleaning and annotation results, and optimize the data processing strategy.
[0046] The present invention provides a data collection, cleaning and annotation processing system and method based on AI technology. It has the following beneficial effects:
[0047] 1. The present invention adopts an AI-based adaptive data cleaning and annotation optimization system, combines information gain and quantum computing to accelerate the annotation process, and achieves an intelligent and efficient data annotation effect; compared with the traditional manual annotation or rule-based annotation methods in the prior art, the present invention significantly reduces manual intervention and annotation errors through an automated optimization strategy, while improving the annotation accuracy and speed, and solves the problems of inefficiency and errors existing in manual annotation.
[0048] 2. The present invention comprehensively optimizes data storage and query by introducing matrix factorization and database index optimization technologies, and achieves the effect of maximizing the utilization of storage space and improving the query response speed; compared with the common redundant storage and inefficient query methods in the prior art, the present invention can significantly reduce data redundancy, improve data storage and retrieval efficiency, and solve the problems of storage and query bottlenecks in traditional methods.
[0049] 3. The present invention adopts deep learning technology and self-learning mechanism to realize dynamic adjustment and optimization of each link in the data processing process; compared with the limitations of fixed algorithms and strategies in the prior art, the present invention makes adaptive adjustments according to the real-time feedback of data through intelligent algorithms such as reinforcement learning, thus ensuring the flexibility of system processing and the accuracy of data, and solving the problem that traditional methods cannot dynamically respond to different data characteristics.
[0050] 4. The present invention combines a high-order differential equation and the variational method to optimize the data cleaning process, effectively removing noise and repairing missing values, achieving a high-precision data cleaning effect. Compared with the relatively simple cleaning algorithms in the prior art, the present invention optimizes the cleaning process by introducing a high-order differential equation, making the data processing process more precise, and solving the deficiencies of ignoring data changes and incomplete noise removal in the existing methods. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 is the system architecture diagram of the present invention;
[0052] Figure 2 is the architecture diagram of the feedback and optimization mechanism module of the present invention;
[0053] Figure 3 is the architecture diagram of the data cleaning module of the present invention;
[0054] Figure 4 is the architecture diagram of the data acquisition module of the present invention;
[0055] Figure 5 is the architecture diagram of the data annotation module of the present invention;
[0056] Figure 6 is the method flow chart of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0057] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0058] Please refer to the attached Figure 1 - attached Figure 5 , the embodiments of the present invention provide a data acquisition, cleaning and annotation processing system based on AI technology. By combining advanced technologies such as deep learning, information gain, and quantum computing, it includes a data acquisition module, a data preprocessing module, a data cleaning module, a data annotation module, a data storage and query optimization module, and a feedback and optimization mechanism module, realizing fully automated processing from data acquisition to annotation, thereby significantly improving the accuracy and efficiency of data processing and solving the problems of manual intervention and low processing efficiency in traditional methods.
[0059] For the data acquisition module, in this embodiment, through intelligent algorithms, it can efficiently obtain data from different forms and types of data sources (such as structured data sources, unstructured data sources, sensor data, etc.), and perform necessary formatting processing to provide a high-quality data basis for subsequent cleaning, annotation, and storage operations. The core goal of the data acquisition module is to ensure the integrity, accuracy, and consistency of the data.
[0060] Taking advantage of deep learning technology, this module can dynamically identify different types of data sources. Specifically, the module can adapt to data of different formats and sources during the acquisition process and perform automated standardization conversion on the data. Through neural network algorithms (such as convolutional neural network CNN, recurrent neural network RNN, etc.), the data acquisition module can process various data types, including but not limited to text data, image data, video data, and sensor data, etc.
[0061] Generally, the acquired data often has different formats and structures, which may include but not limited to text data in JSON or XML format, image data in two-dimensional or three-dimensional structures, etc. For these different types of data, the data acquisition module performs conversion through automated algorithms to unify the data format and provide a consistent data structure for subsequent processing. In some embodiments, the data acquisition module includes the following two sub-modules: a data source identification sub-module and an intelligent acquisition sub-module. The role of the data source identification sub-module is to intelligently identify and classify the data source types according to the characteristics of the data source; while the intelligent acquisition sub-module is responsible for automatically acquiring data from the identified data sources.
[0062] In this embodiment, the data acquisition module uses pre-trained models in deep learning to identify and classify data sources. For example, convolutional neural network (CNN) is used for automatic identification and processing of image data, and recurrent neural network (RNN) is suitable for processing time series data (such as sensor data). These deep learning algorithms can adaptively adjust the acquisition method according to different data source types to ensure the efficient acquisition and accurate conversion of each data type.
[0063] In a possible implementation, the data acquisition module will first identify the type of data source through a deep learning model, such as distinguishing image, text, or time series data. The identification result will be used as input to indicate the subsequent acquisition steps to use appropriate acquisition strategies. For example, for image data, the data acquisition module may use object detection algorithms (such as YOLO, FasterR-CNN, etc.) to identify specific objects in the image. For text data, a pre-trained BERT model may be used for text classification and information extraction. Through these deep learning models, the data acquisition module can dynamically adjust the acquisition strategy when facing a large amount of diverse data sources to improve the acquisition efficiency.
[0064] The function of the data conversion unit is to convert the collected data into a unified standard format. Different data sources often have different format requirements. The data conversion unit applies formatting rules to convert the collected data into the standard format required for subsequent processing. For example, text data can be converted into JSON format or CSV format, and image data can be converted into a standard image format with a unified size and pixel value range. The data conversion unit can not only ensure data consistency but also perform data preprocessing according to the requirements of subsequent modules. Specifically, the data conversion process may include the following operations:
[0065] For text data, natural language processing (NLP) techniques are used for preprocessing, such as word segmentation, stop word removal, key feature extraction, etc. After conversion, the text data will be stored and processed in a unified format.
[0066] For image data, the data conversion unit may perform standardization processing on the image, such as resizing the image, grayscale processing, feature extraction, etc., so that the subsequent cleaning and annotation modules can process it correctly.
[0067] For sensor data or other time-series data, the data conversion unit will perform data standardization processing according to the acquisition frequency and data format requirements.
[0068] In some embodiments, the data conversion unit can also selectively compress or enhance the data according to the requirements of subsequent processing modules to improve the efficiency or accuracy of subsequent processing.
[0069] When performing the data acquisition task, the data acquisition module can not only identify and adapt to data sources but also continuously optimize the acquisition strategy through techniques such as reinforcement learning. For example, when data loss or errors occur in a certain data source in the system, the data acquisition module can automatically adjust the acquisition strategy through a feedback mechanism to reduce errors in future acquisitions. This adaptive optimization ability enables the system to maintain efficient and high-quality data acquisition in the face of constantly changing and diverse data sources.
[0070] For the data preprocessing module, in this embodiment, it ensures that the collected raw data can be converted into a unified format, eliminates inconsistencies, and provides a more reliable data source for the data cleaning module. Through this process, the system can significantly reduce the noise in the data and eliminate duplicate data, laying a solid foundation for subsequent cleaning and annotation tasks.
[0071] In this embodiment, the data preprocessing module first formats the collected raw data. Specifically, the raw data usually comes from multiple data sources with different formats, which may include structured data (such as database tables), semi-structured data (such as JSON and XML files), and unstructured data (such as images and video files). The data preprocessing module unifies these diverse data sources into a standard format by applying different conversion rules for subsequent processing and analysis.
[0072] For example, data formatting refers to converting the raw data from different data source formats into a standard format. In this process, the data preprocessing module selects different conversion methods according to the type of data (such as text, image, time-series data, etc.). Text data usually needs to be tokenized, stop words removed, etc., while image data may need to adjust image size, color depth, etc. For example, assume the original data set D = {x 1 ,x 2 ,x n} contains data in different formats. The goal of data formatting is to convert all elements in the data set into a standard format. For text data, the data formatting process may include the following steps: splitting the text into words or phrases; removing meaningless words (such as "de", "le", "shi"); extracting significant words from the text; data formatting usually involves unit conversion, such as unifying different measurement units (such as converting feet to meters, temperature units to Celsius).
[0073] Data deduplication refers to detecting and deleting duplicate records in the data set. During data collection, especially when data sources are inconsistent, the data set often contains multiple duplicate records. The purpose of deduplication is to reduce redundancy, save storage space, and improve the efficiency of subsequent processing. The deduplication process usually depends on the unique identifier of the data. Assume the data set D contains multiple data points x 1 ,x 2 ,x n . The deduplicated data set D ′ contains the data points with duplicates removed. Specifically, data deduplication can be represented by the following formula:
[0074]
[0075] where D is the original data set, D ′ is the deduplicated data set, x i and x j are data points, \ represents removing duplicate elements from the data set D, i is the index of the data points in the data set D, and j represents the position or serial number of a certain data point in the data set D.
[0076] For the data cleaning module, in this embodiment, it is responsible for further processing the collected raw data. By removing noise, repairing missing values, and removing redundant data, it ensures that the cleaned data can meet the high-quality requirements of subsequent annotation, storage, and analysis. Data cleaning is not only a key link in the entire data processing flow but also the basis for improving data accuracy and consistency. Specifically, this module uses the variational method and higher-order differential equations to optimize and clean the data, effectively removing noise and repairing missing values through a mathematical model to ensure data integrity and high quality.
[0077] In this embodiment, the data cleaning module optimizes the data through the variational method and higher-order differential equations. The core of this module is to construct an objective function for data cleaning through the variational method, and use this objective function to minimize the difference between the data and the cleaned data. First, let the original data set D = {x 1 , x 2 , x n} be a data set containing noise and missing values, and the data set obtained after data cleaning is C(D). The goal is to minimize the loss function in the data cleaning process:
[0078]
[0079] Among them, is the data cleaning loss function, which represents the difference between the original data point x i and the cleaned data point c i . x i is the original data point, which may contain noise or missing values, and c i is the cleaned data point, representing the optimal data after optimization. n is the total number of data points.
[0080] Generally, the goal of data cleaning is to minimize noise and inconsistency as much as possible while retaining as much valid information of the data as possible. By optimizing the above loss function, the data cleaning process can be made more precise and unnecessary data loss can be avoided.
[0081] In this embodiment, the cleaning module introduces higher-order differential equations to model the change rate and acceleration in the data cleaning process. By introducing the second derivative term, the change behavior of the data in the data cleaning process can be described more accurately, avoiding information loss caused by over-smoothing. Specifically, an acceleration term of the data change is added to the optimization objective function of the cleaning module:
[0082]
[0083] Among them, J(C(D)) represents the optimization objective function of data cleaning, is the loss function for data cleaning. α is a tuning factor that controls the balance between smoothness and noise removal during the cleaning process. represents the second derivative of the loss function, which indicates the acceleration of data change and can effectively capture the dynamic change characteristics during the data cleaning process. T is the time or the termination point of data processing, and dt is the small change or increment of the time variable t.
[0084] In one possible implementation, the data cleaning module models the dynamic characteristics of noise removal and missing value repair during the cleaning process through a high-order differential equation. This approach can make the data cleaning process more accurate and avoid cleaning failures or information loss caused by simple averaging or ignoring the acceleration of data change.
[0085] Specifically, to solve the above optimization problem, the data cleaning module uses the L-BFGS (Limited-memory Broyden-Fletcher-Goldfarb-Shanno) algorithm for numerical optimization. L-BFGS is a commonly used optimization algorithm that finds the parameter configuration that minimizes the loss function by iteratively updating the parameters. In actual operation, the data cleaning module minimizes the difference between the cleaned data C(D) and the original data D to achieve data denoising and repair.
[0086] During this process, the L-BFGS algorithm optimizes the cleaning process by calculating the gradient of the loss function and dynamically adjusts the balance factor α to ensure the optimal cleaning effect. The advantage of this method is that it can quickly find the optimal solution through efficient numerical calculations and avoid overfitting or underfitting problems that may exist in traditional cleaning methods.
[0087] According to the characteristics of different datasets, adjust the parameters in the model. Factors such as the structure of the dataset, the noise level of the data, and the distribution of missing values will all affect the data cleaning effect. In some cases, a higher smoothness may be required during the cleaning process to remove noise, while in other cases, excessive smoothness may lead to the loss of valid data. By dynamically adjusting α, the smoothness of the cleaning process can be flexibly controlled according to the characteristics and requirements of different data, thereby obtaining the best cleaning effect.
[0088] For example, use model-based methods (such as K-nearest neighbor method, regression imputation, etc.) to repair missing values. In this case, the data cleaning module first identifies the missing values and then uses machine learning models (such as regression analysis or deep learning methods) to predict the values of the missing data, thereby achieving the effect of repairing the missing data. Compared with traditional data repair methods, this model-based repair method can provide more accurate and reliable filling results.
[0089] For the data annotation module, in this embodiment, by combining information gain and quantum computing, a more intelligent and automated solution is provided. Through the intelligent annotation of data points, the data annotation module not only improves the accuracy of annotation but also significantly speeds up the processing speed. This module is closely connected to the aforementioned data collection and data cleaning modules, ensuring the efficiency and high quality of the data chain from collection to final annotation.
[0090] Specifically, the data annotation module first uses Information Gain (IG) in information theory to evaluate the contribution of each data point to the overall annotation task, ensuring that data points with large amounts of information are preferentially annotated during the annotation process. The formula for calculating information gain is as follows:
[0091] IG(D,Y) = H(Y) - H(Y∣D);
[0092] Among them, IG(D,Y) represents the information gain, that is, reducing the uncertainty of label Y by annotating data D. H(Y) is the entropy of label Y, representing the randomness of the label, and H(Y|D) is the conditional entropy, representing the uncertainty of label Y given the known data D.
[0093] Generally, the larger the information gain, the greater the amount of information of the data point in the annotation task. Therefore, preferentially annotating these data points can improve the annotation efficiency. In some embodiments, the data annotation module calculates the label values of each data point and selects the part with the largest information gain for annotation, thus ensuring the quality and efficiency of annotation.
[0094] To further improve the annotation efficiency, the data annotation module uses quantum computing technology to accelerate the label assignment. Quantum computing can utilize quantum superposition and quantum parallelism to perform parallel search and optimization in a high-dimensional label space. Through the Quantum Annealing algorithm or the Quantum Approximate Optimization Algorithm (QAOA), quantum computing can quickly find the optimal configuration from a large number of possible label combinations.
[0095] In a possible implementation, the data annotation module represents the label selection of each data point through qubits in quantum computing. Quantum computing utilizes the superposition and interference principles of qubits to achieve parallel computing in the label space, greatly improving the annotation speed and computing efficiency. Specifically, quantum computing avoids the local optimum problems that may be encountered in traditional optimization methods by searching for the global optimum solution of the label.
[0096] In this embodiment, the quantum annealing algorithm is applied to solve the label optimization problem. The quantum annealing algorithm gradually finds the global lowest energy state of the system by simulating the spontaneous cooling process of qubits at low temperatures. In the data annotation task, this process is equivalent to finding the optimal label assignment in the label space. The quantum annealing algorithm can effectively avoid falling into local optima and ensure that the label assignment for each data point is globally optimal.
[0097] In actual operation, the quantum annealing algorithm optimizes the objective function in the annotation process by setting appropriate initial conditions and energy functions. Specifically, the energy function can be designed based on factors such as information gain and label consistency to ensure that the optimization process of quantum computing can proceed efficiently according to the actual characteristics of the data.
[0098] The optimization of the data annotation module not only depends on the initial label assignment but also makes dynamic adjustments according to subsequent annotation results. In practical applications, the data annotation module can receive feedback information from the data cleaning module and the annotation process, adjust the annotation strategy and optimization algorithm, and further improve the accuracy and efficiency of annotation.
[0099] For example, during the data cleaning process, the annotation module may pre-annotate some data points and then dynamically adjust the annotation scheme according to the data results after cleaning. This feedback mechanism makes the data annotation process more flexible and intelligent and can adapt to different characteristics and changes of data sources.
[0100] For the data storage and query optimization module, in this embodiment, it is responsible for optimizing the storage of the processed high-quality data and achieving efficient data access in subsequent queries. Through the optimization of data storage, the system can significantly improve the query response speed and reduce unnecessary storage redundancy, thereby providing a more efficient storage solution for large-scale data processing. This module is closely connected to the aforementioned acquisition, cleaning, and annotation modules to ensure the smoothness and efficiency of the entire data processing process.
[0101] In this embodiment, the data storage and query optimization module optimizes the data storage structure through matrix factorization technology. Specifically, this module uses matrix factorization technology to perform a low-dimensional representation of the data. Assuming that the dataset D is a high-dimensional matrix composed of multiple data points and features, traditional storage methods may lead to redundancy and inefficient storage. In this case, matrix factorization can transform the high-dimensional dataset into a low-dimensional representation, thereby reducing the storage space and improving the calculation efficiency. The matrix factorization method decomposes the dataset D into two low-dimensional matrices W and H, defined as:
[0102] D α ≈WH;
[0103] where D αLet the original data matrix be \(X\), representing the entire data set, \(W\) be the feature matrix, containing the latent features of the data, and \(H\) be the label matrix, containing the label information of the data points.
[0104] It reduces the space required for storing high-dimensional data and improves the storage efficiency of data. For example, when dealing with large-scale image data, through matrix factorization, the high-dimensional features of each image can be represented as lower-dimensional feature vectors, thus greatly reducing the storage space requirements. For a large amount of time-series data or sensor data, potential patterns can be extracted through matrix factorization, reducing unnecessary redundant information.
[0105] As an option, the data storage module can combine database index optimization strategies to further improve the data query efficiency. In this embodiment, the data storage and query optimization module accelerates the data query process through various indexing techniques. Common database indexes include hash indexes and B+ tree indexes, which can significantly improve the data search speed. In some embodiments, the system will select the most suitable index method according to the data characteristics to optimize the query efficiency. For example, for large-scale structured data sets, hash indexes can quickly locate the target data, while for complex relational data, B+ tree indexes are more suitable for range queries.
[0106] Specifically, hash indexes map data to different buckets through hash functions, enabling direct location of the target data during query and avoiding full table scans. For relational data, B+ tree indexes can achieve efficient search operations through tree structures, especially suitable for handling range queries and sorting operations. By reasonably configuring indexes, the query time is greatly shortened, especially when dealing with large-scale data, the effect is particularly significant.
[0107] In a possible implementation, the data storage and query optimization module also combines data compression techniques. Data compression techniques are used to reduce the space required for storage. By performing lossless or lossy compression on the data, it is possible to save storage space without affecting the data accuracy. Especially in the case of dealing with large amounts of data such as images and videos, data compression can significantly reduce the storage cost and improve the system performance.
[0108] Common data compression methods include Huffman coding, LZW compression algorithm, etc. These methods can reduce the storage space occupancy by encoding the repeated patterns in the data. For some high-dimensional data, compression techniques can also be combined with matrix factorization techniques to further reduce the storage space while maintaining the data representation ability.
[0109] Specifically, the query optimization of the data storage module adopts an adaptive query algorithm. The query optimization algorithm plays an important role in this module. The system analyzes the characteristics of query requests and dynamically adjusts the query execution strategy to improve query efficiency. For example, when the query data range is small, the system will choose a full table scan to avoid unnecessary index overhead. For large-scale data queries, the system will give priority to using indexes to quickly locate the target data.
[0110] In some embodiments, the data storage module also uses a lazy query mechanism, which starts the query operation only when the user explicitly requests it. This avoids frequent access to data and reduces the impact of query operations on system performance.
[0111] As an option, the data storage and query optimization module can dynamically adjust the storage strategy according to the user access pattern. For example, if some data sets are frequently accessed, the system can preload this data into memory to avoid repeated disk read and write operations, thereby accelerating the query response speed. In addition, the system can also predict future query requests based on the user's historical query behavior and prepare the data that may be needed in advance to further improve query efficiency.
[0112] For the feedback and optimization mechanism module, in this embodiment, through a real-time feedback and self-learning mechanism, it is ensured that the system dynamically adjusts the processing strategy according to the results of data cleaning, annotation, storage and other links. Through the implementation of the adaptive mechanism, this module can optimize the subsequent processing process according to the characteristics and changes of different data sets, ensure the high efficiency of the system and the optimization of data processing. The close connection with the foregoing modules enables the system to perform cyclic optimization during the processing process, thereby improving the overall processing efficiency and data quality.
[0113] In this embodiment, the feedback and optimization mechanism module includes a feedback unit, a self-learning unit and an adjustment unit. The core of the feedback and optimization mechanism module is its feedback loop. Among them, the feedback unit is responsible for receiving the result feedback from each processing module. Specifically, the feedback unit obtains the processing results from the data cleaning, annotation and storage modules, evaluates the accuracy and efficiency of these results, and then transmits the feedback information to the self-learning unit to further promote system optimization.
[0114] Generally, the role of the feedback unit is to monitor the difference between the output and input of each link. Especially during the data cleaning and annotation process, the system will continuously monitor the cleaning effect and the accuracy of the annotation results. Once it is found that there are abnormalities or a decrease in efficiency in data processing, the feedback unit will automatically feedback this information to the self-learning unit to prompt the system to optimize.
[0115] Specifically, the feedback unit calculates the feedback error through the following loss function:
[0116]
[0117] Among them, g(D out , D true ) is the error function of the calculation result, and D out represents the output data set of the model or system, and D true represents the real data set or the target data set, measuring the difference between the processing result and the real data. D out,i is the data output by the module, the i-th data point, and D true,i is the real target data, the i-th data point, ||D out,i - D true,i || 2 represents the square of the error of the i-th data point, and n represents the total number of data points in the data set.
[0118] As an option, the output of the data cleaning and annotation module will be adjusted according to the feedback information. Specifically, the data cleaning module and the annotation module may adjust the algorithm parameters of the cleaning process or the annotation strategy according to the feedback signal. For example, when the system finds that there are errors in the annotated data, the feedback unit will identify the problems in the annotation and feedback them to the annotation module, so as to adjust the annotation strategy and optimize the label assignment rules. In this way, the system can continuously improve the quality of the annotation and cleaning results.
[0119] In a possible implementation, the feedback unit and the self-learning unit cooperate closely. Using the reinforcement learning algorithm, the self-learning unit adopts the reinforcement learning algorithm to continuously learn and adjust the system's parameters and processing strategies from the feedback. The reinforcement learning algorithm optimizes the data processing flow through the reward mechanism and the punishment mechanism. When the output of the system meets the expected result, a reward is given; if the result deviates from the target, a punishment is imposed; this learning mechanism based on rewards and punishments helps the system to gradually improve its processing strategy in actual operation.
[0120] Specifically, assume that in some data cleaning tasks, the parameters initially used by the system may not fully adapt to the characteristics of the data. The self-learning unit will adjust the parameters used in the cleaning process according to the feedback of the processing result, such as adjusting the smoothness factor or the noise removal intensity, so that the system can adopt a more appropriate parameter setting in the next round of data cleaning.
[0121] Specifically, the self-learning unit optimizes and adjusts the system based on reinforcement learning. In the self-learning unit, the core goal of reinforcement learning is to learn the optimal processing strategy according to the historical feedback information. The reinforcement learning algorithm accumulates experience from each feedback and continuously improves the strategy through trial and error. Specifically, the self-learning unit uses the Q-learning algorithm to update the processing strategy, where the Q-value function is used to evaluate the value of each operation. The specific formula is as follows:
[0122]
[0123] where Q(s t , a t ) represents the expected reward for executing action a t in state s t , r t+1 is the reward, representing the feedback effect after the current operation, γ is the discount factor, controlling the influence of future rewards, and α is the learning rate, representing the step size of each update. is the maximum Q-value among all possible actions a t+1 in the next state s ′ , s t represents the state at time step t, a t represents the action selected at time step t, and a ′ represents the action selected in state s t+1 .
[0124] For example, the optimization in the data annotation process can also be adjusted in this way. When the annotation module is annotating, if it is found that the label assignment efficiency of some data points is low, or the annotation result fails to meet the expected effect, the feedback unit will immediately provide relevant feedback information. Based on these feedback results, the self-learning unit adjusts the label assignment strategy or optimizes the calculation method of information gain. In this way, the system can gradually improve the accuracy and efficiency of annotation according to the feedback of each annotation process.
[0125] Specifically, the self-learning unit gradually optimizes the parameters of the cleaning and annotation processes through a reinforcement learning model. Based on reinforcement learning, the self-learning unit can flexibly adjust the cleaning and annotation strategies according to different data inputs and feedback. Whenever the goal in the processing process fails to meet the expectation, the self-learning unit will make adjustments based on historical data and feedback information, exploring the optimal strategy. In this process, the system gradually accumulates experience and optimizes the parameter configuration to ensure the efficiency of the data cleaning and annotation processes.
[0126] In some embodiments, reinforcement learning can be implemented through algorithms such as Q-learning or DeepQ-Network (DQN). In these algorithms, the system evaluates the value of each action (such as adjusting cleaning parameters or selecting a label assignment strategy) and gradually improves the accuracy of the selected action based on historical feedback. This self-learning ability enables the system to automatically adapt and optimize the processing process when facing new data sources.
[0127] The adjustment unit is responsible for adjusting the parameters of the data cleaning and annotation module in real time according to the adjustment strategy of the self-learning unit. The adjustment unit is an important part of the entire feedback and optimization mechanism module. Its role is to update the parameters of the data cleaning and annotation module in real time after the self-learning unit adjusts the cleaning and annotation strategies based on the feedback. Specifically, when the self-learning unit recommends new parameters based on the reinforcement learning algorithm, the adjustment unit will directly adjust the parameter settings of the data processing module according to these recommendations, making the data cleaning and annotation process more efficient.
[0128] For example, during the data cleaning process, if the self-learning unit discovers that a certain data set has a large amount of noise and a higher smoothing factor is required to denoise during the cleaning process, the adjustment unit will immediately update the parameters of the data cleaning module and adjust the smoothing factor to a more appropriate value to improve the cleaning effect.
[0129] Please refer to the appendix Figure 6 The present invention also provides a method for data acquisition, cleaning and annotation processing based on AI technology, including the following steps:
[0130] S1. Automatically collect raw data from multiple data sources;
[0131] S2. Format, deduplicate and preliminarily clean the collected raw data;
[0132] S3. Use variational methods and higher-order differential equations to clean the data, remove noise and repair missing values;
[0133] S4. Optimize data annotation through information gain and quantum computing, and automatically assign labels to each data point;
[0134] S5. Optimize the storage of the cleaned and annotated data, optimize the data storage structure through matrix factorization, and improve the data access efficiency through query optimization algorithms;
[0135] S6. Automatically adjust the data processing process according to the feedback of the cleaning and annotation results, and optimize the data processing strategy.
[0136] For step S1, the system will use technologies such as deep learning and natural language processing (NLP) to automatically identify different types of data sources (structured, unstructured, sensor data, etc.) and perform data collection. The collected data may come from different forms of data sources such as text, images, audio, and video. Therefore, the system must have a high degree of flexibility and adaptability.
[0137] For example, for text data, pre-trained language models (such as BERT, GPT, etc.) may be used for content extraction and formatting. For image or video data, convolutional neural networks (CNNs) may be used for feature recognition and data extraction.
[0138] For step S2, the collected data usually comes from different sources and may have inconsistent formats. Therefore, it is necessary to standardize the format for processing. This step ensures that subsequent processing will not be affected by data format differences. For example, text data may need to be tokenized and stop words removed, and image data may need to be resized and color space adjusted.
[0139] Duplicate removal is an important step to ensure the cleanliness of the dataset. Usually, duplicate items are identified and removed by detecting the unique identifiers of data records. The data after duplicate removal can provide a more reliable basis for subsequent analysis.
[0140] For step S3, by constructing a loss function, the difference between the data and the cleaned data is minimized to remove the noise components in the data. The variational method is usually used for global optimization of the data to ensure that the cleaning process can maximize the retention of useful information.
[0141] By modeling the acceleration and change rate of the data during the cleaning process, the cleaning process is further refined. For example, in the cleaning of time series data, introducing differential equations can better capture the data change trend and reduce the information loss caused by smoothing.
[0142] For step S4, based on information theory, the importance of each data point for the entire annotation task is evaluated by calculating the label entropy and conditional entropy. Data points with larger information gain will be preferentially annotated to ensure that the annotation process can maximize the reduction of label uncertainty, thereby improving the accuracy of annotation.
[0143] For step S5, decomposing the high-dimensional dataset into low-dimensional matrix representations (such as feature matrices and label matrices) can not only reduce the storage space but also improve the calculation efficiency. Matrix decomposition methods are usually used for data dimensionality reduction, such as using techniques like principal component analysis (PCA) or singular value decomposition (SVD).
[0144] To improve the query efficiency, the system accelerates data retrieval through various indexing techniques (such as hash indexing, B+ tree indexing, etc.). According to the data access pattern, the query optimization algorithm will dynamically adjust the storage structure to achieve the best query performance.
[0145] For step S6, feedback is generated in each link of the cleaning and annotation processes, and the system will continuously optimize the processing strategy based on this feedback. For example, when the annotation result does not meet the expectations, the system can automatically adjust the calculation method of information gain or the algorithm parameters of quantum computing.
[0146] The system continuously learns from historical feedback through self-learning algorithms (such as reinforcement learning), adjusts the key parameters in the data processing process, and can flexibly adjust the strategies for data collection, cleaning, and annotation according to the characteristics of different data sets to obtain the best processing effect.
[0147] The method of this embodiment can be used to implement the above system embodiment, and its principle and technical effects are similar, so they will not be elaborated here.
[0148] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principle and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A data collection, cleaning and annotation processing system based on AI technology, characterized in that: include: Data collection module, used to automatically collect raw data from various data sources; Data preprocessing module, used to format, remove duplicates and perform preliminary cleaning on the collected raw data; Data cleaning module, which is used to optimize the data cleaning process based on the calculus of variations and high-order differential equations, remove noise and repair missing values; Data annotation module, used to optimize the data annotation process through information gain and calculation; Data storage and query optimization module, used to optimize data storage structure through matrix decomposition; The feedback and optimization mechanism module is used to automatically adjust the data processing process based on the feedback of cleaning and annotation results.
2. According to the data collection, cleaning and annotation processing system based on AI technology in claim 1, it is characterized in that: The data acquisition module includes an automated data acquisition algorithm, which dynamically identifies and adapts to multiple data sources through deep learning technology, automatically collects data from structured and unstructured data sources, and converts the data into a standard format.
3. According to the data collection, cleaning and annotation processing system based on AI technology in claim 1, it is characterized in that: The data cleaning module includes a variational optimization objective function, which is defined as: in, is the loss function of data cleaning, x i is the original data point, c i is the data point after cleaning, and n is the total number of data points.
4. According to the data collection, cleaning and annotation processing system based on AI technology in claim 1, it is characterized in that: The data annotation module optimizes label allocation by information gain, and the information gain formula is: IG(D,Y)=H(Y)-H(Y|D); Among them, IG(D,Y) represents information gain, H(Y) is label entropy, and H(Y|D) is conditional entropy.
5. According to the data collection, cleaning and annotation processing system based on AI technology in claim 1, it is characterized in that: The data storage and query optimization module decomposes the data set into a potential feature matrix W and a label matrix H by matrix decomposition, and the matrix decomposition is defined as: D α ≈WH; Among them, D α is the original data matrix, W is the feature matrix, and H is the label matrix.
6. The data collection, cleaning and annotation processing system based on AI technology according to claim 1 is characterized in that: The feedback and optimization mechanism module includes: Feedback unit: used to provide feedback on optimization strategies based on the results of data cleaning and labeling; Self-learning unit: used to dynamically adjust the cleaning and labeling process through reinforcement learning algorithms; Adjustment unit: used to adjust the parameters of data cleaning and annotation modules.
7. The data collection, cleaning and annotation processing system based on AI technology according to claim 1 is characterized in that: The data cleaning module includes: Variational Optimization Unit: used to construct and solve optimization problems for data cleaning; High-order differential equation unit: used to model the acceleration during data cleaning; Data repair unit: used to repair missing values according to the results of high-order differential equation optimization.
8. The data collection, cleaning and annotation processing system based on AI technology according to claim 1 is characterized in that: The data acquisition module comprises: Data source identification unit: used to automatically identify and classify different types of data sources; Intelligent collection unit: used to automatically collect data from the identified data source; Data preprocessing unit: used to perform preliminary cleaning and standardization of the collected data.
9. The data collection, cleaning and annotation processing system based on AI technology according to claim 1 is characterized in that: The data annotation module includes: Information gain unit: used to calculate the contribution of each data point to the overall annotation task; Computing unit: used to accelerate the label allocation process; Label adjustment unit: used to optimize label assignment according to information gain and quantum computing results.
10. A data collection, cleaning and annotation processing method based on AI technology, applied to a data collection, cleaning and annotation processing system based on AI technology as claimed in any one of claims 1 to 9, characterized in that: The following steps are involved: Automatically collect raw data from multiple data sources; Format, remove duplicates and perform preliminary cleaning of the collected raw data; Use the calculus of variations and higher-order differential equations to clean the data, remove noise and repair missing values; Optimize data annotation through information gain and quantum computing, and automatically assign labels to each data point; Optimize the storage of cleaned and labeled data, optimize the data storage structure through matrix decomposition, and improve data access efficiency through query optimization algorithms; Automatically adjust the data processing process based on the feedback of cleaning and annotation results to optimize the data processing strategy.
Citation Information
Cited By
Avionics device based on high-precision inertial navigation and signal processing method
CN120521613A
AI big data intelligent management method and system
CN120850065A