Big data cleaning method for AI cloud computing training
By extracting the unified feature vector of multimodal data in AI cloud computing training, dynamically generate cleaning strategies and perform cross-modal collaborative cleaning and cloud-edge collaborative resource scheduling, the problems of low data cleaning efficiency and uneven resource utilization in the existing technology are solved, efficient and accurate data cleaning and resource optimization are achieved, and the overall performance of AI training is improved.
Patent Information
- Application Number
- CN202510599249.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-10
- Publication Date
- 2025-08-22
AI Technical Summary
The existing data cleaning methods are inefficient in AI cloud computing training, uneven resource utilization, and poor data cleaning matches downstream tasks, unable to adapt to the needs of complex multimodal data and different tasks, and lack adaptive mechanisms, resulting in poor cleaning results.
By extracting the unified feature vector of multimodal data, a data cleaning strategy is dynamically generated, and a cross-modal collaborative cleaning and cloud-edge collaborative elastic resource scheduling is adopted, and closed-loop optimization is performed by combining the feedback from downstream training models. The sparse attention mechanism, lightweight convolution network and causal convolution network are used for feature extraction, and the dual-channel Q network and long-term short-term memory network are used for strategy optimization.
It realizes efficient data cleaning under multimodal data, improves the accuracy and resource utilization of data cleaning, ensures that the cleaned data can better adapt to downstream tasks, and improves the efficiency and accuracy of AI training.
Smart Images

Figure CN120524082A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of big data processing technology, and specifically to a big data cleaning method for AI cloud computing training. Background Art
[0002] With the rapid development of big data and artificial intelligence technologies, AI cloud computing training faces enormous data processing demands. Data cleaning, a crucial preprocessing step in AI training, directly impacts model training effectiveness and the accuracy of subsequent applications. Data quality is crucial, especially when faced with massive amounts of multimodal data. Efficient data cleaning and processing has become a core challenge.
[0003] To address these issues, numerous data cleaning methods have been applied to AI cloud computing training. Traditional cleaning methods primarily rely on manual rules or data preprocessing based on specific algorithms. These methods typically use fixed rules and algorithms to remove noise, repair missing data, or standardize data. However, these methods often overlook the interdependencies between data and the specific requirements of downstream tasks, resulting in certain limitations in the processing process.
[0004] However, existing traditional data cleaning methods have several obvious shortcomings. First, existing cleaning methods are mostly static and single-mode, and cannot adapt to complex multimodal data and the needs of different tasks. This means that in scenarios with large amounts of data or diverse data types, the cleaning effect cannot meet expectations and may even lose useful information. Secondly, most current methods ignore the close connection between data cleaning and downstream tasks, resulting in the cleaned data may not meet the performance requirements of specific tasks, reducing the efficiency and accuracy of AI training. Furthermore, existing methods make insufficient use of computing resources and often adopt a fixed resource allocation method, resulting in resource waste or processing overload, and cannot be dynamically adjusted according to task requirements. Finally, the existing cleaning process lacks an adaptive mechanism and cannot optimize the cleaning strategy in real time based on model feedback, resulting in the inability to effectively improve the effect of data cleaning. Summary of the Invention
[0005] In response to the shortcomings of the existing technology, the present invention provides a big data cleaning method for AI cloud computing training, which solves the problems of low efficiency of existing multimodal data cleaning, uneven resource utilization, and poor matching between data cleaning and downstream tasks.
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: a big data cleaning method for AI cloud computing training, comprising the following steps:
[0007] S1. Extract unified feature vectors of multimodal data and simultaneously analyze data quality indicators and resource status information;
[0008] S2. Dynamically generate a data cleaning strategy based on the feature vector, data quality indicators, and resource status information;
[0009] S3. Perform cross-modal collaborative data cleaning according to the data cleaning strategy, and record corresponding cleaning task load information;
[0010] S4. Based on the recorded cleaning task load information, perform cloud-edge collaborative elastic resource scheduling;
[0011] S5. Based on the feedback results of the downstream training model, close the loop and optimize the data cleaning strategy.
[0012] Preferably, the method of extracting multimodal data features includes:
[0013] For text data, the Transformer with sparse attention mechanism is used to extract text features;
[0014] For image data, a lightweight convolutional network is used to extract image features;
[0015] For time series signal data, a causal convolutional network is used to extract time series features.
[0016] Preferably, the dynamically generated cleaning strategy includes:
[0017] Take missing rate, noise rate, feature vector and resource status parameters as input;
[0018] Output optimal cleaning action and parameters based on dual-channel Q network;
[0019] Dynamically adjust data cleaning quality and resource consumption through multi-objective reward functions.
[0020] Preferably, the multi-objective reward function is defined as:
[0021]
[0022] Where R represents the reward value; β is the weight factor; ΔQ d Indicates the improvement in data quality after cleaning; ΔC g Indicates the change in resource consumption; Indicates the maximum acceptable resource consumption.
[0023] Preferably, the cross-modal collaborative data cleaning includes:
[0024] Calculate the alignment loss between features of data from different modalities based on task-aware contrastive learning mechanism;
[0025] According to the gradient direction of the alignment loss, the feature vector is corrected and updated to complete cross-modal consistency cleaning.
[0026] Preferably, the updating method of the feature correction is:
[0027]
[0028] Among them, f i is the original eigenvector; is the cleaned feature vector; η is the step size parameter; is the gradient of the loss function; is the contrastive learning loss function.
[0029] Preferably, the cloud-edge collaborative elastic resource scheduling includes:
[0030] Predict the resource requirements of the cleaning task in the future through the long short-term memory network;
[0031] Based on the prediction results, dynamically select edge computing, cloud burst expansion or hybrid mode for computing resource scheduling.
[0032] Preferably, the closed-loop optimization data cleaning strategy includes:
[0033] Calculate the joint loss function based on the performance of the cleaned data in the downstream training model;
[0034] According to the result of the joint loss function, the parameters of the cleaning strategy generation module are updated in reverse.
[0035] Preferably, the joint loss function is defined as:
[0036]
[0037] in, is the joint loss function; f clean is the cleaned feature vector; f gt is the feature vector corresponding to the true label; γ is the weight coefficient; y pred is the model prediction value; y true is the true label; CrossEntropy(·) is the cross entropy loss function.
[0038] Big data cleaning system for AI cloud computing training, including:
[0039] Feature extraction module, used to extract multimodal unified features;
[0040] Strategy generation module, used to dynamically generate data cleaning strategies;
[0041] The cleaning execution module is used to perform cross-modal collaborative cleaning and record task load;
[0042] Resource scheduling module, used to schedule cloud and edge node resources based on cleaning task load and prediction results;
[0043] A closed-loop optimization module is used to optimize the cleaning strategy based on downstream training feedback.
[0044] The present invention provides a big data cleaning method for AI cloud computing training.
[0045] Beneficial effects:
[0046] 1. This invention utilizes a dynamic data cleaning strategy generation method based on a dual-channel Q network, which can intelligently adjust the cleaning strategy based on data quality and resource status. This method accurately identifies data cleaning priorities when processing multimodal data. Compared to traditional rule-based cleaning methods in the prior art, this invention avoids the limitations of static cleaning strategies and achieves more efficient data cleaning. This significantly improves processing efficiency and cleaning quality, particularly in scenarios with large and multimodal data volumes.
[0047] 2. By introducing a cross-modal collaborative cleaning mechanism, this invention ensures consistency across different types of data within a shared feature space, thereby improving the accuracy of data cleaning. Compared to existing single-modality data processing methods, this invention solves the problem of feature inconsistency in multimodal data cleaning, reduces the risk of data loss or information bias, and improves the overall quality of multimodal data processing.
[0048] 3. This invention uses LSTM-based resource demand forecasting, combined with flexible resource scheduling via cloud-edge collaboration, to significantly improve computing resource utilization. Compared to existing fixed resource allocation solutions, this invention dynamically adjusts resource scheduling, avoiding resource waste or overload, achieving more flexible and efficient resource allocation and reducing system operating costs.
[0049] 4. The closed-loop optimization mechanism of this invention continuously optimizes data cleaning strategies by incorporating feedback from downstream training models. This feedback-based adjustment approach addresses the mismatch between data cleaning and downstream tasks in existing technologies. It enables real-time adjustments to the data cleaning process based on model requirements, ensuring that cleaned data is better suited to downstream tasks, ultimately improving overall system accuracy and task completion efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] Figure 1 Schematic diagram of the method flow of the present invention;
[0051] Figure 2 It is the system framework structure diagram of the present invention. DETAILED DESCRIPTION
[0052] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the present specification. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0053] Please see the attached Figure 1 , an embodiment of the present invention provides a big data cleaning method for AI cloud computing training, comprising:
[0054] S1. Extract unified feature vectors of multimodal data and simultaneously analyze data quality indicators and resource status information;
[0055] In this embodiment, step S1 involves extracting a unified feature vector from the multimodal data and simultaneously analyzing data quality indicators, including missingness rate and noise rate. Step S1 is a key starting point for the entire data cleaning method, aiming to provide the necessary input data and data quality assessment for subsequent cleaning strategy generation, thereby supporting the accuracy and efficiency of subsequent cleaning decisions.
[0056] When processing multimodal data, the data may come from a variety of sources and formats, including text, images, and time-series signals. To ensure effective processing of data from each modality, this paper designs appropriate feature extraction methods for different modalities. These methods not only extract the most representative features in the data but also adapt to the characteristics of various data types, ensuring that the features used throughout the data cleaning process effectively reflect the underlying patterns in the data.
[0057] For text data, this embodiment uses the Transformer model with a sparse attention mechanism for feature extraction. Transformer is one of the current mainstream models for processing natural language. It primarily uses a self-attention mechanism to calculate the correlation between elements in the input sequence, thereby generating context-aware feature representations. Compared to traditional sequence processing models (such as RNNs or LSTMs), Transformer can better capture global dependencies.
[0058] In this embodiment, a sparse attention mechanism is employed to improve model computational efficiency. Specifically, the sparse attention mechanism limits the scope of each attention calculation, focusing only on a few key positions in the sequence, rather than fully connecting every pair of positions. This reduces computing resource consumption and increases processing speed while maintaining high expressiveness. The final feature vector of the text data is extracted using this mechanism, effectively capturing the key semantic information in the text.
[0059] For image data, this embodiment uses a lightweight convolutional neural network (CNN) for feature extraction. Traditional convolutional neural networks extract local features from images through a series of convolutional and pooling layers, but when faced with high-resolution images, traditional CNNs often face the problem of large computational load.
[0060] To address this issue, this embodiment designs a lightweight convolutional neural network. This network reduces the model size by reducing the number of convolutional layers and the number of parameters in each layer. Despite this, this lightweight network can still effectively extract local features and global information from the image. During the image feature extraction process, the network extracts the spatial features of the image through convolution and pooling operations, and outputs a high-dimensional image feature vector through a fully connected layer.
[0061] For time series signal data, this embodiment uses a causal convolutional network for feature extraction. A causal convolutional network is a network structure designed specifically for time series data. Its main feature is that the convolution operation relies only on past input data, ensuring that the temporal order of the time series signal is not destroyed.
[0062] Specifically, the causal convolutional network gradually extracts the dynamic features of time series data through a series of causal convolutional layers. In each convolutional layer, the convolution operation only acts on historical data points and does not involve future data points. This ensures that the model does not leak future information when performing time series prediction. In this way, the causal convolutional network can effectively capture the temporal dependencies of signals and generate feature vectors with temporal regularity.
[0063] While extracting features, this example evaluates the quality of each data type, primarily by calculating the missing rate and noise rate. These rates are important indicators for assessing data integrity and accuracy, and are crucial for developing subsequent data cleaning strategies.
[0064] The missing rate indicates the proportion of missing values in a data set. Each data sample may contain some missing attribute values, which need to be processed during the cleaning process. The missing rate is calculated by counting the number of missing values in the data set. The specific calculation formula is:
[0065]
[0066] Among them, p is the missing rate; R is the amount of missing data; T is the total amount of data.
[0067] p reflects the integrity of the data. Data with a higher missing rate need to be processed first in the subsequent cleaning process; R represents the number of all missing values in the dataset; and T represents the number of samples in the dataset.
[0068] The noise rate indicates the proportion of outliers or erroneous information in the data. Noise can arise from various errors during data collection, transmission, or storage, and can severely impact data quality. The noise rate is calculated by comparing the difference between the actual data values and the expected data values.
[0069] Specifically, the noise rate is calculated by the following formula:
[0070]
[0071] Among them, x t Indicates the actual collected data value; x p It represents the expected value calculated based on the predetermined model or historical data; all is the total amount of data, including the number of all samples in the data set. r The noise rate reflects the possible errors and inconsistencies in the data. Data with greater noise need to be cleaned first.
[0072] After feature extraction and quality assessment, this embodiment provides each data sample's feature vector and its quality indicators (including missing rate and noise rate) to the subsequent data cleaning strategy generation module. These feature vectors contain key information about the data, while the quality indicators provide a reference for whether the data needs cleaning.
[0073] Specifically, samples with lower data quality (i.e., high missing or noise rates) will be assigned a higher cleaning priority in subsequent steps to ensure they are adequately cleaned. At the same time, the quality of the feature vectors ensures the expressive power of the data, allowing the system to accurately determine the potential value of the data based on these feature vectors.
[0074] Step S1 employs specific network architectures tailored to different data types, ensuring effective processing of diverse data modalities such as text, images, and time series signals. Furthermore, the assessment of missing and noise rates provides a basis for subsequent data cleaning decisions, ensuring that the cleaning strategy prioritizes lower-quality data, thereby optimizing data quality. This step lays a solid foundation for the entire data cleaning process.
[0075] S2. Dynamically generate data cleaning strategies based on feature vectors, data quality indicators, and resource status information;
[0076] In this embodiment, the core of step S2 is the dynamic generation of a data cleaning strategy. This process combines the feature vectors extracted in step S1, data quality indicators (missing rate, noise rate), and the current resource status of the system. It uses reinforcement learning to generate optimal cleaning actions through a dual-channel Q network and a multi-objective reward function to balance the relationship between cleaning quality and resource consumption.
[0077] In this embodiment, the generation of cleaning strategies relies on a reinforcement learning model based on a dual-channel Q-network (DoubleQ-Learning). This model uses two Q functions to estimate the Q value of each action, reducing the bias in Q-value estimation and thus more accurately evaluating the utility of each cleaning action. The system generates the optimal cleaning strategy by selecting the action corresponding to the maximum Q value.
[0078] During the cleaning strategy generation process, the system first takes the missing rate, noise rate, feature vector, and resource status parameters as input. Combining data quality and resource status, it dynamically generates the optimal cleaning action and its parameters for each cleaning task. These inputs represent the data quality assessment, data feature representation, and the current system resource status. Based on these inputs, the system can make decisions and select appropriate cleaning actions based on the current environment and historical data.
[0079] Missing rate: This indicates the proportion of missing values in a dataset. Data with a high missing rate should be cleaned first to ensure data integrity.
[0080] Noise rate: Indicates the proportion of noise in the data. Noisy data has a significant impact on model training and data analysis, so data with a high noise rate should be cleaned first.
[0081] Feature vectors: Feature vectors are high-dimensional vector representations of the various types of data extracted in step S1, containing key information about the data. These features are the basis for cleaning decisions.
[0082] Resource status parameters: Resource status includes the computing resources currently available to the system, such as memory, CPU usage, bandwidth, etc. Through this resource status information, the system can determine whether more resources need to be allocated to certain cleaning tasks or whether the priority of the tasks needs to be adjusted.
[0083] In this example, the dual-channel Q network optimizes the decision-making process for cleaning tasks. The Q network takes the current state and action as input and predicts the Q value for each action. This value represents the long-term reward of performing an action in a given state. To avoid overestimating Q values, the dual-channel Q network uses two independent Q functions, estimating the Q value separately and updating the other using one of the two Q functions.
[0084] The Q value of each cleaning action is estimated according to the following formula:
[0085]
[0086] in:
[0087] Q(s t ,a t ) means in state s t Next, perform action a t The Q value after .
[0088] r t is the action a executed at the current time step t t The instant reward obtained is calculated based on the cleaning effect and resource consumption, and is determined by a multi-objective reward function.
[0089] γ is a discount factor that balances the weight of future rewards with current rewards. The larger the discount factor, the greater the impact of future rewards on the decision.
[0090] is in the next state s t+1 Next, all possible actions a t+1 The corresponding maximum Q value. It indicates the maximum long-term reward that the system can obtain after taking the best action based on the current state.
[0091] During the training process, the dual-channel Q network gradually optimizes the cleaning strategy by continuously updating the Q value, enabling the system to select cleaning actions that maximize long-term returns.
[0092] In order to comprehensively consider the quality of data cleaning and the resources consumed, this embodiment uses a multi-objective reward function to optimize the cleaning strategy. The reward function is designed to balance two goals: on the one hand, improving the quality of data after data cleaning, and on the other hand, reducing the resource consumption during the data cleaning process. The specific reward function is as follows:
[0093]
[0094] in:
[0095] R represents the reward value of the current cleaning action. The larger the value, the better the balance between data quality improvement and resource consumption.
[0096] β is a weight factor used to adjust the relative importance between data quality improvement and resource consumption. Through this factor, the system can flexibly adjust according to demand and give priority to a certain goal.
[0097] ΔQ dIndicates the improvement in data quality after cleaning. Typically, this metric is calculated by comparing quality indicators (such as accuracy and completeness) before and after data cleaning.
[0098] is the data quality assessment value before cleaning. It is a quality indicator before data cleaning, usually composed of indicators such as data missing rate and noise rate.
[0099] ΔC g Indicates the change in resources consumed by the cleaning task. This metric is usually calculated based on the usage of CPU, memory, and network bandwidth.
[0100] The maximum acceptable resource consumption is usually limited by the system's hardware configuration or task scheduling strategy.
[0101] Through this reward function, the system can consider the balance between improving data quality and resource consumption when executing each cleaning task. A higher reward value indicates that the cleaning strategy consumes less computing resources while improving data quality.
[0102] During the cleaning strategy generation process, the system utilizes a dual-channel Q network for dynamic adjustment and optimization based on the current state, historical data, and real-time feedback. Each cleaning task decision calculates a Q value based on the current environmental state (i.e., data quality and system resources) and selects the optimal cleaning action based on the maximum Q value. This process relies on historical data experience and feedback, and the system optimizes the cleaning strategy by continuously updating the Q value during training.
[0103] At the same time, through the dynamic adjustment of the multi-objective reward function, the system can flexibly find a balance between data cleaning quality and resource consumption. The optimization goal of the cleaning strategy is to maximize the comprehensive value of the reward function, ensuring the data cleaning effect while avoiding excessive resource consumption.
[0104] Through step S2 in this embodiment, the system dynamically generates the optimal cleaning strategy based on data characteristics, quality assessment, and current resource status. The dual-channel Q network optimizes the decision-making process by accurately evaluating the Q-values of cleaning actions. The multi-objective reward function comprehensively considers both data quality improvement and resource consumption, ensuring efficient execution of cleaning tasks while avoiding resource waste. This approach enables intelligent optimization of large-scale data cleaning tasks, dynamically adjusting cleaning strategies based on actual conditions, thereby improving data quality while maximizing resource utilization.
[0105] S3. Perform cross-modal collaborative data cleaning based on the data cleaning strategy and record the corresponding cleaning task load information.
[0106] In this embodiment, step S3 involves executing cross-modal collaborative data cleaning based on the previously generated data cleaning strategy, and recording the corresponding cleaning task load information during this process. This step is a key step in implementing intelligent cleaning operations. Through cross-modal collaborative cleaning, data from different modalities can be efficiently cleaned in coordination with each other. Furthermore, recording cleaning task load information provides an important basis for subsequent resource scheduling.
[0107] In the cross-modal collaborative cleaning process, the present invention adopts a method based on contrastive learning mechanism to ensure that different modal data can achieve good alignment, and on this basis, perform data cleaning. Contrastive learning is a method of learning feature representation by maximizing the similarity between data. In this embodiment, cross-modal data cleaning requires not only feature correction within a single modality data, but also ensuring that the data between each modality has good alignment, so as to achieve a more efficient cleaning effect.
[0108] Cross-modal collaborative cleaning is achieved in this embodiment through the following steps:
[0109] Inter-modal alignment: Based on the feature vectors of the different modal data, the system first calculates the alignment loss between the modal data. The alignment loss reflects the similarity of the different modal data in the shared feature space. The goal is to make the feature vectors of different modalities as close as possible within the same space.
[0110] The alignment loss is calculated as follows:
[0111]
[0112] in, and represent the feature vector of the i-th data sample in text and image modalities respectively, and N is the total number of data samples.
[0113] By minimizing the alignment loss, the system is able to map data from different modalities into a shared feature space, ensuring consistency between them.
[0114] Inter-modal feature correction: Based on the alignment loss calculation, the system further adjusts the features of different modalities to make their representations more consistent in the shared space. By minimizing the gradient of the alignment loss, the system effectively adjusts the feature vectors to correct for differences between modalities. This feature correction process is implemented through a contrastive learning mechanism, where the optimization goal is to make the features of the same data sample across different modalities as similar as possible, thereby improving data cleaning quality.
[0115] The update formula for feature correction is as follows:
[0116]
[0117] Where η is the learning rate, is the gradient of the alignment loss. Through this process, the system can optimize the data features, making the representation of data of different modalities more consistent in the feature space, thereby improving the data cleaning effect.
[0118] In the feature correction process, the feature vector f i Adjustments are made using the following update formula:
[0119]
[0120] in:
[0121] f i Represents the original feature vector. It is the data feature extracted from step S1, and the correction is performed on this basis.
[0122] Represents the cleaned feature vector. It is the new feature obtained after feature correction and is used to characterize the cleaned data.
[0123] η is the learning rate, which controls the step size of each update and affects the amplitude of feature correction.
[0124] is the gradient of the loss function. It represents the error of the model in the current state and determines the adjustment amount in the correction direction.
[0125] It is a comparative learning loss function used to measure the similarity or difference between feature vectors and ensure that the features of different modal data are aligned in the shared feature space.
[0126] The general form of the contrastive learning loss function is as follows:
[0127]
[0128] in, and are the corrected feature vectors of data samples i and j, respectively. The loss function measures their similarity in the feature space. By minimizing the contrastive learning loss, the system ensures that the feature vectors of similar data samples are closer.
[0129] The update process of feature correction introduces the Sigmoid function, that is, The Sigmoid function, as a nonlinear activation function, compresses the gradient during the update process, preventing overcorrection or gradient explosion and ensuring stable and smooth feature correction. Through this mechanism, the update step size adjusts as the gradient changes, avoiding instability in the presence of large gradients.
[0130] Learning the gradient of the loss function by contrast We can determine the update direction of the feature vector. This gradient indicates the amount of adjustment required for the data sample in the feature space and modifies the features in the direction of minimizing the loss function. Each iteration will be updated according to this gradient until the desired feature cleaning effect is achieved.
[0131] In cross-modal cleaning, the goal of feature correction is to maintain consistency in the shared feature space between data from different modalities. For example, when cleaning data from different modalities, such as images, text, and time series data, their feature vectors should be aligned as closely as possible. Guided by a contrastive learning loss function, the feature vectors are corrected to cluster similar data samples across modalities. This process can effectively improve the quality of data cleaning, particularly in multimodal data processing, by enhancing the interrelationships between different modalities.
[0132] During the data cleaning process, in addition to correcting and aligning data features, this embodiment also requires the system to record the load information of each cleaning task.
[0133] Load information mainly includes the computing resources consumed during task execution, memory usage, execution time, etc. Recording this information is crucial for subsequent resource scheduling, task optimization, and performance evaluation.
[0134] Load information is recorded in the following ways:
[0135] Computing resource consumption: The system monitors the computing resource usage of each cleaning task, including processor (CPU or GPU) load, memory usage, and the number of computing cycles required to execute the task. This data helps assess the computational cost of each cleaning task and provides a basis for subsequent resource scheduling.
[0136] Execution time: Records the execution time of each cleaning task from start to finish. This data can help the system evaluate the efficiency of the cleaning task and provide a reference for subsequent optimization.
[0137] Task priority: Based on data quality (such as missing rate, noise rate) and system resource conditions, the system will assign a priority to each cleaning task. Tasks with higher priorities will be executed first when resources are sufficient.
[0138] The record format of the cleaning task load information is as follows:
[0139] Task Load={CPU usage,Memory usage,Execution time,Priority}
[0140] By recording load information in real time, the system can continuously optimize task scheduling strategies, rationally allocate computing resources, avoid resource overload, and ensure efficient completion of cleaning tasks.
[0141] When executing cleaning tasks, the system makes scheduling decisions based on real-time load information. Especially when system resources are limited, the system can dynamically adjust the order of cleaning tasks based on task priority. Tasks with higher priorities will be executed first, ensuring that tasks with poor data quality are cleaned promptly.
[0142] Furthermore, by analyzing historical task load data, the system can identify bottlenecks in cleaning tasks and optimize them. For example, if a cleaning task consumes a large amount of resources during execution, the system may adjust the execution method of the task or allocate more resources to speed up the task.
[0143] In step S3 of this embodiment, the update method of feature correction is based on the contrastive learning mechanism and is combined with the gradient compression strategy of the Sigmoid function. This method ensures that cross-modal data can be effectively aligned in the shared feature space by continuously optimizing the feature vectors of the data samples, thereby improving the quality of data cleaning. Through the above feature correction process, the present invention can maintain a high cleaning accuracy when processing different types of data and ensure that the performance of the data after cleaning meets expectations.
[0144] S4. Based on the recorded cleaning task load information, perform cloud-edge collaborative elastic resource scheduling;
[0145] In this embodiment, step S4 implements flexible resource scheduling for cloud-edge collaboration based on cleaning task load information. This is a key step in ensuring the efficient execution of large-scale data cleaning tasks. By accurately predicting and flexibly scheduling resource requirements, the system can effectively utilize computing resources and avoid resource waste while meeting performance requirements. This embodiment uses a resource demand prediction and dynamic scheduling strategy based on a long short-term memory network (LSTM) to maximize the utilization efficiency of system resources.
[0146] During cloud-edge collaborative resource scheduling, the system flexibly allocates computing resources through two resource pools: a cloud resource pool and an edge computing resource pool. This scheduling strategy makes decisions based on real-time load information for each cleaning task and the system's current resource status. Specifically, when data cleaning tasks involve massive amounts of data and high computing requirements, the system prioritizes cloud resources; when tasks are latency-sensitive, the system prefers edge computing resources.
[0147] During the data cleaning process, the system records the load information of each cleaning task, including but not limited to computing resource consumption, memory usage, bandwidth consumption, and execution time. This load information provides key input for subsequent resource scheduling.
[0148] Specifically, the system uses a long short-term memory (LSTM) network to predict resource requirements based on historical resource consumption data from cleaning tasks. LSTM is particularly well-suited for modeling time series data and can capture the long-term dependencies of cleaning task resource requirements. By using LSTM, the system can accurately predict resource requirements for a period of time before task execution, enabling preemptive resource scheduling decisions.
[0149] The input of LSTM is the resource consumption data of historical time steps, and the output is the prediction of future resource demand. The prediction formula is:
[0150]
[0151] in:
[0152] Represents the system's predicted value of cleaning task resource consumption at time step t.
[0153] C g (tk) represents the actual resource consumption data at historical time step tk.
[0154] k is the number of historical time steps, representing the model learning from past information to predict resource demand.
[0155] Through this prediction, the system can accurately assess the resources required for each cleaning task and make scheduling decisions accordingly.
[0156] Based on the resource requirements predicted by LSTM and the characteristics of the cleaning task, the system will adopt the following two scheduling strategies for resource allocation:
[0157] Cloud Scheduling: When tasks require high computational resources or have large datasets, the system prioritizes cloud-based processing. Cloud resources offer greater computing and storage capabilities, making them suitable for processing large datasets. In these cases, the system uses a resource prediction model to determine whether to assign tasks to the cloud, ensuring efficient processing.
[0158] Edge Scheduling: When tasks require low-latency processing or data processing needs to be close to the data source, the system selects edge computing nodes for execution. Edge computing nodes offer low latency and geographical proximity, making them suitable for time-sensitive tasks. Edge computing nodes typically have limited resources, so flexible scheduling is required based on the actual needs of the task and available resources.
[0159] In the scheduling process between cloud and edge computing resource pools, the priority of tasks is an important decision-making factor. Tasks with higher priorities (for example, tasks with poor data quality or high cleaning timeliness requirements) will be given priority in obtaining resources.
[0160] Specifically, the system assigns priorities to tasks based on the following factors:
[0161] Data quality indicators: For data with a high missing rate or a large noise rate, the system will give higher priority to cleaning.
[0162] Latency requirements: For tasks that require fast processing (such as real-time data stream processing), the system will allocate higher priority resources to them and minimize latency.
[0163] Computing requirements: For tasks with high computing requirements, the system will give priority to allocating cloud resources to ensure that the tasks can be completed within a limited time.
[0164] The priority of a task will directly affect the priority of its resource scheduling. Specifically, the resource allocation process will be carried out in order of priority to ensure that high-priority tasks can obtain resources in a timely manner.
[0165] As tasks progress, the system monitors each task's resource consumption in real time and dynamically adjusts resources based on this data. If a task requires more resources than expected during execution, the system can dynamically adjust its resource allocation, even migrating the task from an edge computing node to the cloud, or from the cloud to an edge computing node with sufficient resources.
[0166] The decision to migrate a task is based on the following circumstances:
[0167] Insufficient resources: When a computing node (cloud or edge computing node) is short of resources, the system will migrate tasks to a node with more abundant resources.
[0168] Changes in latency requirements: If the latency requirements of a task change (for example, the task processing time becomes longer, or needs to be completed in a shorter time), the system will adjust the resource scheduling strategy based on actual needs.
[0169] To further optimize resource scheduling, this embodiment designs a feedback mechanism. The system dynamically adjusts resource scheduling strategies by monitoring task execution feedback in real time. The content of the feedback mechanism includes but is not limited to:
[0170] Task execution time: If the execution time of some tasks exceeds expectations, the system will re-evaluate its resource allocation strategy and consider whether more resources need to be allocated.
[0171] Task completion efficiency: The system evaluates the completion efficiency of tasks and adjusts the priority and resource allocation of tasks based on the relationship between resource consumption and task progress.
[0172] Through continuous feedback optimization, the system can implement a closed-loop feedback mechanism for resource scheduling, ensuring that each cleaning task is efficiently executed within a reasonable resource range.
[0173] Through step S4 in this embodiment, the system can flexibly schedule computing resources based on the load information of the cleaning task. Using the LSTM model to predict resource requirements enables the system to make resource scheduling decisions in advance. At the same time, a cloud-edge collaborative scheduling method is adopted to allocate tasks to the most suitable resource pool (cloud or edge computing node) based on the computing requirements and latency requirements of the task. Dynamic resource adjustment and task migration functions further ensure that the system can optimize resource allocation according to actual conditions during operation, improving the execution efficiency and resource utilization of data cleaning tasks.
[0174] S5. Based on the feedback results of the downstream training model, close the loop and optimize the data cleaning strategy.
[0175] In this embodiment, step S5 involves closed-loop optimization of the data cleaning strategy based on feedback from the downstream training model. This process combines the downstream model's data processing performance with adjustments to the data cleaning process, allowing the data cleaning strategy to be dynamically optimized based on the training model's feedback, thereby ensuring that the quality of the cleaned data can effectively support downstream tasks and improving the efficiency and accuracy of data processing. Through this closed-loop feedback mechanism, the system can adaptively adjust the cleaning strategy and continuously optimize the data cleaning effect through the backpropagation algorithm.
[0176] In step S5, the optimization of the data cleaning strategy depends on the feedback of the downstream training model. Specifically, when the data cleaning is completed and input into the downstream model, the model will show the corresponding training effect according to the quality of the input data. This effect is usually measured by the loss function of the model, such as the training loss (L train ) or prediction accuracy, etc. The system evaluates the effect of data cleaning by comparing the difference between the actual results and the expected targets.
[0177] The key to the feedback process is adjusting the data cleaning strategy through model feedback. This feedback mechanism ensures that the data cleaning process is not only one-way, but also iteratively optimized according to task requirements, thereby continuously improving data quality and ensuring that the data can better meet the needs of downstream tasks.
[0178] To achieve closed-loop optimization, this embodiment employs a joint loss function. This function combines the quality loss of the data cleaning process with the performance loss of downstream tasks, thereby balancing the goals of data cleaning and downstream tasks. Specifically, the joint loss function combines the data cleaning loss and the downstream model loss through a weighted summation, allowing the system to consider both data quality and downstream task performance when optimizing the data cleaning strategy.
[0179] The specific form of the joint loss function is:
[0180]
[0181] in:
[0182] f clean Represents the cleaned data feature vector, f gt is the target feature vector before cleaning. This loss measures the impact of the data cleaning process on the data features, ensuring that data cleaning does not destroy the core features of the original data.
[0183] γ is a weight factor used to adjust the weight between data cleaning loss and downstream task loss. By adjusting this factor, the system can flexibly balance the quality of data cleaning and the performance requirements of downstream tasks.
[0184] ||f clean -fg t || 2 The data cleaning loss measures the difference between the cleaned features and the standard features. This loss ensures that key information is not excessively lost during the cleaning process.
[0185] CrossEntropy(y pred ,y true ) is the cross entropy loss in downstream tasks (such as classification tasks), y pred is the predicted value of the model, y true is the true label. This loss measures the accuracy of the downstream model after using the cleaned data, ensuring that the cleaned data can effectively support downstream tasks.
[0186] During data cleaning strategy optimization, the system updates the strategy parameters through backpropagation of the joint loss function. Specifically, the system calculates the gradient of the joint loss function and updates the various cleaning strategy parameters based on this gradient information. Through the backpropagation algorithm, the system can dynamically adjust various operations in the data cleaning process (such as data repair, denoising, and data augmentation) based on model feedback.
[0187] The parameter update formula is as follows:
[0188]
[0189] in:
[0190] Δθ represents the update amount of the data cleaning strategy parameters.
[0191] η is the learning rate, which controls the step size of each update.
[0192] is the gradient of the joint loss function, which indicates the optimization direction of the data cleaning strategy. By minimizing the joint loss, the system can adjust the data cleaning process to achieve the best balance between data quality and downstream task performance.
[0193] Through the above process, the system is able to gradually optimize the data cleaning strategy to maximize the comprehensive benefits of data quality and downstream task performance.
[0194] In step S5, the optimization of the data cleaning strategy is a dynamic process. The system will continuously adjust the strategy based on the feedback from downstream tasks. This adjustment process relies on the feedback mechanism, and its basic operation is as follows:
[0195] Model feedback evaluation: The system evaluates the effectiveness of data cleaning strategies through feedback from trained models. If the model performs poorly, it indicates a problem in the data cleaning process. The system will adjust the cleaning strategy, such as adding data augmentation or fixing missing data.
[0196] Task feedback influences strategy adjustments: By analyzing the loss and accuracy of model training, the system can determine the effectiveness of the data cleaning strategy. If the cleaned data cannot effectively support downstream tasks, the system will adjust the cleaning parameters and optimize the data cleaning process.
[0197] Balancing data quality and task performance: The core goal of the feedback mechanism is to balance the quality of data cleaning with the execution performance of downstream tasks. By adjusting the γ weighting factor, the system can flexibly adjust the cleaning strategy to meet different task requirements. For example, when data cleaning has a significant impact on downstream tasks, the system will increase data cleaning optimization efforts; when downstream tasks have higher performance requirements, the system will prioritize optimizing task performance.
[0198] In step S5, this embodiment implements a closed-loop optimization process for the data cleaning strategy using feedback from the downstream training model. By combining a loss function with a feedback mechanism, the system enables dynamic adjustment of the data cleaning process based on task requirements. Through the backpropagation algorithm, the system adjusts the various parameters of the cleaning strategy based on the performance of the training model, ensuring a balance between data quality and downstream task performance. Furthermore, this feedback mechanism makes the entire data cleaning process adaptive, enabling continuous optimization as task requirements and data characteristics change, ultimately achieving optimal coordination between data cleaning and task performance.
[0199] The big data cleaning system for AI cloud computing training described below and the big data cleaning method for AI cloud computing training described above can be referenced to each other.
[0200] See also Figure 2 The present invention also provides a big data cleaning system for AI cloud computing training, including:
[0201] Feature extraction module, used to extract multimodal unified features;
[0202] Strategy generation module, used to dynamically generate data cleaning strategies;
[0203] The cleaning execution module is used to perform cross-modal collaborative cleaning and record task load;
[0204] Resource scheduling module, used to schedule cloud and edge node resources based on cleaning task load and prediction results;
[0205] A closed-loop optimization module is used to optimize the cleaning strategy based on downstream training feedback.
[0206] The system of this embodiment can be used to execute the above method embodiments, and its principles and technical effects are similar, so they will not be repeated here.
[0207] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A big data cleaning method for AI cloud computing training, characterized in that: The following steps are involved: S1. Extract unified feature vectors of multimodal data and simultaneously analyze data quality indicators and resource status information; S2. Dynamically generate a data cleaning strategy based on the feature vector, data quality indicators, and resource status information; S3. Perform cross-modal collaborative data cleaning according to the data cleaning strategy, and record corresponding cleaning task load information; S4. Based on the recorded cleaning task load information, perform cloud-edge collaborative elastic resource scheduling; S5. Based on the feedback results of the downstream training model, close the loop and optimize the data cleaning strategy.
2. The big data cleaning method for AI cloud computing training according to claim 1, characterized in that: The method of extracting multimodal data features includes: For text data, the Transformer with sparse attention mechanism is used to extract text features; For image data, a lightweight convolutional network is used to extract image features; For time series signal data, a causal convolutional network is used to extract time series features.
3. The big data cleaning method for AI cloud computing training according to claim 1, characterized in that: The dynamically generated cleaning strategy includes: Take missing rate, noise rate, feature vector and resource status parameters as input; Output optimal cleaning action and parameters based on dual-channel Q network; Dynamically adjust data cleaning quality and resource consumption through multi-objective reward functions.
4. The big data cleaning method for AI cloud computing training according to claim 3, characterized in that: The multi-objective reward function is defined as: Where R represents the reward value; β is the weight factor; ΔQ d Indicates the improvement in data quality after cleaning; ΔC g Indicates the change in resource consumption; Indicates the maximum acceptable resource consumption.
5. The big data cleaning method for AI cloud computing training according to claim 1, characterized in that: The cross-modal collaborative data cleaning includes: Calculate the alignment loss between features of data from different modalities based on task-aware contrastive learning mechanism; According to the gradient direction of the alignment loss, the feature vector is corrected and updated to complete cross-modal consistency cleaning.
6. The big data cleaning method for AI cloud computing training according to claim 1, characterized in that: The update method of the feature correction is: Among them, f i is the original eigenvector; is the cleaned feature vector; η is the step size parameter; is the gradient of the loss function; is the contrastive learning loss function.
7. The big data cleaning method for AI cloud computing training according to claim 1, characterized in that: The cloud-edge collaborative elastic resource scheduling includes: Predict the resource requirements of the cleaning task in the future through the long short-term memory network; Based on the prediction results, dynamically select edge computing, cloud burst expansion or hybrid mode for computing resource scheduling.
8. The big data cleaning method for AI cloud computing training according to claim 1, characterized in that: The closed-loop optimization data cleaning strategy includes: Calculate the joint loss function based on the performance of the cleaned data in the downstream training model; According to the result of the joint loss function, the parameters of the cleaning strategy generation module are updated in reverse.
9. The big data cleaning method for AI cloud computing training according to claim 8, characterized in that: The joint loss function is defined as: in, is the joint loss function; f clean is the cleaned feature vector; f gt is the feature vector corresponding to the true label; γ is the weight coefficient; y pred is the model prediction value; y true is the true label; CrossEntropy(·) is the cross entropy loss function.
10. A big data cleaning system for AI cloud computing training, applied to the big data cleaning method for AI cloud computing training according to any one of claims 1 to 9, characterized in that: include: Feature extraction module, used to extract multimodal unified features; Strategy generation module, used to dynamically generate data cleaning strategies; The cleaning execution module is used to perform cross-modal collaborative cleaning and record task load; Resource scheduling module, used to schedule cloud and edge node resources based on cleaning task load and prediction results; A closed-loop optimization module is used to optimize the cleaning strategy based on downstream training feedback.