AI Corpus Data Automatic Screening and Acquisition Method Based on Big Data Analysis
By dynamically evaluating data sources and constructing a spatiotemporal selection model, the method ensures high-quality data is selected and adapted to AI training needs, improving data stability and reducing noise, thus enhancing AI training efficiency.
Patent Information
- Application Number
- CN202510435440.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-04-09
AI Technical Summary
The existing AI corpus data acquisition and screening methods lack dynamic adaptability and cannot adaptively adjust according to changes in data source and changes in AI training requirements, resulting in the adoption of low-quality or redundant data and the error removal of high-quality data, making it difficult for data quality to remain stable for a long time.
By building a dynamic source confidence evaluation mechanism, dynamic reliability parameters are generated, combined with spatiotemporal weighted screening models and intelligent decision vectors, real-time monitoring and multi-dimensional screening of data sources are realized, acquisition frequency, cleaning strategies are dynamically adjusted, low-quality data flows are blocked, and coordinated allocation of data processing nodes is optimized.
It significantly improves the stability and overall quality of data quality, reduces the impact of noise, improves the generalization ability and stability of AI training, ensures the intelligent evolution and targetedness of the data supply chain, and reduces waste of computing resources.
Smart Images

Figure CN119961577B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and particularly to an automatic screening and acquisition method for AI corpus data based on big data analysis. Background Art
[0002] With the rapid development of artificial intelligence technology, the dependence on high-quality corpus data is increasing day by day. The core lies in the quality, timeliness, and multi-modal adaptation ability of the data. Therefore, how to efficiently obtain, screen, and optimize large-scale corpus data has become the key to improving the AI training effect. However, there are still many problems in the current AI corpus data acquisition and screening methods, which are difficult to meet the needs of large-scale AI training tasks.
[0003] First of all, most of the existing data screening methods rely on fixed rules and lack dynamic adaptation ability. Traditional data screening is usually based on preset quality thresholds or simple keyword matching rules, and cannot be adaptively adjusted according to changes in data sources, fluctuations in its own quality, and changes in AI training requirements. As a result, low-quality, redundant, or irrelevant data is wrongly adopted or high-quality data is miseliminated. In addition, the data source credibility evaluation method is relatively single, and the spatio-temporal dynamic characteristics of the data are not fully considered, making it difficult to maintain stable data quality for a long time. Summary of the Invention
[0004] The present invention provides an automatic screening and acquisition method for AI corpus data based on big data analysis.
[0005] The automatic screening and acquisition method for AI corpus data based on big data analysis includes the following steps:
[0006] S1. Dynamic evaluation of source credibility: Real-time monitor the update frequency, content evolution path, and abnormal fluctuation characteristics of each data source, and generate dynamic credibility parameters including credibility decay rate, semantic deviation degree, and abnormal propagation coefficient;
[0007] S2. Multi-dimensional screening modeling: Based on the dynamic credibility parameters, construct a spatio-temporal weighted screening model, and synchronously fuse the real-time training requirement characteristics to generate a screening decision vector including quality threshold, modality weight, and timeliness sensitivity;
[0008] S3. Adaptive acquisition execution: Generate a multi-dimensional acquisition instruction set according to the dynamic credibility parameters and the screening decision vector, and control the distributed acquisition nodes to execute:
[0009] Dynamically adjust the acquisition frequency;
[0010] Trigger heterogeneous data cleaning;
[0011] Block low-quality data streams;
[0012] After executing the above, the screened corpus data is obtained.
[0013] Optionally, S1 specifically includes:
[0014] S11, Update frequency monitoring: By sliding a time window to statistically analyze the release interval time series of each data source, calculate the update interval (frequency) fluctuation coefficient, and generate a credibility decay rate based on the interval fluctuation coefficient, where the fluctuation coefficient and the decay rate are exponentially positively correlated;
[0015] S12, Content evolution tracking: Adopt dynamic semantic baseline comparison technology, perform LSTM modeling on the historical content of the data source to generate a semantic evolution baseline, and calculate the cosine similarity deviation between the current content and the semantic evolution baseline in real time to generate a semantic deviation degree;
[0016] S13, Abnormal fluctuation detection: Construct a data source association graph, and use a graph neural network to analyze the abnormal propagation path across data sources in real time, and statistically analyze the abnormal input edge weight and output edge diffusion speed of the target data source to generate an abnormal propagation coefficient;
[0017] S14, Parameter fusion calculation: Normalize the credibility decay rate, semantic deviation degree, and abnormal propagation coefficient, and construct a three-dimensional feature vector of dynamic credibility parameters.
[0018] The said S2 specifically includes:
[0019] S21, Spatiotemporal weight assignment: Based on the dynamic credibility parameters, construct a three-dimensional weight matrix, and assign weights to the three dimensions of time weight, semantic weight, and topological weight respectively;
[0020] S22, Training requirement fusion:
[0021] Perform double fusion of real-time training requirement features and the spatiotemporal weight matrix:
[0022] Linear feature interaction: Multiply the model training requirement features with the weights of each dimension by a diagonal matrix;
[0023] Nonlinear correlation capture: Extract the deep correlation between the weight matrix and the requirement features through an activation function;
[0024] Fusion weight generation: Weightedly superimpose the linear transformation result and the nonlinear interaction result to form a dynamic weight that adapts to the model requirements;
[0025] S23, Decision vector generation: Convert the fusion weight into an operable screening decision parameter.
[0026] Optionally, the time weight is adjusted inversely according to the credibility decay rate, and the data source with a higher decay rate is given a lower time weight;
[0027] The semantic weight uses an exponential decay function to process the semantic deviation degree, and the data source with a large deviation degree reduces its semantic weight;
[0028] The topological weight is reversely normalized based on the anomaly propagation coefficient, and the data source with strong anomaly propagation ability reduces its topological weight.
[0029] Optionally, the generation of the decision vector in S23 specifically includes:
[0030] Dynamic setting of the quality threshold: Combining time and topological weight, an adaptive standard deviation method is used to generate a dynamic quality threshold that changes with the data quality distribution;
[0031] Balanced distribution of modal weights: Based on semantic weights and modal requirements, the acquisition ratio of the data source is generated through normalization processing;
[0032] Control of timeliness sensitivity: Use the Sigmoid function to convert the time parameter into a data update frequency control instruction.
[0033] Optionally, S3 specifically includes:
[0034] S31, dynamically adjust the acquisition frequency, and based on the timeliness sensitivity parameter in the screening decision vector, dynamically generate a data acquisition frequency instruction;
[0035] S32, trigger heterogeneous data cleaning, jointly use the modal weight and semantic deviation degree parameters in the screening decision vector, execute a hierarchical cleaning strategy, allocate different cleaning resources according to the modal weight, and give priority to ensuring the data quality of high-weight modalities;
[0036] S33, block low-quality data streams, and based on the dynamic quality threshold, set dynamic blocking rules:
[0037] Combine the dynamic quality threshold with the safety redundancy coefficient to generate a dynamic blocking threshold;
[0038] Real-time calculate the anomaly propagation degree of the data stream, and block immediately when the anomaly propagation degree exceeds the blocking threshold.
[0039] Optionally, the hierarchical cleaning strategy adopts three-level progressive cleaning:
[0040] Basic cleaning: Standardize the format and remove redundancy;
[0041] Deep cleaning: Cross-modal alignment and contradiction resolution;
[0042] Reconstruction cleaning: Adversarial sample detection and semantic reconstruction.
[0043] Optionally, the dynamic generation of the data acquisition frequency instruction in S31 specifically includes:
[0044] Adjust the basic acquisition frequency proportionally according to the magnitude of the timeliness sensitivity value;
[0045] Fine-tune in combination with the credibility attenuation rate, and implement acquisition frequency reduction compensation for data sources with high attenuation rates;
[0046] When the timeliness sensitivity reaches the critical value, start the pulse enhancement mode for burst data acquisition.
[0047] Optionally, the S3 also includes distributed node collaboration, and dynamically selects the optimal node type according to the maximum value of each screening decision vector component, specifically including:
[0048] Classify and allocate data processing nodes based on the quality assessment, modal weight, and timeliness sensitivity in the screening decision vector. By comparing the values of the screening decision vector, select the dominant factor and allocate the processing node type; if the quality assessment factor is dominant, allocate it to the quality-priority node to ensure the rigor of data screening; if the modal weight is dominant, allocate it to the modal-priority node to make the data meet the modal requirements of different AI tasks; if the timeliness sensitivity is dominant, allocate it to the timeliness-priority node to improve the data update speed; if the values of the three are similar, use the comprehensive balance node for processing to ensure that multiple aspects of requirements are taken into account. In addition, in the case of a high risk of abnormal data propagation, the data will be directed to the abnormal filtering node to perform additional detection or blocking measures.
[0049] Advantages of the present invention:
[0050] 1. By constructing a dynamic source credibility evaluation mechanism, comprehensively monitoring the update frequency, content evolution trajectory, and abnormal propagation characteristics of data sources, generating dynamic credibility parameters including credibility attenuation rate, semantic deviation degree, and abnormal propagation coefficient, and using this parameter to construct a spatio-temporal weighted screening model and combining the characteristics of AI training requirements to generate an accurate screening decision vector, this mechanism realizes the quantitative evaluation of the reliability of data sources, content consistency, and abnormal diffusion risk, avoids low-quality, counterfeit, or data with abnormal propagation from entering the training set, significantly improves the overall quality stability of data, reduces the impact of data noise on the AI training process, and improves the generalization ability and stability of the model.
[0051] 2. By constructing a training requirement vector and fusing it with the data screening weight matrix, dynamically optimizing data acquisition, cleaning, and blocking strategies, and using the screening decision vector to guide the frequency regulation, modal weight allocation, and abnormal data blocking of data acquisition nodes, enabling the data stream to adaptively adjust according to the changes in AI training tasks, thereby realizing the intelligent evolution of the data supply chain. This method ensures that optimal data can be continuously obtained, avoids redundant and inefficient data from occupying computing resources, while improving the pertinence and effectiveness of AI training data, and reducing data mismatch problems during the training process.
[0052] 3. An intelligent decision vector-driven data node collaborative optimization strategy is adopted. According to the real-time calculation results of data quality, modality requirements, and timeliness sensitivity, the optimal data processing node is automatically matched. Through the intelligent division of node types, including quality-priority, modality-priority, timeliness-priority, comprehensive balance, and anomaly filtering nodes, the accurate allocation and processing of data are realized, enabling different types of data to be processed according to the optimal strategy, optimizing the allocation mode of data flow in computing resources, and reducing data redundancy and computing waste. Brief Description of the Drawings
[0053] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required in the description of the embodiments or the prior art. Obviously, the following drawings are only for the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0054] Figure 1 It is a schematic flowchart of the method according to an embodiment of the present invention. Detailed Embodiments
[0055] The following will describe the present invention in detail with reference to the drawings and specific embodiments. At the same time, it should be noted here that in order to make the embodiments more detailed, the following embodiments are the best and preferred embodiments. For some well-known technologies, those skilled in the art can also adopt other alternative methods for implementation; and the drawings are only for more specifically describing the embodiments, and are not intended to specifically limit the present invention.
[0056] As Figure 1 shown, the method for automatically screening and obtaining AI corpus data based on big data analysis includes the following steps:
[0057] S1. Dynamic evaluation of source credibility: Real-time monitor the update frequency, content evolution path, and abnormal fluctuation characteristics of each data source, and generate dynamic credibility parameters including credibility decay rate, semantic deviation degree, and abnormal propagation coefficient;
[0058] S2. Multi-dimensional screening modeling: Based on the dynamic credibility parameters, construct a spatio-temporal weighted screening model, and synchronously fuse the real-time training requirement characteristics to generate a screening decision vector including quality threshold, modality weight, and timeliness sensitivity;
[0059] S3. Adaptive acquisition execution: Generate a multi-dimensional acquisition instruction set according to the dynamic credibility parameters and the screening decision vector, and control the distributed acquisition nodes to execute:
[0060] Dynamically adjust the acquisition frequency;
[0061] Trigger heterogeneous data cleaning;
[0062] Block low-quality data streams;
[0063] After performing the above, filtered corpus data is obtained.
[0064] S1 specifically includes:
[0065] S11, Update frequency monitoring: By sliding a time window to statistically analyze the release interval time series of each data source, calculate the update interval fluctuation coefficient, and generate a credibility decay rate based on the interval fluctuation coefficient, where the interval fluctuation coefficient and the decay rate are exponentially positively correlated;
[0066] S111, Set the time window of the data source , Calculate the release interval time series: , represents the th data release time within the time window, that is, the current latest release time point, represents the th data release time within the time window, represents the th time point, ;
[0067] S112, Calculate the interval fluctuation coefficient : , where represents the standard deviation of the time interval, represents the mean of the time interval.
[0068] S113, Calculate the credibility decay rate : , where is the decay sensitivity adjustment factor (set range ).
[0069] S12, Content evolution tracking: Adopt dynamic semantic baseline comparison technology to perform LSTM modeling on the historical content of the data source to generate a semantic evolution baseline, and calculate the cosine similarity deviation between the current content and the semantic evolution baseline in real time to generate a semantic deviation degree;
[0070] S121, Let the historical content sequence of the data source be , represents the th data content, and use LSTM to calculate the hidden state: , where is the text embedding function, represents the data content at the th time point, represents the hidden state of LSTM at time represents the time The hidden state of the LSTM at a certain time represents the text embedding vector of the data content ;
[0071] S122, calculating the dynamic semantic baseline : ;
[0072] S123, calculating the semantic deviation degree between the current content and the baseline : , where is the embedding vector of the current content, and the cosine similarity is used to measure the semantic change
[0073] S13, anomaly fluctuation detection: constructing a data source association graph, analyzing the anomaly propagation path across data sources in real time through a graph neural network, and statistically calculating the anomaly input edge weight and output edge diffusion speed of the target data source to generate an anomaly propagation coefficient
[0074] S131, setting the data source association graph , where is the data source set is the edge set of data reference relationships is the edge weight, representing the historical anomaly propagation times
[0075] S132, calculating the anomaly propagation coefficient : , where respectively represent the input edge set and the output edge set is the time decay factor respectively represent the input edge weight and the output edge weight: ;
[0076] is the normalization coefficient: , is the current time is the anomaly occurrence time represents the input edge entering the target data source node represents the output edge emitted from the target data source node
[0077] S14, parameter fusion calculation: performing spatio-temporal normalization processing on the credibility decay rate, semantic deviation degree, and anomaly propagation coefficient to generate a three-dimensional feature vector of dynamic credibility parameters .
[0078] S2 specifically includes:
[0079] S21, spatio-temporal weight allocation: constructing a three-dimensional weight matrix based on the dynamic credibility parameters
[0080] ;
[0081] Among them, is the time weight, is the semantic weight, is the topological weight, is the domain adaptation weight (the initial value is set by expert experience), is the smoothing factor, , preventing the denominator from being zero, is the semantic shift sensitivity adjustment parameter, , is the historical maximum anomaly propagation coefficient.
[0082] S22, Training requirement fusion: Real-time obtain the training requirement feature vector of the target Al model :
[0083] ;
[0084] Among them, is the norm of the loss function gradient, , is the norm of the modal attention weight, , is the modal attention matrix, representing the distribution of the attention weights of different modalities in the multi-modal input, is the rate of change of the validation set accuracy over time, , represents the rate of change of the validation set accuracy with respect to time, measuring the performance improvement or decline trend over the training time, is the increment of the training time, used to calculate the performance improvement or decline trend;
[0085] Execute requirement-weight fusion: , among them, is a diagonal matrix to ensure that each training requirement component acts on the corresponding weight respectively, is the fusion intensity coefficient to control the fusion degree, , is the filtered weight matrix after fusion, is the activation function, is the transpose of the training requirement feature vector;
[0086] S23, Decision vector generation: Generate a filtered decision vector through normalization mapping:
[0087] ;
[0088] Among them, is the quality threshold, is the modal weight, is the aging sensitivity;
[0089] S231, Quality threshold calculation: ;
[0090] Dynamically set the quality threshold: , where, are the historical quality mean and standard deviation respectively, , where is the inverse cumulative distribution function of the standard normal distribution, is the quality decision factor, a normalized weight factor used to measure the relative quality of data, is an actual numerical threshold used to decide whether data passes the screening, affected by ;
[0091] S232, Modal weight assignment: , where, is the real-time demand intensity of the th mode;
[0092] S233, Aging sensitivity control , where, is the slope factor, controlling the change rate, , is the offset, , adjusting the reference point of the aging sensitivity;
[0093] are the fused time, topology, and semantic weights respectively.
[0094] S3 specifically includes:
[0095] S31, Dynamic acquisition frequency regulation: Based on the aging sensitivity in the screening decision vector Generate the acquisition frequency instruction: , where, is the adjusted sampling frequency, is the reference acquisition frequency, is the aging sensitivity, is the credibility decay rate, used here as a correction factor, is the adjustment coefficient, , controlling 's influence in the acquisition frequency regulation, is the standard Sigmoid function: .
[0096] S32, Intelligent data cleaning trigger: Jointly use the modal weight in the decision vector and the semantic offset Calculate the data cleaning level: , where is the data cleaning level, is the modal weight, which comes from the decision vector of S2, is the semantic deviation degree, is the semantic deviation reference value for normalization, is the floor function to ensure that the cleaning level is an integer.
[0097] S33, real-time data flow blocking: Based on the quality threshold in the decision vector Calculate the data flow blocking condition:
[0098] ;
[0099] where is the data blocking flag, 0 means not blocking, 1 means blocking, is the abnormal propagation coefficient, is the dynamic quality threshold, is the safety redundancy coefficient, , which is used to adjust the threshold of the blocking condition.
[0100] S34, node collaboration: Map the three elements of the decision vector to node allocation:
[0101] , where represents the data processing node type, returns the index corresponding to the maximum value, , , respectively represent data quality, modal weight, and timeliness sensitivity.
[0102] Decision logic of NodeType:
[0103] If is the largest: Prioritize the quality-first node to ensure that the selected data meets the high-quality standard.
[0104] If is the largest: Allocate to the modal-priority node to match the modal requirements of the AI task.
[0105] If is the largest: Adopt the timeliness-priority node to ensure the real-time nature of the data.
[0106] If the values of the three are similar: Allocate to the comprehensive balance node to optimize data selection among multiple factors.
[0107] The present invention covers any alternatives, modifications, equivalent methods and solutions made to the essence and scope of the present invention. For the public to have a thorough understanding of the present invention, specific details are described in detail in the following preferred embodiments of the present invention, and those skilled in the art can fully understand the present invention without the description of these details. In addition, well-known methods, processes, procedures, components and circuits, etc. are not described in detail in order to avoid unnecessary confusion to the essence of the present invention.
[0108] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention.
Claims
1. An AI corpus data automatic screening and acquisition method based on big data analysis, characterized in that, It includes the following steps: S1. Dynamic evaluation of source credibility: Real-time monitor the update frequency, content evolution path and abnormal fluctuation characteristics of each data source, and generate dynamic credibility parameters including credibility attenuation rate, semantic deviation degree, and abnormal propagation coefficient; S2. Multi-dimensional screening and modeling: Based on the dynamic credibility parameters, construct a spatio-temporal weighted screening model, and synchronously fuse the real-time training requirement characteristics to generate a screening decision vector including quality threshold, modal weight, and timeliness sensitivity; Specifically including: S21. Spatio-temporal weight assignment: Based on the dynamic credibility parameters, construct a three-dimensional weight matrix, and assign weights to the three dimensions of time weight, semantic weight, and topological weight respectively; S22. Fusion of training requirements: Double fusion of real-time training requirement characteristics and spatio-temporal weight matrix: Linear feature interaction: Multiply the model training requirement characteristics with the weights of each dimension by a diagonal matrix; Non-linear correlation capture: Extract the deep correlation between the weight matrix and the requirement characteristics through an activation function; Fusion weight generation: Weightedly superimpose the linear transformation result and the non-linear interaction result to form a dynamic weight adapted to the model requirements; S23. Generation of decision vector: Convert the fusion weight into an operable screening decision parameter; The time weight is adjusted inversely according to the credibility attenuation rate, and the data source with a higher attenuation rate is given a lower time weight; The semantic weight uses an exponential decay function to process the semantic deviation degree, and the data source with a large deviation degree reduces its semantic weight; The topological weight is inversely standardized based on the abnormal propagation coefficient, and the data source with strong abnormal propagation ability reduces its topological weight; S3. Adaptive acquisition execution: Generate a multi-dimensional acquisition instruction set according to the dynamic credibility parameters and the screening decision vector, and control the distributed acquisition nodes to execute: Dynamically adjust the acquisition frequency; Trigger heterogeneous data cleaning; Block low-quality data streams; After performing the above steps, the screened corpus data is obtained.
2. The method for automatically screening and obtaining AI corpus data based on big data analysis according to claim 1, wherein Specifically, S1 includes: S11. Update frequency monitoring: Calculate the update interval fluctuation coefficient by statistically analyzing the release interval time series of each data source through a sliding time window, and generate a credibility attenuation rate based on the interval fluctuation coefficient. The interval fluctuation coefficient and the attenuation rate are exponentially positively correlated; S12. Content evolution tracking: Adopt dynamic semantic baseline comparison technology, perform LSTM modeling on the historical content of the data source to generate a semantic evolution baseline, and calculate the cosine similarity deviation between the current content and the semantic evolution baseline in real time to generate a semantic deviation degree; S13. Abnormal fluctuation detection: Construct a data source association graph, and analyze the abnormal propagation path across data sources in real time through a graph neural network. Statistically analyze the abnormal input edge weight and output edge diffusion speed of the target data source to generate an abnormal propagation coefficient; S14. Parameter fusion calculation: Normalize the credibility attenuation rate, semantic deviation degree, and abnormal propagation coefficient, and construct a three-dimensional feature vector of dynamic credibility parameters.
3. The method for automatically screening and obtaining AI corpus data based on big data analysis according to claim 1, characterized in that, Specifically, the generation of the decision vector in S23 includes: Dynamic setting of quality threshold: Combine time and topological weights, and use the adaptive standard deviation method to generate a dynamic quality threshold that changes with the data quality distribution; Balanced distribution of modal weights: Based on semantic weights and modal requirements, generate the acquisition ratio of the data source through normalization processing; Ageing Sensitivity Control: Using the Sigmoid function to convert the time parameter into a data update frequency control instruction.
4. The method for automatically screening and obtaining AI corpus data based on big data analysis according to claim 1, wherein The specific steps of S3 are as follows: S31: Dynamically adjust the acquisition frequency. Based on the ageing sensitivity parameter in the screening decision vector, dynamically generate a data acquisition frequency instruction. S32: Trigger heterogeneous data cleaning. Combine the modality weight and semantic deviation parameter in the screening decision vector to execute a hierarchical cleaning strategy. Allocate different cleaning resources according to the modality weight to prioritize ensuring the data quality of high-weight modalities. S33: Block low-quality data streams. Set a dynamic blocking rule based on the dynamic quality threshold: Combine the dynamic quality threshold with the safety redundancy coefficient to generate a dynamic blocking threshold. Real-time calculate the abnormal propagation degree of the data stream. When the abnormal propagation degree exceeds the blocking threshold, block it immediately.
5. The method for automatically screening and obtaining AI corpus data based on big data analysis according to claim 4, wherein The hierarchical cleaning strategy adopts a three-level progressive cleaning: Basic cleaning: Standardize the format and remove redundancy. Deep cleaning: Cross-modal alignment and contradiction resolution. Reconstruction cleaning: Adversarial sample detection and semantic reconstruction.
6. The method for automatically screening and obtaining AI corpus data based on big data analysis according to claim 4, characterized in that, The specific steps of dynamically generating the data acquisition frequency instruction in S31 are as follows: Adjust the basic acquisition frequency proportionally according to the magnitude of the ageing sensitivity value. Perform fine-tuning in combination with the credibility decay rate, and implement acquisition frequency reduction compensation for data sources with a high decay rate. When the ageing sensitivity reaches the critical value, start the pulse enhancement mode for burst data acquisition.
7. The method for automatically screening and obtaining AI corpus data based on big data analysis according to claim 4, wherein, S3 also includes distributed node collaboration. Dynamically select the optimal node type according to the maximum value of each component of the screening decision vector. The specific steps are as follows: Classify and allocate data processing nodes according to the quality assessment, modality weight, and ageing sensitivity in the screening decision vector. By comparing the values of the screening decision vector, select the dominant factor and allocate the processing node type. If the quality assessment factor is dominant, allocate it to the quality-priority node to ensure the rigor of data screening. If the modality weight is dominant, allocate it to the modality-priority node to make the data meet the modality requirements of different AI tasks. If the ageing sensitivity is dominant, allocate it to the ageing-priority node to improve the data update speed. If the values of the three are similar, use the comprehensive balance node for processing to ensure that multiple requirements are taken into account.
Citation Information
Patent Citations
Data set construction method and system for training professional field large model
CN119204266A
Building intelligent operation and maintenance management system and method based on big data analysis
CN119692814A