A redundant data preprocessing method for heterogeneous computing platforms

CN122821075APending Publication Date: 2026-09-25BEIJING SHENGXUN TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610861636.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-15
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0005]针对现有技术中去重与调度步骤割裂的问题,本发明提供一种用于异构计算平台的冗余数据预处理方法,以解决上述背景技术中提出的一个或多个问题

Benefits of technology

[0048]1.通过将冗余关系从样本识别阶段延伸至任务调度阶段,使得冗余关系同时驱动样本保留决策与异构任务分配,实现去重准确性、资源效率和稀有保留在同一轮处理中的同步优化,克服了现有技术中去重步骤与调度步骤相互割裂的问题;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122821075A_ABST
    Figure CN122821075A_ABST
Patent Text Reader

Abstract

The present application belongs to the field of data processing, and mainly relates to a redundant data preprocessing method for a heterogeneous computing platform. The method generates a local redundant candidate set by jointly cutting, block description extracting and roughly sorting long video data to be processed. Then, the local redundant candidate set is subjected to cross-modal compression representation generation, data block aggregation, propagation calculation and propagation modeling to obtain a sample reservation decision result. Further, based on the sample reservation decision result, a device affinity matrix and a task allocation result are generated, and parallel preprocessing is performed among a central processing unit, a graphics processing unit, a data processor and a field programmable gate array. The method simultaneously uses the redundant relationship for sample reservation decision and heterogeneous task allocation, can reduce cross-device data moving overhead, inhibit rare sample deletion error and improve long video training data preprocessing efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of data processing and mainly relates to a method for preprocessing redundant data for heterogeneous computing platforms. Background Technology

[0002] In multimodal long video data preprocessing scenarios, the original training data typically includes video frames, audio streams, subtitle text, optical character recognition text, and source metadata. This data exhibits near-duplication within the same modality, semantic duplication across modalities, and structural duplication across time periods. It also contains pseudo-difference samples formed by differences in transcoding, shot segmentation, and frame extraction. Furthermore, different devices on heterogeneous computing platforms differ in terms of computing power, data migration costs, and parallel processing characteristics. Under the combined influence of these complex redundancy types and heterogeneous hardware differences, how to avoid instability in redundant recognition, control cross-device data migration costs, and prevent the accidental deletion of critical low-frequency samples has become a pressing technical problem to be solved in the preprocessing stage of this field.

[0003] To address the aforementioned issues, existing technologies typically first perform slicing, format unification, and fingerprint generation on multimodal long video data. Then, similarity calculations or semantic representation extraction are performed on candidate data blocks to identify redundant relationships such as duplicate or near-duplicate correspondences between different data blocks. After generating deduplication results, relevant processing tasks are assigned to different computing devices for execution based on pre-configured resource strategies. Some technical solutions also introduce hash-accelerated deduplication, cross-modal duplicate detection, and quality filtering to improve the efficiency of large-scale data preprocessing.

[0004] However, existing technologies employing a deduplication-then-scheduling approach suffer from a disconnect between the deduplication and scheduling steps. Firstly, existing technologies only use redundancy relationships as the basis for identifying and deleting duplicate samples, failing to utilize these relationships in subsequent scheduling stages such as task splitting, processing path selection, and device affinity allocation. This neglects the optimization significance of redundancy relationships for processing paths during the scheduling process, resulting in the deduplication results failing to effectively support the scheduling process. Furthermore, the deduplication-then-scheduling approach also leads to the accidental deletion of rare samples. In many identical scenarios, the difference between rare and ordinary samples is minimal, making them easily removed by deduplication. Even though some existing technologies protect rare samples by adding them to a whitelist after scheduling, this additional repair method does not incorporate rare sample protection into the deduplication judgment and scheduling decision-making processes. This indicates that existing technologies struggle to simultaneously achieve scheduling optimization and key sample retention during redundancy reduction, highlighting the disconnect between the deduplication and scheduling steps. Therefore, this invention proposes a preprocessing optimization method that couples the deduplication and scheduling steps and incorporates rare sample protection into both processes. Summary of the Invention

[0005] To address the problem of the separation of deduplication and scheduling steps in existing technologies, this invention provides a method for preprocessing redundant data on heterogeneous computing platforms, thereby solving one or more of the problems mentioned in the background.

[0006] To solve the above problems, the present invention employs the following technology:

[0007] In a first aspect, the present invention proposes a method for preprocessing redundant data for heterogeneous computing platforms, comprising:

[0008] The process involves: acquiring long video data to be processed to obtain raw data; performing coarse sorting on the raw data to generate a set of locally redundant candidates; performing compression characterization on the locally redundant candidate set to generate a cross-modal compressed characterization set; performing data block aggregation based on the cross-modal compressed characterization set to generate a set of refined candidate clusters; performing propagation calculation on the set of refined candidate clusters to obtain redundant propagation records; performing propagation modeling based on the redundant propagation records to generate sample retention decision results; performing affinity matrix generation based on the sample retention decision results to obtain a device affinity matrix; performing task allocation based on the device affinity matrix to obtain task allocation results; and performing data preprocessing based on the task allocation results to obtain preprocessing execution results.

[0009] In some examples, the original data is coarsely sorted to generate a set of locally redundant candidates, including:

[0010] Based on the processing requirements of the training task, set joint splitting conditions, perform joint splitting on the original data, and obtain the original data block set;

[0011] A unified block description set is obtained by extracting descriptions from the original data block set.

[0012] A coarse screening is performed based on the unified block description set to obtain a locally redundant candidate set.

[0013] In some examples, the locally redundant candidate set is compressed to generate a cross-modal compressed representation set, including:

[0014] Feature extraction is performed based on a set of locally redundant candidates to obtain a joint feature vector;

[0015] A linear projection is performed based on the joint feature vector to obtain the cross-modal compressed representation of the current data block.

[0016] In some examples, data block aggregation is performed based on the cross-modal compressed representation set to generate a refined candidate cluster set, including:

[0017] Similarity calculation is performed based on cross-modal compressed representation sets to obtain sparse similarity relation sets;

[0018] A connection graph is constructed based on a sparse similarity set to obtain a data block similarity graph;

[0019] Connectivity component partitioning is performed based on the data block similarity graph to obtain initial candidate clusters;

[0020] A refined set of candidate clusters is obtained by further refining the initial candidate clusters.

[0021] In some examples, selection is based on the initial candidate clusters, including:

[0022] Cluster centers are obtained by calculating the cluster centers based on the initial candidate clusters;

[0023] Similarity filtering is performed based on cluster centers to obtain the data blocks to be verified.

[0024] Similarity retrieval and connectivity merging are re-executed for all data blocks to be verified to generate supplementary candidate clusters. All the obtained supplementary candidate clusters are recorded as a fine candidate cluster set.

[0025] In some examples, propagation calculations are performed on the refined candidate cluster set to obtain redundant propagation records, including:

[0026] Similarity component calculation is performed based on refined candidate clusters to obtain similarity component data;

[0027] Protection values ​​are calculated based on refined candidate clusters to obtain protection values ​​for rare samples;

[0028] A redundancy propagation score is obtained by weighting similarity component data and rare sample protection values.

[0029] Statistical analysis is performed based on redundancy propagation scores, and a redundancy judgment threshold is set.

[0030] The rare sample protection value, propagation score record, and redundancy judgment threshold are associated and recorded as a redundancy propagation record.

[0031] In some examples, protection values ​​are calculated based on refined candidate clusters to obtain protection values ​​for rare samples, including:

[0032] Pre-trained models are used to identify features of fine-grained candidate clusters, resulting in event label frequency tables, scene label frequency tables, text phrase frequency tables, and modality combination frequency tables.

[0033] The inverse frequencies of the event label frequency table, scene label frequency table, text phrase frequency table, and modality combination frequency table are calculated separately, and the inverse frequencies are weighted and summed to obtain the rare sample protection value.

[0034] In some examples, propagation modeling is performed based on the redundant propagation records to generate sample retention decision results, including:

[0035] A redundant propagation graph is obtained by generating a graph structure based on redundant propagation records.

[0036] Based on the generated redundant propagation graph, connected component partitioning is performed to generate propagation subgraphs;

[0037] Using the representative data block as a constraint, the current propagation subgraph is reduced to obtain the sample retention decision result.

[0038] In some examples, affinity matrix generation is performed based on the sample retention decision results to obtain a device affinity matrix, including:

[0039] Based on the sample retention decision results, subtasks are split into preprocessing subtask sets;

[0040] Device affinity calculations are performed based on the preprocessed subtask set to obtain the device affinity matrix.

[0041] In some examples, task allocation is performed based on the device affinity matrix to obtain task allocation results, including:

[0042] Device priority ranking is performed based on device affinity matrix to determine the target execution device;

[0043] Tasks are sorted and divided into execution batches based on the target execution device;

[0044] The target execution device and execution batch are associated and recorded as task allocation results.

[0045] In a second aspect, the present invention provides an electronic device, including a memory and a processor, wherein the memory stores a computer program, characterized in that: when the processor executes the computer program, it implements the steps of the method in the first aspect.

[0046] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, characterized in that: when the computer program is executed by a processor, it implements the steps of the method in the first aspect.

[0047] The beneficial effects of this invention are:

[0048] 1. By extending the redundancy relationship from the sample identification stage to the task scheduling stage, the redundancy relationship simultaneously drives the sample retention decision and heterogeneous task allocation, achieving synchronous optimization of deduplication accuracy, resource efficiency and rare retention in the same round of processing, overcoming the problem of the deduplication step and scheduling step being separated in the existing technology.

[0049] 2. By introducing a rare sample protection value in the redundancy determination process, low-frequency events, rare scenarios, long-tail texts and low-coverage modalities are included in the retention constraints, so that rare sample protection is directly embedded in the redundancy propagation calculation and retention decision process, avoiding the accidental deletion of key samples in the preprocessing stage.

[0050] 3. By combining the acceleration benefits of preprocessing subtasks, data migration costs, available computing power of devices, and queue pressure to generate a device affinity matrix, and performing task allocation and batch partitioning accordingly, different types of preprocessing subtasks can be executed in parallel on heterogeneous computing platforms according to their compatibility, thereby reducing cross-device migration overhead and improving preprocessing efficiency. Attached Figure Description

[0051] Figure 1 This is an exemplary method flowchart of a redundant data preprocessing method for a heterogeneous computing platform provided by an embodiment of the present invention;

[0052] Figure 2 This is a comparison chart showing the effect of a redundant data preprocessing method for heterogeneous computing platforms provided by an embodiment of the present invention compared with traditional methods. Detailed Implementation

[0053] To make the technical means, creative features, and achieved objectives and effects of this invention easier to understand, the invention is further described below with reference to specific embodiments. However, the following embodiments are merely preferred embodiments of this invention and not all of them. Other embodiments obtained by those skilled in the art based on the embodiments described herein without creative effort are all within the protection scope of this invention. Unless otherwise specified, the experimental methods in the following embodiments are conventional methods, and the materials and reagents used in the following embodiments are commercially available unless otherwise specified.

[0054] Example 1 Figure 1 As shown in the flowchart of the present invention, this embodiment provides a method for preprocessing redundant data for heterogeneous computing platforms, specifically including the following steps:

[0055] Step S1: Collect the raw data to be processed; perform coarse sorting on the raw data to generate a locally redundant candidate set;

[0056] Specifically, a predefined data interface is used to access data sources; preset data acquisition rules are set, such as data type and data range; and raw multimodal long video training data is obtained through the data interface according to the data acquisition rules.

[0057] The acquired raw multimodal long video training data is coarsely sorted, including:

[0058] First, joint segmentation is performed based on the original multimodal long video training data to obtain a set of original data blocks. Specifically, the joint segmentation conditions are set according to the processing requirements of the current training task. The joint segmentation conditions include at least shot segmentation conditions, auxiliary text segmentation conditions, and duration segmentation conditions. Different dominant segmentation conditions can be adopted for different training tasks. For example, when the training task takes video content changes as the main processing basis, shot segmentation conditions are used as the dominant segmentation conditions, and segmentation is performed at each shot change position. When the duration of the single video block formed after segmentation exceeds the preset value... When the duration threshold is reached, a duration segmentation condition is introduced to further divide the current video block into smaller blocks to limit the length of the video block. When the training task takes the integrity of the auxiliary text as the main processing basis, the start and end times of the auxiliary text are used as the dominant segmentation condition, and segmentation is performed according to the boundaries of the auxiliary text. If there is no auxiliary text within the current time range, segmentation is performed according to the shot segmentation condition. When the duration of the video block formed after segmentation according to the auxiliary text boundary or shot boundary exceeds the preset duration threshold, a duration segmentation condition is introduced to further divide the current video block into smaller blocks to obtain the original data block set.

[0059] The duration segmentation condition is achieved by setting a preset duration parameter. The duration parameter is set according to different dominant needs. It is 1.2 times the median length of the dominant need data block. For example, in the shot segmentation condition, the length of all data blocks is counted, and 1.2 times the median is set as the duration parameter.

[0060] Then, description extraction is performed based on the original data block set to obtain a unified block description set: The video frame range, audio segment range, subtitle text range, optical character recognition text range, and source metadata corresponding to each original data block are read one by one; a modal identifier is set for the current original data block, indicating whether the original data block contains video, audio, subtitles, and optical character recognition results; subsequently, the start and end times of the current original data block are read to generate a time range; and the data source of the original data block is recorded as the source identifier; the video encoding format, resolution, frame rate, bit rate, audio sampling rate, and number of channels corresponding to the current original data block are extracted. One or more of the following are used to form a transcoding parameter field; the subtitle text content and optical character recognition text content falling within the time range of the current original data block are extracted to generate a subtitle fragment field and an optical character recognition fragment field, respectively; one or more of the following are counted in the current original data block: the number of video frames, audio duration, number of subtitle characters, number of optical character recognition characters, average brightness, and audio energy, to generate basic statistics; finally, the modality identifier, time range, source identifier, transcoding parameters, subtitle fragment, optical character recognition fragment, and basic statistics are written into the block description record corresponding to the current original data block in a unified order to generate a unified block description set;

[0061] Coarse screening is performed based on a unified block description set to obtain a locally redundant candidate set: For each data block, its coarse screening features are extracted. The coarse screening features include at least one or more of the following: downsampled hash values ​​of video frames within the data block, energy envelope fingerprints of audio segments, and locally sensitive hash values ​​of subtitles and optical character recognition text; The above coarse screening features are concatenated in a preset order to generate a coarse fingerprint of fixed length. The length of the coarse fingerprint is calibrated to 64 bits or 128 bits according to the target false merging rate.

[0062] Based on the calculated coarse fingerprint results and the preset local window width, the cosine similarity of the coarse fingerprints of each data block is calculated within the current window. If the coarse fingerprint similarity of two data blocks exceeds the preset coarse screening threshold, the corresponding data block pair is marked as a candidate redundant pair. After traversing all windows, transitive closure merging is performed on all candidate redundant pairs to form several local redundant candidate clusters. The data blocks in each candidate cluster are candidate redundancies to each other, resulting in a local redundant candidate set.

[0063] The preset local window width is calibrated based on the lens switching distribution and sampling frequency, and is preferably taken as the 75th percentile value of the lens boundary interval of the most recent 5 batches; the coarse screening threshold is initially set to 0.9 and is dynamically adjusted according to the false positive rate of historical batches. When the false positive rate is too high, the coarse screening threshold is increased, and when the false negative rate is too high, the coarse screening threshold is decreased.

[0064] Step S2: Generate compressed representations based on the locally redundant candidate set to obtain a cross-modal compressed representation set; aggregate data blocks based on the cross-modal compressed representation set to obtain a refined candidate cluster set;

[0065] Specifically, compressed representation generation based on locally redundant candidate sets includes:

[0066] Feature extraction is performed based on the local redundancy candidate set to obtain the joint feature vector: First, each data block in the local redundancy candidate set is sent to the processing pipeline on the graphics processor side, and the original features of different modes are extracted according to the modality identifier corresponding to each data block.

[0067] For video frames, sampled frames are extracted from the current data block at preset frame intervals, and each sampled frame is scaled to a uniform resolution and then input into a video feature extraction model, such as MobileNetV2, to obtain the frame-level features corresponding to each sampled frame. Then, average pooling is performed on each frame-level feature in time order to obtain the original feature vector of the video modality.

[0068] For an audio segment, the audio segment corresponding to the current data block is resampled to a uniform sampling rate, and the Mel frequency cepstral coefficients are extracted. Then, the data is input into an audio feature extraction model, such as YAMNet, to obtain the original feature vector of the audio modality.

[0069] For subtitle text, the subtitle segments corresponding to the current data block are concatenated into a text sequence in chronological order, and after word segmentation, they are input into a text feature extraction model, such as DistilBERT. Average pooling is performed on the output term features to obtain the original feature vector of the subtitle modality.

[0070] For optical character recognition text, the optical character recognition results corresponding to the current data block are concatenated into a text sequence in chronological order, and after word segmentation, they are input into the same text feature extraction model as the subtitle text. Average pooling is performed on the output term features to obtain the original feature vector of the optical character recognition modality.

[0071] After obtaining the original feature vectors of each modality, the original feature vectors of the video modality, audio modality, subtitle modality, and optical character recognition modality corresponding to the current data block are concatenated in order to form a joint feature vector. For missing modalities, the zero-value vectors of the corresponding dimensions are used to fill in the missing features.

[0072] Linear projection is performed based on the joint feature vector to compress the data into a preset low-dimensional space, resulting in a cross-modal compressed representation corresponding to the current data block. The above process is repeated for all data blocks in the local redundant candidate set to obtain a cross-modal compressed representation set. The compressed representation dimension is calibrated according to the principle that the cumulative fidelity after dimensionality reduction is not less than 95%, and can be set to 128 or 256.

[0073] Data block aggregation based on cross-modal compressed representation sets includes:

[0074] Similarity calculation is performed based on the cross-modal compressed representation set to obtain a sparse similarity set: on the GPU side, the similarity between all feature vectors is calculated by calculating cosine similarity; based on the approximate nearest neighbor index, a preset number of candidate nearest neighbor vectors are retrieved for each feature vector, and candidate nearest neighbor vectors with similarity lower than a preset first threshold are removed to generate a sparse similarity set;

[0075] The preset number can be set according to the total amount of data in the current batch or the estimated average size of fine candidate clusters. The preset number is equal to 1.5 to 2.5 times the average cluster size, with a lower limit of not less than 20 and an upper limit of not more than 300. The preset first threshold is set according to the effective similarity distribution of the current batch, and the value is the 75th percentile of the effective similarity of the current batch.

[0076] A connection graph is constructed based on a sparse similarity set to obtain a data block similarity graph: each data block is used as a graph node, and data block pairs with similarity reaching or exceeding the first threshold are used as connection edges between nodes to construct the data block similarity graph.

[0077] Based on the data block similarity graph, connected component partitioning is performed to obtain initial candidate clusters: the disjoint-set data structure algorithm is used to merge directly connected or indirectly connected through intermediate nodes into the same initial candidate cluster;

[0078] Based on the initial candidate clusters, a refined set of candidate clusters is obtained: After obtaining each initial candidate cluster, the mean vector is calculated for the feature vectors corresponding to the data blocks within each initial candidate cluster, and the mean vector is determined as the cluster center of the current initial candidate cluster; then the cosine similarity between the feature vectors corresponding to each data block within the cluster and the cluster center is calculated. When the similarity between a data block and the cluster center is lower than a preset second threshold, the data block is removed from the current initial candidate cluster, and the removed data block is recorded as a data block to be verified.

[0079] The second threshold is set to 0.8 times the average similarity within the current candidate cluster;

[0080] Similarity retrieval and connectivity merging are re-executed for all data blocks to be verified to generate supplementary candidate clusters: when the similarity between the cluster center of a supplementary candidate cluster and the cluster center of an existing candidate cluster reaches or exceeds a preset third threshold, the supplementary candidate cluster is merged into the corresponding existing candidate cluster; supplementary candidate clusters that do not meet the merging conditions are retained as independent candidate clusters.

[0081] The third threshold is set to 0.9 times the average similarity within the existing candidate clusters.

[0082] Record the cluster identifier, cluster center, list of data block indexes within the cluster, and cluster size for each cluster. Each cluster obtained is a fine candidate cluster. Each cluster contains several data blocks that are highly similar at the cross-modal semantic level. Record all the fine candidate clusters as a fine candidate cluster set.

[0083] Furthermore, when the number of data blocks contained in a candidate cluster exceeds the preset cluster size threshold, similarity retrieval, connectivity merging, and intra-cluster correction are recursively performed on the candidate cluster to split the candidate cluster into multiple sub-clusters. The cluster size threshold is set according to the current GPU memory capacity and the maximum allowable processing volume of a single batch aggregation task, such as setting it to 1000 to ensure the efficiency of subsequent processing steps.

[0084] Step S3: Perform propagation calculations based on the refined candidate cluster set to obtain redundant propagation records; perform propagation modeling based on the redundant propagation records to generate sample retention decision results;

[0085] Specifically, propagation computation based on a refined set of candidate clusters includes:

[0086] Similarity component calculation is performed based on refined candidate clusters to obtain similarity component data: within each refined candidate cluster, candidate data block pairs are selected as propagation calculation objects, and similarity component calculation is performed on the candidate data block pairs to calculate structural similarity, semantic similarity, and temporal overlap respectively.

[0087] The structural similarity is calculated as follows: read the coarse fingerprint results corresponding to each candidate data block, calculate the fingerprint matching ratio between the two, and normalize the fingerprint matching ratio to the structural similarity.

[0088] Semantic similarity is calculated based on cross-modal compressed representations. Specifically, the cross-modal compressed representations of the current candidate data blocks are read, the cosine similarity between them is calculated, and the cosine similarity is determined as the semantic similarity.

[0089] The temporal overlap is calculated based on the time range of each candidate data block pair. Specifically, the start and end times of the current candidate data block pair are read, and the proportion of the overlap duration of the two time intervals to the duration of the shorter data block is calculated. When there is no overlap between the two time intervals, the time interval between the two time intervals is calculated, and twice the duration of the shorter data block is determined as the adjacent decay interval. When the time interval is greater than or equal to the adjacent decay interval, the temporal proximity is recorded as 0. When the time interval is less than the adjacent decay interval, the temporal proximity is calculated by subtracting the ratio of the time interval to the adjacent decay interval from 1. The obtained temporal proximity is determined as the temporal overlap.

[0090] Based on the refined candidate clusters, the protection value is calculated to obtain the protection value of rare samples: First, based on all data blocks that enter the propagation calculation step in the current batch, the event label frequency table, scene label frequency table, text phrase frequency table and modality combination frequency table are generated respectively;

[0091] The event tag frequency table is generated as follows: the subtitle segments and optical character recognition segments corresponding to each data block are segmented into words and input into a semantic recognition model specifically trained for the event semantic recognition task to obtain event semantic tags corresponding to each text segment; the semantic recognition model is trained through historical labeled samples, in which the event categories corresponding to each text segment are pre-labeled, and the model version after training is recorded in the model configuration table; synonym tag normalization is performed on the event semantic tags output by the semantic recognition model, the frequency of each event semantic tag in all data blocks of the current batch is counted, and the event semantic tags in the bottom 10% of the frequency ranking are identified as low-frequency event tags;

[0092] The scene tag frequency table is generated as follows: scene recognition is performed on the video keyframes corresponding to each data block, and the scene tag corresponding to the current data block is obtained by using a visual semantic matching model (such as the CLIP model); the frequency of occurrence of scene tags corresponding to all data blocks is counted, and the scene tags in the bottom 10% of the frequency ranking are identified as rare scene tags.

[0093] The text phrase frequency table is generated as follows: the subtitle fragments and optical character recognition fragments corresponding to each data block are segmented into words, and Chinese word segmentation models (such as Jieba) are used to obtain candidate text phrases; after normalizing the candidate text phrases according to the synonym expression normalization table, the frequency of each candidate text phrase in all data blocks of the current batch is counted, and the text phrases in the last 10% of the frequency ranking are determined as long-tail text phrases.

[0094] The modality combination frequency table is generated as follows: the modality identifiers corresponding to each data block are combined and merged, the frequency of each modality combination in all data blocks of the current batch is counted, and the modality combinations that rank in the bottom 10% of the frequency ranking are determined as low-coverage modality combinations.

[0095] After obtaining low-frequency event labels, rare scene labels, long-tail text phrases, and low-coverage modal combinations, the set of low-frequency event labels, rare scene labels, long-tail text phrases, and low-coverage modal combinations that the current candidate data block matches is determined. The reciprocal of the frequency of each label, phrase, or combination in the current batch is taken and summed to obtain the event inverse frequency value, scene inverse frequency value, text inverse frequency value, and modal inverse frequency value. Finally, the event inverse frequency value, scene inverse frequency value, text inverse frequency value, and modal inverse frequency value are weighted by 0.35, 0.25, 0.25, and 0.15, respectively, and summed to obtain the rare sample protection value for the current candidate data block. ;

[0096] The redundancy propagation score is obtained by weighting similarity component data and rare sample protection values: The redundancy propagation score for the current candidate data block is calculated by weighting structural similarity, semantic similarity, temporal overlap, and rare sample protection values. The calculation formula is as follows:

[0097] ;

[0098] in, Scoring for redundancy propagation, For structural similarity, For semantic similarity, For time overlap, The value represents the protection value for rare samples, and a, b, c, and d represent the structural similarity weight, semantic similarity weight, temporal overlap weight, and rare sample protection value weight, respectively.

[0099] In this formula, the structural similarity weight 'a' is set to 0.25, the semantic similarity weight 'b' to 0.35, the temporal overlap weight 'c' to 0.15, and the rare sample protection value weight 'd' to 0.25. The semantic similarity weight is higher than the structural similarity weight and the temporal overlap weight, which is used to enhance the determination of cross-modal semantic repetition. The rare sample protection value is introduced as a subtraction term in the formula to reduce the redundancy propagation score of data block pairs with retention value.

[0100] Repeat the above weighted calculation process for all candidate data block pairs in the current batch, and record the structural similarity, semantic similarity, temporal overlap, rare sample protection value and redundancy propagation score of each candidate data block pair as a propagation score record;

[0101] Statistical analysis is performed based on redundancy propagation scores, and a redundancy determination threshold is set. The redundancy propagation scores of all candidate data blocks in the current batch are aggregated to generate a score set. Then, the scores in the score set are sorted by value, and the median is taken. Subsequently, the absolute difference between each score value and the median is calculated to obtain an absolute deviation set. The absolute deviation values ​​in the absolute deviation set are then sorted by value, and the median value is taken as the median absolute deviation of the redundancy propagation score for the current batch. After obtaining the median and median absolute deviation, a weighted average is performed to obtain the redundancy determination threshold for the current batch. The calculation formula is as follows:

[0102] ;

[0103] in, The calculated redundancy threshold, This represents the median of all redundancy propagation scores for the current batch. This represents the median absolute deviation of the total redundancy propagation score of the current batch relative to the median. Indicates the threshold generation coefficient;

[0104] The threshold generation coefficient k is set based on the risk of accidental deletion and the risk of missed detection in the current batch. When the current batch has a higher requirement for retaining rare samples, the value of k is increased; when the current batch has a higher requirement for redundancy reduction, the value of k is decreased. Preferably, k is between 0.8 and 1.5.

[0105] The candidate data block pair identifier, candidate data block identifier, rare sample protection value, propagation score record and redundancy judgment threshold calculated for the same batch are associated and recorded as a redundancy propagation record;

[0106] Propagation modeling is performed based on the obtained redundant propagation records, including:

[0107] A graph structure is generated based on the redundancy propagation record to obtain a redundancy propagation graph: each data block is used as a graph node, and candidate data block pairs are used as candidate connections. When the redundancy propagation score of a candidate data block pair reaches or exceeds the corresponding redundancy judgment threshold, a propagation connection edge is established between the two graph nodes corresponding to the candidate data block pair, and the redundancy propagation score is written into the edge weight of the corresponding propagation connection edge; otherwise, no propagation connection edge is established, and the candidate data block pair is recorded as a non-propagation candidate pair.

[0108] Based on the generated redundant propagation graph, connected component partitioning is performed to generate propagation subgraphs: For each propagation subgraph, firstly, the rare sample protection value corresponding to each data block in the current propagation subgraph in the candidate data block pairs formed with other data blocks is calculated, and the maximum value among all rare sample protection values ​​corresponding to the current data block is determined as the block-level protection reference value of the data block; then, the time range, modal identifier, and source identifier corresponding to each data block in the current propagation subgraph are read, and sorted according to the block-level protection reference value, and the data block with the highest block-level protection reference value is determined as the representative reserved data block; when two or more data blocks have the same block-level protection reference value, they are further sorted according to time range integrity, modal integrity, and source stability in that order, and the data block with the best sorting result is determined as the representative reserved data block;

[0109] Using the representative retained data block as a constraint, the current propagation subgraph is subjected to a deletion decision to obtain the sample retention decision result: For all data blocks other than the representative retained data block, the edge weights, structural similarity, and semantic similarity of the corresponding propagation connection edges between the current data block and the representative retained data block are read; when there is a propagation connection edge between the current data block and the representative retained data block, and the edge weight of the propagation connection edge meets the deletion confirmation condition, and the structural similarity and semantic similarity both meet the preset deletion similarity conditions, the current data block is determined as a redundant deleted data block; when any condition is not met, the current data block is determined as an additional retained data block.

[0110] The deletion confirmation condition is: the weight of the propagation connection edge between the current data block and the representative retained data block is in the top 50% of the weight of all propagation connection edges in the current propagation subgraph; the preset deletion similarity condition is: the structural similarity between the current data block and the representative retained data block is not less than 0.8, and the semantic similarity is not less than 0.85.

[0111] After determining the deletion of all data blocks within the current propagation subgraph, a mapping relationship is established between each redundant deleted data block and the representative retained data block. The identifiers of the representative retained data block, the additional retained data block, and the redundant deleted data block corresponding to the current propagation subgraph are recorded. Subsequently, the retention decision results corresponding to all propagation subgraphs are summarized, and the decision type, representative mapping identifier, and decision basis for each data block are written into the sample retention decision record. The decision type includes at least representative retention, additional retention, and redundant deletion. Finally, all sample retention decision records are summarized to generate the sample retention decision result.

[0112] Step S4: Generate an affinity matrix based on the sample retention decision results to obtain the device affinity matrix; perform task allocation based on the device affinity matrix to obtain the task allocation results;

[0113] Specifically, affinity matrix generation is performed based on the sample retention decision results, including:

[0114] Based on the sample retention decision results, subtasks are decomposed to obtain a set of preprocessing subtasks: Subtasks are decomposed according to the decision type in the sample retention decision results. Taking representative retention, additional retention, and redundancy removal as examples, the decomposition results in representative retention subtasks, additional retention subtasks, and redundancy removal subtasks. It can also include index update subtasks and data migration subtasks. The decision type is set according to the usage requirements. The decomposed subtasks are associated and recorded as a set of preprocessing subtasks.

[0115] Device affinity calculation is performed based on the preprocessed subtask set to obtain the device affinity matrix. The basic information of the subtasks is analyzed based on the obtained preprocessed subtasks, including subtask type, input data volume, input position, and output position.

[0116] The subtask type is derived from the decision type; the input data volume is derived from the set of data blocks associated with the current subtask; the input location is derived from the current storage location or device location of the input data of the current subtask; and the output location is derived from the processing target of the subtask.

[0117] Next, the expected acceleration benefits, data migration costs, available computing power, and current queue pressure of each preprocessing subtask on the central processing unit, graphics processing unit, data processor, and field-programmable gate array are calculated separately. Among them, the expected acceleration benefits are calculated based on the average execution time of the same subtask type on different devices in the historical execution records; the data migration costs are calculated based on the current subtask input data volume and the data transmission path latency between the input location and the target execution device location; the available computing power is calculated based on the current device's idle computing resources ratio and bandwidth idle ratio; and the current queue pressure is calculated based on the current device's corresponding task queue length and average waiting time.

[0118] Subsequently, a weighted calculation is performed on the expected acceleration benefits, data migration costs, available computing power on the device, and current queue pressure to obtain the device affinity score of the current preprocessing subtask on each device. The calculation formula is as follows:

[0119] ;

[0120] in, For the p-th preprocessing subtask, give it a device affinity score on the d-th device. Let p be the expected speedup gain of the p-th preprocessing subtask on the d-th device. For the available computing power of the device, The data migration cost of the p-th preprocessing subtask on the d-th device. To give the current queue pressure of device d, , , , These are the weights for device affinity score, expected acceleration benefits, available computing power of the device, and data migration cost.

[0121] Each weight is set according to the processing objective of the current batch: when the current batch prioritizes processing throughput, the weight for expected acceleration benefits is increased. And the available computing power weight of the device The value of is determined by increasing the data migration cost weight when the current batch prioritizes cross-device migration suppression. The value of is determined by increasing the pressure weight of the current queue when the current batch prioritizes device load balancing. The values ​​of w1, w2, w3, and w4 are preferably set to 0.35, 0.20, 0.30, and 0.15.

[0122] After obtaining the device affinity scores of each preprocessing subtask on each device, all device affinity scores corresponding to the same preprocessing subtask are arranged by device type to generate the device affinity vector corresponding to that preprocessing subtask. Then, the device affinity vectors corresponding to all preprocessing subtasks are summarized and arranged by subtask identifier to generate a device affinity matrix. In the device affinity matrix, the rows correspond to each preprocessing subtask, the columns correspond to the central processing unit, graphics processing unit, data processor, and field-programmable gate array, and each matrix element corresponds to the device affinity score of the corresponding preprocessing subtask on the corresponding device.

[0123] Task allocation based on device affinity matrix includes:

[0124] Device priority ranking is performed based on the device affinity matrix to determine the target execution device: The device affinity score corresponding to each preprocessing subtask is read from the device affinity matrix; for each preprocessing subtask, it is sorted from high to low according to the device affinity score corresponding to the CPU, GPU, data processor, and FPGA, and the device with the highest device affinity score is determined as the priority execution device for the current preprocessing subtask; when two or more devices have the same device affinity score, they are further sorted according to data migration cost from low to high and current queue pressure from low to high, and the device with the optimal sorting result is determined as the target execution device for the current preprocessing subtask.

[0125] Task sorting and batch division based on target execution device: After determining the target execution device for each preprocessing subtask, preprocessing subtasks corresponding to the same target execution device are summarized, and the execution order is arranged according to the input position, input data volume, and output position of each preprocessing subtask; when multiple preprocessing subtasks correspond to the same target execution device, and their input positions are consistent and meet the conditions for merging execution, these multiple preprocessing subtasks are merged into the same execution batch; when multiple preprocessing subtasks correspond to the same target execution device but have different input positions, or when the input data volume after merging exceeds the single batch processing limit of the current device, these multiple preprocessing subtasks are split into multiple independent execution batches.

[0126] Finally, write the subtask identifier, target execution device identifier, execution batch identifier, data transfer identifier, and queue order to all preprocessing subtasks, and summarize all the written results to generate task allocation results.

[0127] Step S5: Perform data preprocessing based on the task allocation results to obtain the preprocessing execution results;

[0128] Specifically, parallel preprocessing is performed based on the task allocation results, including:

[0129] First, read the target execution device identifier, execution batch identifier, data transfer identifier, queue order, input data block identifier set, input position, and output position corresponding to each preprocessing subtask in the task allocation result; then, according to the target execution device identifier, distribute all preprocessing subtasks to the execution queues corresponding to the central processing unit, graphics processing unit, data processor, and field programmable gate array respectively, and start parallel execution according to the queue order;

[0130] For preprocessing subtasks with data migration flags, the data migration operation is performed first. Specifically, based on the input location and the target execution device location, the corresponding data transmission path is invoked to transfer the input data block, intermediate results, or mapping record corresponding to the current preprocessing subtask to the storage area corresponding to the target execution device. After the data migration is completed, a migration completion flag is written, and the current preprocessing subtask is set to an executable state. For preprocessing subtasks whose input location is the same as the target execution device location, they are directly entered into the execution queue of the target execution device.

[0131] On each device side, the preprocessing subtasks entering the execution queue are processed according to their subtask types:

[0132] For the representative retention subtask, the representative retention data block is written into the standard sample result set, and the corresponding fine candidate cluster identifier, propagation subgraph identifier and source identifier are recorded;

[0133] For the additional retention subtask, the additional retention data block is recorded in the supplementary sample results, and the association between it and the representative retention data block is recorded;

[0134] For the redundancy reduction subtask, the redundancy reduction data block identifier, the corresponding representative retained data block identifier, and the reduction basis are written into the reduction mapping record;

[0135] For the index update subtask, the update represents the mapping index between the retained data blocks and the redundant removed data blocks, the fine-grained candidate cluster index, and the propagation subgraph index;

[0136] For the data migration subtask, complete the transmission and writing of input data and intermediate results between devices;

[0137] For log write-back subtasks, record the data block identifier, decision type, execution device, execution start time, execution end time, and execution status corresponding to the current preprocessing subtask.

[0138] During execution, the task execution status of each device is monitored in real time. Specifically, this involves: reading the number of completed tasks, the number of failed tasks, the remaining queue length, and the execution time for each device in the current batch; when a preprocessing subtask is executed successfully, the current preprocessing subtask is set to the execution completed state, and the corresponding output result is written to the preset output location; when a preprocessing subtask fails, the failure reason code is recorded, and the current preprocessing subtask is set to the execution failed state; when the remaining queue length for a device exceeds the preset queue threshold, or the execution time of a preprocessing subtask exceeds the preset time threshold, the corresponding preprocessing subtask is recorded as a task to be adjusted, and an execution exception record is written.

[0139] After all preprocessing subtasks in the current batch have been executed, the output results of each device are summarized to generate the preprocessing execution results. The preprocessing execution results include at least the standard sample result set, the supplementary sample result set, the deleted mapping records, the index update results, the log write-back results, and the execution exception records. Finally, the above results are associated and recorded according to the current batch identifier to obtain the preprocessing execution results.

[0140] like Figure 2 The comparison chart of the effects of this invention and traditional methods is shown in the figure. The black bars in the figure represent the effect of this invention, and the gray bars represent the effect of traditional methods. This invention introduces a rare sample protection value in the redundancy propagation calculation, making it easier to retain sample pairs containing rare elements, thereby improving the rare sample retention capability compared to traditional methods. By coupling deduplication and scheduling in the same round of processing, serial waiting is avoided, and subtasks are allocated to the optimal devices using the expected acceleration benefits, thereby improving preprocessing efficiency. By adding the current queue pressure item to the device affinity score and supporting dynamic backpressure control, tasks are moved from high-load devices to idle devices, thus optimizing resource utilization compared to traditional methods.

[0141] Example 2: After obtaining the preprocessing execution results, backpressure control is performed based on the preprocessing execution results to dynamically optimize the accuracy of preprocessing task allocation for subsequent batches. The process data of the current batch's preprocessing execution results is monitored, and backpressure control is performed, including:

[0142] When the device queue level corresponding to the graphics processor exceeds the preset high water level threshold, the weight of data migration cost and queue pressure in the device affinity calculation is increased, and the device affinity score corresponding to the graphics processor is reduced; if the length of the graphics processor task queue reaches 85% of the queue capacity, the weight of data migration cost is adjusted from 0.30 to 0.40, the weight of queue pressure is adjusted from 0.15 to 0.25, and the device affinity score corresponding to the graphics processor is multiplied by a suppression coefficient of 0.85.

[0143] When the data back-to-work latency exceeds the preset back-to-work latency threshold, the preprocessing subtasks involving cross-device movement will be re-split and preferentially assigned to the device corresponding to the current location of the input data for execution. If the average data back-to-work latency from the graphics processor to the central processing unit in the current batch reaches 18ms, exceeding the preset back-to-work latency threshold of 15ms, the joint subtask that originally needed to be back-to-work from the graphics processor to the central processing unit for mapping write-back will be split into a graphics processor-side feature write-back subtask and a central processing unit-side index update subtask, and the former will be retained for execution on the graphics processor.

[0144] When the replay results of the accidental deletion show that the current batch's accidental deletion rate is higher than the preset accidental deletion rate threshold, the weight corresponding to the rare sample protection value is increased, and the redundancy judgment threshold generation coefficient is increased; if the current batch's accidental deletion rate reaches 4.8%, exceeding the preset accidental deletion rate threshold of 3%, the weight of the rare sample protection value is adjusted from 0.25 to 0.35, and the redundancy judgment threshold generation coefficient k is adjusted from 1.0 to 1.2.

[0145] When the average waiting time of a task exceeds the preset waiting time threshold, the number of target execution tasks corresponding to the currently high-load device is reduced, and some preprocessing subtasks in subsequent batches are adjusted to be executed on idle devices; if the average waiting time of the task queue corresponding to the data processor reaches 220ms, exceeding the preset waiting time threshold of 150ms, the 20 preprocessing subtasks originally planned to be allocated to the data processor in the next batch are reduced to 12, and the remaining 8 preprocessing subtasks are adjusted to be executed on central processing units and field programmable gate arrays with an idle rate of more than 40%.

[0146] Write the adjusted weight parameters, threshold parameters, and task allocation parameters into the corresponding scheduling parameter record for the next batch.

[0147] Example 3: This example provides an electronic device, including: a memory and a processor; the memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions to implement the method proposed in the above examples.

[0148] The electronic device can be a terminal, comprising a processor, memory, communication interface, display screen, and input device connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input device can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.

[0149] This embodiment provides a storage medium on which a computer program is stored. When the program is executed by a processor, it implements the method proposed in the above embodiment. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0150] Example 4: The training batch of the "dangerous action recognition model" to be built on a short video platform was selected as the processing object; the batch corresponds to 120 original long video data segments with a total duration of 14,400 seconds, which are derived from 8 urban public area surveillance videos and 12 user-uploaded videos; each video segment is accompanied by an audio stream, automatically generated subtitle text, and on-screen optical character recognition text; the heterogeneous computing platform includes 1 central processing unit, 2 graphics processing units, 1 data processor and 1 field-programmable gate array, where the central processing unit is responsible for index write-back and log merging, the graphics processing unit is responsible for cross-modal feature extraction and similarity calculation, the data processor is responsible for batch transfer and parallel writing, and the field-programmable gate array is responsible for coarse fingerprint calculation.

[0151] First, joint segmentation was performed on 120 original long video segments: Since the training task for this batch primarily focuses on changes in video content, shot segmentation was determined as the dominant segmentation condition, with supplementary text segmentation as a secondary alignment condition, and duration segmentation as a length constraint. After calculating the initial shot segment lengths for this batch, the median was 10 seconds, so the duration parameter was set to 12 seconds. After segmentation, a total of 1,260 original data blocks were obtained. Subsequently, a unified data processing protocol was generated for each of the 1,260 original data blocks. Block description record; taking four data blocks as an example: Data block B017 comes from monitoring source C2, with a time range of 00:03:12 to 00:03:20, a resolution of 1920×1080, a frame rate of 25 frames per second, an audio sampling rate of 16kHz, a subtitle clip of "Someone climbed over the fence", an optical character recognition clip of "East Gate Passage", a video frame count of 200, an audio duration of 8 seconds, and a subtitle character count of 7; Data block B018 comes from user source U5, with a time range of 00:01:08 to At 00:01:16, the subtitle fragment is "Man climbs over fence," and the optical character recognition fragment is empty; data block B019 comes from monitoring source C6, with a time range of 00:09:41 to 00:09:49, and the subtitle fragment is "Pedestrian crosses the guardrail"; data block B020 comes from monitoring source C2, with a time range of 00:03:21 to 00:03:29, and the subtitle fragment is "Someone climbs over the guardrail"; after the unified block description is generated, coarse fingerprints are calculated for each data block, including B017 and B018. The coarse fingerprint similarities of B019, B020, and B017 were 0.93, 0.91, and 0.92, respectively, all higher than the coarse screening threshold of 0.90. Therefore, the three were grouped into the same local redundancy candidate cluster L12. The coarse fingerprint similarity of B020 and B017 was 0.87, lower than the coarse screening threshold, so they were not included in L12 during the coarse screening stage. Finally, after coarse sorting, 1,260 original data blocks formed 146 local redundancy candidate clusters, covering a total of 832 candidate data blocks. The remaining 428 data blocks were directly retained as non-candidate blocks.

[0152] A cross-modal compressed representation was generated for 832 candidate data blocks within 146 locally redundant candidate clusters. Taking L12 as an example, video, audio, subtitle, and optical character recognition features were extracted from B017, B018, B019, and B020, which subsequently entered the verification process. These features were then concatenated into a joint feature vector and compressed into a 128-dimensional cross-modal compressed representation. The semantic cosine similarity between B017 and B018 was set to 0.94, between B017 and B019 to 0.92, and between B017 and B020 to 0.79. After constructing a similarity graph for all vectors within the L12 cluster, the initial candidate cluster C12 contained a total of 4 data blocks. Next, calculate the cosine similarity between the cluster center and each data block, obtaining: B017 is 0.95, B018 is 0.93, B019 is 0.91, and B020 is 0.72. Since the average similarity within C12 is 0.8775, the second threshold is set to 0.8775 × 0.8 = 0.702. B020 is higher than 0.702, so it is not removed in this round. Then look at another initial candidate cluster C23, which contains 5 data blocks with an average similarity of 0.88, corresponding to a second threshold of 0.704. Among them, data block B331 has a similarity of only 0.66 with the cluster center, so B331 is removed and recorded as a data block to be verified. After re-searching all the data blocks to be verified, B331 was reconnected and merged with two other originally scattered data blocks to form a supplementary candidate cluster SC7. Finally, this batch was refined from 146 locally redundant candidate clusters to 121 fine candidate clusters, with cluster sizes ranging from 2 to 17 and an average cluster size of 6.88.

[0153] Propagation calculations are performed on the refined candidate clusters: Taking candidate data block pairs B017 and B018 in the refined candidate cluster C12 as an example, after reading their coarse fingerprints, the structural similarity is calculated to be 0.90; after reading their 128-dimensional compressed representations, the semantic similarity is calculated to be 0.94; after reading their time ranges, they are located in videos from different sources and there is no temporal overlap, with the time interval recorded as 124 seconds; the shorter data block has a duration of 8 seconds, so the nearest decay interval is 16 seconds. Since 124 seconds is greater than 16 seconds, the time overlap is recorded as 0. Then, low-frequency event labels, rare scene labels, long-tail text phrases, and low-coverage modal combinations are statistically analyzed for all candidate blocks in this batch. The results show that the event label "climbing over the guardrail" appears 12 times in 832 candidate blocks, with an inverse frequency of 1 / 12 = 0.0833; the scene label "nighttime East Gate passage" appears 8 times, with an inverse frequency of 0.125; the text phrase "East Gate passage" appears 5 times, with an inverse frequency of 0.2; and the modal combination "video + subtitle + optical character recognition" appears 104 times, with an inverse frequency of... The value is 0.0096. After weighting by 0.35, 0.25, 0.25, and 0.15, the rare sample protection value of this data block pair is: 0.35×0.0833+0.25×0.125+0.25×0.2+0.15×0.0096=0.1118. The redundancy propagation score is calculated with weights a=0.25, b=0.35, c=0.15, and d=0.25. The propagation score of this data block pair is: 0.25×0.90+0.35×0.94+0.15×0-0.25×0.1118=0.5261.

[0154] Continuing with the calculation for another set of data blocks B017 and B019 in C12: Assuming a structural similarity of 0.88, a semantic similarity of 0.92, a temporal overlap of 0, and a rare sample protection value of 0.1040, the propagation score is 0.25×0.88 + 0.35×0.92 + 0.15×0 - 0.25×0.1040 = 0.5150. Then, the calculation is completed for all 2,964 candidate data block pairs in the same batch of refined candidate clusters, yielding the propagation score. The data is divided into sets; after sorting the set, the median is 0.472 and the median absolute deviation is 0.061; the threshold generation coefficient k=1.2 is taken, and the redundancy judgment threshold T=0.472+1.2×0.061=0.5452 is obtained accordingly; it can be seen that the propagation score of 0.5261 corresponding to B017 and B018 is lower than the threshold, and the propagation score of 0.5150 corresponding to B017 and B019 is also lower than the threshold. Therefore, no propagation connection edge is established for these two sets of data blocks. Looking at another data block pair, B441 and B446, in the propagation subgraph, let their structural similarity be 0.96, semantic similarity be 0.97, temporal overlap be 0.85, and protection value be 0.020. Then the propagation score is 0.25×0.96+0.35×0.97+0.15×0.85-0.25×0.02=0.7010, which is higher than the threshold of 0.5452. Therefore, a propagation connection edge is established and the edge weight is recorded as 0.7010. After forming a redundant propagation graph, a total of 89 propagation subgraphs are obtained.

[0155] In the propagation modeling phase, representative retained data blocks are determined for each propagation subgraph. Taking propagation subgraph G31 as an example, it contains data blocks B441, B446, B452, and B460. After calculating the maximum protection value among the candidate pairs associated with each of the four data blocks, the values ​​are: B441 is 0.036, B446 is 0.021, B452 is 0.064, and B460 is 0.030. Therefore, B452, which has the highest block-level protection reference value, is determined as the representative retained data block. Subsequently, deletion judgments are performed on the remaining three data blocks one by one. The propagation edge weight between B441 and B452 is 0.688, which is in the top 50% of the edge weight ranking within G31. Furthermore, the structural similarity is 0.91 and the semantic similarity is 0.93, satisfying the deletion confirmation condition and the deletion similarity condition. Therefore, B441 is designated as a redundant pruning data block; the edge weight between B446 and B452 is 0.612, with a structural similarity of 0.85 and a semantic similarity of 0.88, and is also designated as a redundant pruning data block; the edge weight between B460 and B452 is 0.571, with a structural similarity of 0.78 and a semantic similarity of 0.90. Since the structural similarity does not reach 0.8, B460 is determined as an additional retained data block; after modeling all 89 propagation subgraphs, the sample retention decision results for this batch are as follows: 89 representative retained data blocks, 173 additional retained data blocks, and 570 redundant pruning data blocks; thus, of the 832 data blocks that originally entered the candidate processing chain, 262 entered the retention set and 570 entered the pruning mapping set.

[0156] Based on the sample retention decision results, a set of preprocessing subtasks is generated: 89 representative retention data blocks are divided into 89 representative retention subtasks, 173 additional retention data blocks are divided into 173 additional retention subtasks, and 570 redundant pruning data blocks are divided into 570 redundancy pruning subtasks. Then, 89 index update subtasks and 47 data migration subtasks are generated, resulting in a total of 968 preprocessing subtasks in this batch. Subsequently, the computing device affinity of each subtask is scored; taking representative retention subtask T052 as an example... Its input data size is 1.6GB, the input location is in the graphics processor's video memory area, and the output location is the training sample storage pool; according to historical execution records, the expected speedup gains of T052 on the central processing unit, graphics processor, data processor, and field-programmable gate array are 0.42, 0.91, 0.58, and 0.33, respectively; the available computing power of the device is 0.76, 0.63, 0.81, and 0.55, respectively; and the data migration cost is 0.12, 0.03, 0.18, and 0.25, respectively; when The front queue pressures were set to 0.31, 0.47, 0.28, and 0.19, respectively. Weighted scores of 0.35, 0.20, 0.30, and 0.15 were applied to calculate device affinity scores. T052 scored 0.2855, 0.4310, 0.2960, and 0.1420 on the four device categories, respectively. Therefore, the graphics processor (GPU) was identified as the target execution device. Taking the index update subtask T611 as an example, its input data size was only 28MB, and its output location was in the central processing unit (CPU) side index library. The CPU received the highest score after calculation, thus identifying it as the target execution device. After completing the device affinity calculations for all 968 subtasks, a 968×4 device affinity matrix was generated. Further sorting resulted in the GPU handling 402 subtasks, the CPU handling 311 subtasks, the data processor handling 181 subtasks, and the field-programmable gate array (FPGA) handling 74 subtasks. Execution batches were then divided based on input position consistency and batch size limits, ultimately forming 86 execution batches.

[0157] Parallel preprocessing is performed based on the task allocation results: First, 86 execution batches are distributed to the corresponding execution queues according to the target execution devices, with 34 batches allocated to the graphics processor queue, 27 batches to the central processing unit queue, 17 batches to the data processor queue, and 8 batches to the field-programmable gate array queue. For the 47 subtasks with data migration markers, inter-device migration is performed first. Taking data migration subtask M09 as an example, it needs to write 0.9GB of intermediate results from the graphics processor's video memory to the central processing unit's memory. The actual migration latency is 11 milliseconds, which is lower than the 15 millisecond migration control threshold set for this batch. Therefore, a migration completion marker is written and it is converted to an executable state. Subsequently, each device executes in parallel according to the batches: the representative reserved subtask writes 89 data blocks to the standard sample result set; the additional reserved subtask writes 173 data blocks to the supplementary sample. The result set; the redundancy pruning subtask writes 570 "pruned block - representative block" mappings into the pruning mapping record; the index update subtask updates 89 cluster indexes and propagation subgraph indexes; the log write-back subtask records the device, start and end times, and status of all 968 subtasks; the actual execution results of this batch are: 964 subtasks completed, 4 subtasks failed, with failure reason codes of 2 "source block verification failure" and 2 "target storage write timeout"; the average batch execution time is 4.8 seconds, with the longest batch on the graphics processor side being 6.1 seconds and the longest batch on the CPU side being 5.4 seconds; after execution, the preprocessed execution results are summarized, including 262 standard sample result sets, 173 supplementary sample result sets, 570 pruning mapping records, 89 index update records, 968 log write-back records, and 4 execution exception records.

[0158] The batch results were compared with a control scheme that used deduplication before scheduling: Under the same original data scale, the control scheme ultimately retained 247 samples, including 11 samples of low-frequency dangerous actions that were mistakenly deleted; in this embodiment, the final effective retained samples entering the training pool were 262, including all samples of low-frequency dangerous actions, and the total amount of cross-device data transfer decreased from 312GB in the control scheme to 226GB, a reduction of 27.56%; the total preprocessing time decreased from 38.4 seconds to 29.7 seconds; and the average idle waiting time of the graphics processor decreased from 6.8 seconds to 3.9 seconds. Therefore, in this specific scenario, by continuing to use the redundancy propagation results for sample retention decisions and device affinity allocation, redundancy compression, rare sample protection, and heterogeneous device scheduling optimization can be achieved simultaneously.

[0159] The foregoing has shown and described the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The embodiments and descriptions in the specification are merely preferred examples and are not intended to limit the invention. Various changes and modifications can be made to the invention without departing from its spirit and scope, and all such changes and modifications fall within the scope of the present invention as claimed. The scope of protection of the present invention is defined by the appended claims and their equivalents.

[0160] Based on the foregoing description in conjunction with the accompanying drawings, those skilled in the art will understand that the embodiments of this application can also be implemented by software programs. Therefore, this application also provides a computer-readable storage medium. This computer-readable storage medium stores computer-readable instructions thereon, which, when executed by one or more processors, implement the method described above in conjunction with the accompanying drawings.

[0161] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computing device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0162] It should be noted that although the operations of the method of this application are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. On the contrary, the steps depicted in the flowchart can be performed in a different order. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.

[0163] It should be understood that when the terms "first," "second," "third," and "fourth," etc., are used in the claims, specification, and drawings of this application, they are used only to distinguish different objects and not to describe a specific order. The terms "comprising" and "including" as used in the specification and claims of this application indicate the presence of the described features, integrals, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components, and / or collections thereof.

[0164] It should also be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the application. As used in this specification and claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in this specification and claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes such combinations.

[0165] Although the embodiments of this application are described above, the content is merely an example adopted for the purpose of facilitating understanding of this application and is not intended to limit the scope and application scenarios of this application. Any person skilled in the art described in this application may make any modifications and changes in the form and details of the implementation without departing from the spirit and scope disclosed in this application, but the scope of patent protection of this application shall still be determined by the scope defined in the appended claims.

Claims

1. A method for preprocessing redundant data for heterogeneous computing platforms, characterized in that, include: Collect the long video data to be processed to obtain the raw data; The original data is coarsely sorted to generate a locally redundant candidate set; The locally redundant candidate set is compressed and characterized to generate a cross-modal compressed characterization set; Data block aggregation is performed based on the cross-modal compressed representation set to generate a refined candidate cluster set; Propagation calculations are performed on the refined candidate cluster set to obtain redundant propagation records; propagation modeling is then performed based on the redundant propagation records to generate sample retention decision results. Based on the sample retention decision results, affinity matrix generation is performed to obtain the device affinity matrix; Task allocation is performed based on the device affinity matrix to obtain the task allocation result; Based on the task allocation results, data preprocessing is performed to obtain the preprocessing execution results.

2. The method according to claim 1, characterized in that, The original data is coarsely sorted to generate a locally redundant candidate set, including: Based on the processing requirements of the training task, set joint splitting conditions, perform joint splitting on the original data, and obtain the original data block set; A unified block description set is obtained by extracting descriptions from the original data block set. A coarse screening is performed based on the unified block description set to obtain a locally redundant candidate set.

3. The method according to claim 1, characterized in that, The locally redundant candidate set is compressed to generate a cross-modal compressed representation set, including: Feature extraction is performed based on a set of locally redundant candidates to obtain a joint feature vector; A linear projection is performed based on the joint feature vector to obtain the cross-modal compressed representation of the current data block.

4. The method according to claim 1, characterized in that, Data block aggregation is performed based on the cross-modal compressed representation set to generate a refined candidate cluster set, including: Similarity calculation is performed based on cross-modal compressed representation sets to obtain sparse similarity relation sets; A connection graph is constructed based on a sparse similarity set to obtain a data block similarity graph; Connectivity component partitioning is performed based on the data block similarity graph to obtain initial candidate clusters; A refined set of candidate clusters is obtained by further refining the initial candidate clusters.

5. The method according to claim 4, characterized in that, The selection process is based on the initial candidate clusters, including: Cluster centers are obtained by calculating the cluster centers based on the initial candidate clusters; Similarity filtering is performed based on cluster centers to obtain the data blocks to be verified. Similarity retrieval and connectivity merging are re-executed for all data blocks to be verified to generate supplementary candidate clusters. All the obtained supplementary candidate clusters are recorded as a fine candidate cluster set.

6. The method according to claim 1, characterized in that, Propagation calculations are performed on the refined candidate cluster set to obtain redundant propagation records, including: Similarity component calculation is performed based on refined candidate clusters to obtain similarity component data; Protection values ​​are calculated based on refined candidate clusters to obtain protection values ​​for rare samples; A redundancy propagation score is obtained by weighting similarity component data and rare sample protection values. Statistical analysis is performed based on redundancy propagation scores, and a redundancy judgment threshold is set. The rare sample protection value, propagation score record, and redundancy judgment threshold are associated and recorded as a redundancy propagation record.

7. The method according to claim 6, characterized in that, Based on the refined candidate clusters, the protection value is calculated to obtain the protection value of rare samples, including: Pre-trained models are used to identify features of fine-grained candidate clusters, resulting in event label frequency tables, scene label frequency tables, text phrase frequency tables, and modality combination frequency tables. The inverse frequencies of the event label frequency table, scene label frequency table, text phrase frequency table, and modality combination frequency table are calculated separately, and the inverse frequencies are weighted and summed to obtain the rare sample protection value.

8. The method according to claim 1, characterized in that, Based on the redundant propagation records, propagation modeling is performed to generate sample retention decision results, including: A redundant propagation graph is obtained by generating a graph structure based on redundant propagation records. Based on the generated redundant propagation graph, connected component partitioning is performed to generate propagation subgraphs; Using the representative data block as a constraint, the current propagation subgraph is reduced to obtain the sample retention decision result.

9. The method according to claim 1, characterized in that, Based on the sample retention decision results, affinity matrix generation is performed to obtain the device affinity matrix, including: Based on the sample retention decision results, subtasks are split into preprocessing subtask sets; Device affinity calculations are performed based on the preprocessed subtask set to obtain the device affinity matrix.

10. The method according to claim 1, characterized in that, Task allocation is performed based on the device affinity matrix to obtain task allocation results, including: Device priority ranking is performed based on device affinity matrix to determine the target execution device; Tasks are sorted and divided into execution batches based on the target execution device; The target execution device and execution batch are associated and recorded as task allocation results.