Multi-modal data digital processing process monitoring method

By identifying and splitting multimodal data, combining clustering analysis and kernel density estimation, and dynamically adjusting the resource allocation ratio, the problems of multimodal data processing delay and unbalanced resource utilization in the existing technology are solved, and more efficient and stable data processing is achieved.

CN119988332AInactive Publication Date: 2025-05-13BEIJING LIUJINSUIYUE TECH CO LTD

Patent Information

Application Number
CN202510485604.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-17
Publication Date
2025-05-13
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing multimodal data processing methods lack adaptive adjustment capabilities, resulting in the problems of data processing delay and unbalanced resource utilization in high-load environments.

Method used

A multimodal data digital processing process monitoring method is proposed, through file type identification and splitting, cluster analysis and kernel density estimation are used to determine data characterization, dynamically adjust the resource allocation ratio of parallel channels, monitor the processing rate in real time, and optimize and adjust according to historical data.

Benefits of technology

It improves the parallelism and efficiency of data processing, avoids resource waste and processing bottlenecks, ensures load balancing of different channels, and improves overall processing efficiency and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119988332A_ABST
    Figure CN119988332A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a multi-modal data digital processing process monitoring method, which comprises the following steps: acquiring a to-be-processed file, identifying a file type, and splitting the to-be-processed file into a plurality of to-be-processed segments according to the file type; and extracting each to-be-processed segment, performing clustering processing, determining representation data, obtaining a representation total data volume, and determining a distribution proportion of the parallel channels. And acquiring the real-time processing rate of each channel in the parallel channels to judge whether the distribution proportion is adjusted or not. When it is judged that the distribution proportion is adjusted, the data set is determined to be compared with a historical processing data set, an adjustment coefficient is determined according to a comparison result to adjust the distribution proportion, and processing is conducted according to the adjusted distribution proportion. According to the method and the device, by identifying and splitting the type of the to-be-processed file, analysis and parallel processing of different types of multi-modal data are realized, and the overall processing efficiency and stability are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular to a method for monitoring a multi-modal data digital processing process. Background Art

[0002] With the rapid development of digital technology, the demand for processing multimodal data such as audio and video is growing. Multimodal data usually contains multiple information types such as text + voice and text + image. Due to its complex data structure and large amount of information, how to efficiently process and analyze it has become the focus of current technical research.

[0003] Existing multimodal data processing methods usually adopt a fixed channel allocation strategy to input different types of data into independent processing channels, or rely on static resource allocation mechanisms for processing. However, due to the changes in data flow and the uneven distribution of computing resources, traditional methods have certain limitations in high-load environments. There is a lack of effective ways to optimize the data splitting granularity for the splitting and processing of different file types, resulting in uneven data block division, which affects the parallelism of subsequent processing. The lack of adaptive adjustment capabilities makes it impossible to dynamically optimize according to changes in data load, resulting in low resource utilization in some channels, while other channels may be overloaded and cause data processing delays.

[0004] Therefore, it is necessary to design a multimodal data digital processing process monitoring method to solve the problems existing in current technology. Summary of the invention

[0005] In view of this, the present invention proposes a multimodal data digitization processing process monitoring method, aiming to solve the problems of high data processing delay and low data processing efficiency caused by the lack of adaptive adjustment capability in the current multimodal data digitization processing process.

[0006] The present invention proposes a multi-modal data digital processing process monitoring method, comprising:

[0007] Collecting files to be processed and identifying file types, and splitting the files to be processed into a number of segments to be processed according to the file types; the file types include text, video, and text, audio;

[0008] Extracting each of the to-be-processed segments for clustering processing, obtaining clustering results, performing kernel density estimation on the data in each cluster according to the clustering results to determine the characterization data, and obtaining the total characterization data volume of the to-be-processed segments;

[0009] Determining the allocation ratio of the parallel channels according to the total data volume representing all the fragments to be processed;

[0010] Collecting the real-time processing rate of each channel in the parallel channels, and determining whether to adjust the allocation ratio according to the real-time processing rate;

[0011] When it is determined that the allocation ratio needs to be adjusted, the total data volume and rate ratio representing all the fragments to be processed are taken as a data set, and the data set is compared with the historical processing data set. The adjustment coefficient is determined according to the comparison result to adjust the allocation ratio, and processing is performed with the adjusted allocation ratio.

[0012] Furthermore, when extracting each of the to-be-processed segments for clustering processing, the process includes:

[0013] When the to-be-processed segment is text, the text is converted into word vectors based on Word2Vec, the sentence structure is extracted based on dependency syntax, and the word vectors and sentence results are used as clustering features for clustering analysis;

[0014] When the to-be-processed segment is a video, a color histogram and an optical flow feature of the video are obtained, and the color histogram and the optical flow feature are used as clustering features for clustering analysis;

[0015] When the segment to be processed is audio, the Mel-frequency cepstral coefficients and fundamental frequency of the audio are obtained, and the Mel-frequency cepstral coefficients and fundamental frequency are used as clustering features for clustering analysis.

[0016] Furthermore, when extracting each of the to-be-processed segments for clustering processing, the process further includes:

[0017] S1: Initialize K centroids and assign the clustering features to the nearest centroids to form K clusters;

[0018] S2: Recalculate the centroid of each cluster;

[0019] S3: Repeat S1 and S2 until the centroid no longer changes or the specified number of iterations is reached.

[0020] Furthermore, when performing kernel density estimation on the data in each cluster according to the clustering result to determine the characterization data, it includes:

[0021] Gaussian mixture distribution is used to determine the characterization data in each cluster;

[0022] The data corresponding to the highest frequency and the lowest frequency of the kernel density are used as the characterization data.

[0023] Furthermore, when determining the allocation ratio of the parallel channels according to the total amount of data representing all the fragments to be processed, it includes:

[0024] The ratio of the total data amounts representing different fragments to be processed in the to-be-processed file is determined according to the total data amounts representing the fragments to be processed, and the allocation ratio of the parallel channels is determined according to the ratio of the total data amounts representing.

[0025] Further, judging whether to adjust the allocation ratio according to the real-time processing rate includes:

[0026] Obtaining the total transmission channel rate corresponding to different to-be-processed segments of the to-be-processed file according to the real-time processing rate, and obtaining a total rate ratio, and determining whether to adjust the allocation ratio according to the total rate ratio;

[0027] When the total rate ratio satisfies a first condition, determining not to adjust the allocation ratio;

[0028] When the total rate ratio satisfies a second condition, it is determined to adjust the allocation ratio.

[0029] Furthermore, comparing the data set with the historically processed data set, and determining an adjustment coefficient according to the comparison result to adjust the allocation ratio, includes:

[0030] The historical processing data set includes a plurality of historical data sets and a plurality of historical allocation ratios, and each of the historical data sets corresponds to a historical allocation ratio;

[0031] When there is data in the historical processing data set whose similarity with the data set is greater than or equal to the similarity threshold, determining a similar data set and adjusting the allocation ratio by determining an adjustment coefficient according to the similar data set;

[0032] When the similarity between the historical data sets in the historical processing data set and the data set is less than the similarity threshold, the historical data set corresponding to the maximum similarity is determined, and the historical allocation ratio corresponding to the historical data set is obtained, and a first adjustment coefficient is obtained according to the historical allocation ratio and the allocation ratio, and the allocation ratio is adjusted according to the first adjustment coefficient, and the adjusted allocation ratio is the product of the first adjustment coefficient and the allocation ratio.

[0033] Further, determining the same type of data set and determining the adjustment coefficient according to the same type of data set to adjust the allocation ratio includes:

[0034] Classify all data in the historical processing data set whose similarity with the data set is greater than or equal to a similarity threshold into the same type of data set;

[0035] When there is only one historical data set among the data sets of the same type, obtaining a first adjustment coefficient according to the historical allocation ratio and the allocation ratio corresponding to the historical data set, and adjusting the allocation ratio according to the first adjustment coefficient;

[0036] When the historical data set in the same type of data set is not unique, obtaining an average of historical allocation ratios of the historical data set, classifying data in the same type of data set whose historical allocation ratio is greater than or equal to the average of the historical allocation ratio into a first set, classifying data in the same type of data set whose historical allocation ratio is less than the average of the historical allocation ratio into a second set, and determining the adjustment coefficient according to the first set, the second set and the average of the historical allocation ratio;

[0037] Obtain a first variance between all historical allocation ratios in the first set and the mean of the historical allocation ratios, and obtain a second variance between all historical allocation ratios in the second set and the mean of the historical allocation ratios, sum the first variance and the second variance and take the square root to obtain the square root of the variance, and sum half of the square root of the variance with the mean of the historical allocation ratio to obtain the adjustment coefficient, the adjusted allocation ratio being the product of the allocation ratio and the adjustment coefficient.

[0038] Furthermore, after obtaining a first adjustment coefficient according to the historical allocation ratio and the allocation ratio, and adjusting the allocation ratio according to the first adjustment coefficient, the method further includes:

[0039] Collecting the second real-time processing rate of the parallel channel within a preset time period, and obtaining the estimated processing time of different to-be-processed segments of the to-be-processed file based on a short-term prediction model, and determining whether to modify the first adjustment coefficient according to the estimated processing time;

[0040] When the value of the difference between the estimated processing durations of different to-be-processed segments is greater than the duration threshold, determining to modify the first adjustment coefficient;

[0041] When the value of the difference between the estimated processing durations of the different to-be-processed segments is less than or equal to the duration threshold, it is determined that the first adjustment coefficient is not to be corrected, and the first adjustment coefficient is used as the adjustment coefficient to perform processing at the adjusted allocation ratio.

[0042] Further, when the difference between the estimated processing durations of different to-be-processed segments is greater than a duration threshold, it is determined that the first adjustment coefficient is to be corrected, including:

[0043] The difference in estimated processing time of different to-be-processed segments is obtained, and a correction coefficient is determined according to the difference in estimated processing time to correct the adjustment coefficient. When the difference in estimated processing time is less than zero, the correction coefficient takes a value of (0.8, 1); when the difference in estimated processing time is greater than zero, the correction coefficient takes a value of (1, 1.2), and the correction coefficient is proportional to the difference in estimated processing time.

[0044] Compared with the prior art, the beneficial effect of the present invention is that: by identifying and splitting the types of files to be processed, the parsing and parallel processing of different types of multimodal data (including text, video and text audio) are realized. By identifying and splitting the file types, the data distribution is more balanced, the parallelism of processing is improved, and the waste of resources and processing bottlenecks caused by uneven data block division are avoided. The data fragments are classified by cluster analysis method, and the data characterization information is accurately extracted in combination with the kernel density estimation method, so that the data characterization is more representative and the accuracy of subsequent resource allocation is improved. Through the channel allocation mechanism based on the characterization of the total data volume, the dynamic allocation of parallel channel resources is realized to ensure the load balancing of different channels, avoid processing delays caused by data overload in some channels, and prevent inefficient utilization caused by idleness of some channel resources. By real-time monitoring of the processing rate of each channel, based on the comparison and analysis of the data set and the historical processing data set, the allocation ratio is adjusted to ensure that the optimal resource utilization can be maintained under the dynamic load environment, thereby improving the overall processing efficiency and stability. Compared with traditional fixed allocation or static rule adjustment methods, it can adapt to changes in data flow, making data processing more flexible and efficient, and is suitable for intelligent processing scenarios of high concurrency and large-scale multimodal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] Various other advantages and benefits will become apparent to those of ordinary skill in the art by reading the detailed description of the preferred embodiments below. The accompanying drawings are only for the purpose of illustrating the preferred embodiments and are not to be considered as limiting the present invention. Moreover, the same reference symbols are used throughout the accompanying drawings to represent the same components. In the accompanying drawings:

[0046] Figure 1 A flowchart of a method for monitoring a multimodal data digitization process provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0047] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although exemplary embodiments of the present disclosure are shown in the accompanying drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments described herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present disclosure and to fully convey the scope of the present disclosure to those skilled in the art. It should be noted that, in the absence of conflict, the embodiments of the present invention and the features described in the embodiments can be combined with each other. The present invention will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.

[0048] In some embodiments of the present application, see Figure 1 As shown, a multimodal data digital processing process monitoring method includes:

[0049] S100: Collect the files to be processed and identify the file types, and split the files to be processed into a number of segments to be processed according to the file types. The file types include text video and text audio.

[0050] S200: extract each to-be-processed segment for clustering processing, obtain clustering results, perform kernel density estimation on the data in each cluster according to the clustering results to determine the characterization data, and obtain the total characterization data volume of the to-be-processed segment.

[0051] S300: Determine the allocation ratio of parallel channels according to the total representation data volume of all the fragments to be processed.

[0052] S400: collecting the real-time processing rate of each channel in the parallel channel, and determining whether to adjust the allocation ratio according to the real-time processing rate.

[0053] S500: When it is determined that the allocation ratio needs to be adjusted, the total data volume and rate ratio representing all the fragments to be processed are taken as a data set, the data set is compared with the historical processing data set, the adjustment coefficient is determined according to the comparison result to adjust the allocation ratio, and processing is performed with the adjusted allocation ratio.

[0054] Specifically, by collecting the files to be processed and identifying their types (such as text video, text audio), the files are split into several to-be-processed segments according to the data characteristics, such as text video is split into text segments and video segments, ensuring that subsequent processing can be carried out in a more fine-grained manner. The identification and splitting strategy of file types can avoid the problem of uneven data division and improve the processing efficiency and parallelism of data. In the data analysis stage, the clustering method is used to classify each data segment so that similar data can be classified and processed. Clustering analysis is used to reduce data redundancy and improve the rationality of data organization. On this basis, kernel density estimation is introduced to model the probability distribution of each type of data, determine the characterization data, and obtain the total characterization data volume of the entire data segment. Compared with the traditional histogram method, kernel density estimation can more accurately characterize the distribution of data and make data characterization more accurate. Based on the calculated total characterization data volume of all the segments to be processed, the allocation ratio of parallel channels is determined, that is, the resource allocation weights of different channels for processing different types of data are determined. Traditional methods usually adopt a fixed allocation strategy, while this scheme maintains a better channel utilization rate under different data flow conditions through adaptive dynamic allocation. During the real-time processing process, the real-time processing rate of each parallel channel is collected, and it is determined whether the allocation ratio needs to be adjusted. If it is detected that the processing rate is not in line with expectations, it indicates that the utilization rate of the current channel resources is unbalanced, and the adaptive optimization stage is entered. In this stage, the total data volume representing all data fragments and the real-time rate ratio of each channel constitute a data set, and it is compared with the historical processing data set. Based on the comparison and analysis results, the adjustment coefficient is calculated, and the channel allocation ratio is optimized and adjusted accordingly to ensure that the data stream can be processed in the best way. The adjusted allocation ratio is used to continue data processing, making the system resource allocation more intelligent and dynamic.

[0055] It can be understood that through file type identification and splitting, the parallelism of data processing is improved, avoiding resource waste and processing bottlenecks caused by uneven data block division; combining cluster analysis and kernel density estimation, accurate modeling of data representation is achieved, providing high-quality data support for channel resource allocation; through an adaptive channel allocation mechanism, it is possible to monitor and dynamically adjust the resource allocation ratio in real time to ensure load balancing of different channels, avoiding processing delays due to data overload or resource waste due to channel idleness. By comparing the current data set with the historical data set, optimization and adjustment are performed to further improve stability and data processing efficiency. Compared with traditional fixed or static rule allocation methods, this embodiment is more flexible in processing high-concurrency, large-scale multimodal data, and can effectively adapt to dynamic changes in data streams, achieving more efficient and stable multimodal data digital processing process monitoring.

[0056] In some embodiments of the present application, when extracting each to-be-processed segment for clustering processing, the process includes:

[0057] When the fragment to be processed is text, the text is converted into word vectors based on Word2Vec, the sentence structure is extracted based on dependency syntax, and the word vectors and sentence results are used as clustering features for clustering analysis.

[0058] When the segment to be processed is a video, a color histogram and an optical flow feature of the video are obtained, and the color histogram and the optical flow feature are used as clustering features for clustering analysis.

[0059] When the segment to be processed is audio, the Mel-frequency cepstral coefficients and fundamental frequency of the audio are obtained, and the Mel-frequency cepstral coefficients and fundamental frequency are used as clustering features for clustering analysis.

[0060] In some embodiments of the present application, when extracting each to-be-processed segment for clustering processing, the process further includes:

[0061] S1: Initialize K centroids and assign clustering features to the nearest centroids to form K clusters.

[0062] S2: Recalculate the centroid of each cluster.

[0063] S3: Repeat S1 and S2 until the centroid no longer changes or the specified number of iterations is reached.

[0064] The centroid calculation formula is as follows:

[0065]

[0066] Among them, Mk represents the centroid of the kth cluster, represents the number of clustering features in the kth cluster, and Ti represents the i-th clustering feature.

[0067] Specifically, when the clip to be processed is text, Word2Vec is used to convert the text into word vectors, thereby converting discrete text data into high-dimensional vector representations, so that semantically similar words are closer in the vector space. In order to retain the text structure information, dependency syntax analysis is used to extract the grammatical structure of the sentence. Dependency syntax can parse the grammatical relations of the text such as subject, predicate, object, attributive, adverbial, and complement. When clustering, it not only focuses on the similarity at the lexical level, but also identifies the relevance of syntactic structures. Word vectors and grammatical structure features are used as clustering features for clustering analysis to improve the semantic clustering accuracy of text data. When the clip to be processed is a video, color histogram and optical flow features are used as clustering features. Color histogram is a statistical analysis of the color distribution of video frames, which can effectively describe the overall color characteristics of video content and is suitable for distinguishing video data of different scenes. Optical flow features can describe the movement of objects in the video. By calculating the motion vectors between adjacent frames, dynamic information of the video content can be extracted. Combining color histogram and optical flow features, static visual features and dynamic motion features of the video can be considered simultaneously in the clustering process, so that similar video clips can be reasonably classified. When the clip to be processed is audio, the solution uses Mel-frequency cepstral coefficients and fundamental frequency as clustering features. Mel-frequency cepstral coefficients can extract the spectral features of audio and are widely used in speech recognition and audio classification. By performing a short-time Fourier transform on the audio signal and then processing the spectrum through a Mel filter, the low-frequency features are made more recognizable and suitable for the classification of speech signals. The fundamental frequency represents the basic vibration frequency of the audio and can be used to distinguish different timbres and speech patterns. Combining the Mel-frequency cepstral coefficients and fundamental frequency information can more comprehensively characterize the audio data, making the clustering effect more accurate.

[0068] It is understandable that different feature extraction methods are used for different types of data to ensure the accuracy of clustering analysis. Word2Vec and dependency syntax analysis are used for text data, so that both the semantic and structural information of the text can be effectively utilized; video data combines color histogram and optical flow features to achieve a comprehensive analysis of static and dynamic information; audio data uses MFCC and fundamental frequency to better identify the spectrum and pitch characteristics of audio signals. The K-means clustering algorithm is used for data classification, so that data can be reasonably classified according to feature similarity, improving the accuracy and efficiency of subsequent data processing. Compared with the traditional single feature clustering method, it can perform adaptive optimization for different data types, improving the accuracy and stability of multimodal data processing.

[0069] In some embodiments of the present application, when performing kernel density estimation on the data in each cluster according to the clustering result to determine the characterization data, it includes: using Gaussian mixture distribution to determine the characterization data in each cluster:

[0070] The kernel density function expression is:

[0071]

[0072] Where n represents the number of representation data in each cluster, h represents the smoothing bandwidth, represents the i-th characterization data in each cluster, and x represents the value of a certain frequency point to be estimated.

[0073] The data corresponding to the highest frequency and the lowest frequency of the kernel density are used as the characterization data.

[0074] In some embodiments of the present application, when determining the allocation ratio of parallel channels based on the total data volume of representations of all fragments to be processed, it includes: determining the ratio of the total data volumes of representations of different fragments to be processed in the file to be processed based on the total data volume of representations of the fragments to be processed, and determining the allocation ratio of parallel channels based on the ratio of the total data volumes of representations.

[0075] Specifically, the highest frequency point of kernel density indicates the most common eigenvalues ​​in the data distribution, representing the typical data pattern of the cluster. The lowest frequency point of kernel density is located at the edge of the data distribution and can be used to describe the boundary characteristics of the data, making the representation of the data more comprehensive.

[0076] It can be understood that by estimating the kernel density of the Gaussian mixture distribution, the characterization data of the data cluster is extracted, making the data feature extraction more representative and improving the accuracy of the processing. Based on the calculation and proportional allocation strategy of the total data volume, the allocation of computing resources can be dynamically adjusted to ensure the load balance of parallel channels and avoid the problem of idle or overloaded computing resources. The adaptive allocation mechanism improves computing efficiency and stability.

[0077] In some embodiments of the present application, when determining whether to adjust the allocation ratio based on the real-time processing rate, it includes: obtaining the total rate of the transmission channel corresponding to different to-be-processed fragments of the to-be-processed file based on the real-time processing rate, and obtaining the total rate ratio, and determining whether to adjust the allocation ratio based on the total rate ratio.

[0078] Specifically, when the total rate ratio satisfies the first condition, it is determined that the allocation ratio is not adjusted. When the total rate ratio satisfies the second condition, it is determined that the allocation ratio is adjusted.

[0079] Specifically, the first condition is 0.9≤total rate ratio≤1.1, the second condition is total rate ratio>1.1, or the second condition is total rate ratio<0.9.

[0080] In some embodiments of the present application, a data set is compared with a historically processed data set, and an adjustment coefficient is determined based on the comparison result to adjust the allocation ratio, including: the historically processed data set includes several historical data sets and several historical allocation ratios, and each historical data set corresponds to a historical allocation ratio.

[0081] Specifically, when there is data in the historical processing data set whose similarity with the data set is greater than or equal to the similarity threshold, a similar data set is determined and an adjustment coefficient is determined based on the similar data set to adjust the allocation ratio. When the similarity between the historical data set in the historical processing data set and the data set is less than the similarity threshold, the historical data set corresponding to the maximum similarity is determined, and the historical allocation ratio corresponding to the historical data set is obtained. The first adjustment coefficient is obtained based on the historical allocation ratio and the allocation ratio, and the allocation ratio is adjusted based on the first adjustment coefficient. The adjusted allocation ratio is the product of the first adjustment coefficient and the allocation ratio.

[0082] In some embodiments of the present application, determining a similar data set and adjusting the allocation ratio based on the adjustment coefficient of the similar data set includes: classifying all data in the historical processing data set whose similarity with the data set is greater than or equal to the similarity threshold into the similar data set.

[0083] Specifically, when there is only one historical data set in the same type of data set, the first adjustment coefficient is obtained according to the historical allocation ratio and the allocation ratio corresponding to the historical data set, and the allocation ratio is adjusted according to the first adjustment coefficient. When the historical data set in the same type of data set is not unique, the mean of the historical allocation ratio of the historical data set is obtained, and the data with a historical allocation ratio greater than or equal to the mean of the historical allocation ratio in the same type of data set is classified into the first set, and the data with a historical allocation ratio less than the mean of the historical allocation ratio in the same type of data set is classified into the second set, and the adjustment coefficient is determined according to the first set, the second set and the mean of the historical allocation ratio.

[0084] Specifically, obtain the first variance of all historical allocation ratios in the first set and the average of the historical allocation ratios, and obtain the second variance of all historical allocation ratios in the second set and the average of the historical allocation ratios. Sum the first variance and the second variance and take the square root to obtain the square root of the variance, and sum half of the square root of the variance with the average of the historical allocation ratio to obtain the adjustment coefficient. The adjusted allocation ratio is the product of the allocation ratio and the adjustment coefficient.

[0085] It is understandable that the method of obtaining historical processing data sets depends on the continuous recording and summarization of past processing tasks. During the data processing process, the feature information of each type of fragment to be processed, the allocation ratio of parallel channels, the processing rate and the final processing effect are monitored in real time, and these data are stored as historical processing data sets. Specifically, first, the input data of each task to be processed is feature extracted, including the semantic vector of the text, the spectral characteristics of the audio, the image histogram of the video, etc.; secondly, the allocation ratio and processing efficiency of different channels during the execution of the task are recorded; finally, combined with indicators such as task completion time and resource utilization, a complete historical data record is formed, and data cleaning and optimization are performed regularly to retain representative historical data sets to ensure that they can provide effective optimization references in subsequent similar task processing.

[0086] It is understandable that dynamically adjusting computing resources according to real-time rates improves channel utilization and prevents resource overload or idleness. Using historical processing experience makes allocation adjustments more accurate and reduces uncertainty in the adjustment process. By setting rate ratio thresholds and smoothing adjustment coefficients, the adjustment strategy is ensured to be stable, avoiding frequent allocation of computing resources due to short-term fluctuations and affecting overall performance. Reasonable resource allocation optimization balances the load of parallel channels and improves overall data processing capabilities.

[0087] In some embodiments of the present application, a first adjustment coefficient is obtained based on the historical allocation ratio and the allocation ratio, and after the allocation ratio is adjusted according to the first adjustment coefficient, it also includes: collecting the second real-time processing rate of the parallel channel within a preset time period, and obtaining the estimated processing time of different to-be-processed fragments of the to-be-processed file based on the short-term prediction model, and determining whether to correct the first adjustment coefficient based on the estimated processing time.

[0088] Specifically, when the value of the difference between the estimated processing durations of different to-be-processed segments is greater than the duration threshold, it is determined that the first adjustment coefficient is to be corrected. When the value of the difference between the estimated processing durations of different to-be-processed segments is less than or equal to the duration threshold, it is determined that the first adjustment coefficient is not to be corrected, and the first adjustment coefficient is used as the adjustment coefficient, and the processing is performed at the adjusted allocation ratio.

[0089] In some embodiments of the present application, when the value of the difference between the estimated processing times of different fragments to be processed is greater than a time threshold, it is determined that the first adjustment coefficient is to be corrected, including: obtaining the difference between the estimated processing times of different fragments to be processed, determining a correction coefficient based on the difference in the estimated processing times to correct the adjustment coefficient, when the difference in the estimated processing time is less than zero, the correction coefficient takes a value of (0.8, 1), when the difference in the estimated processing time is greater than zero, the correction coefficient takes a value of (1, 1.2), and the correction coefficient is proportional to the difference in the estimated processing time.

[0090] It is understandable that after completing the first step of adjustment based on historical data, the second real-time processing rate of the parallel channel will be continuously monitored within a preset period of time, and the estimated processing time of different to-be-processed segments will be calculated using a short-term prediction model. The prediction model can be trained based on historical data, such as using time series analysis or machine learning methods to model factors such as the current data processing rate and computing resource utilization to obtain a processing time prediction value.

[0091] After predicting the estimated processing time of different segments, calculate their time difference and compare it with the preset time threshold. The time threshold is a natural number greater than zero. The value of the difference in the estimated processing time is the absolute value of the difference in the estimated processing time. If the value of the difference in the estimated processing time exceeds the time threshold, it indicates that there is still room for optimization in the current allocation ratio and the adjustment coefficient needs to be further corrected; if the value of the difference in the time is lower than or equal to the threshold, keep the original adjustment coefficient unchanged to ensure that the adjusted allocation ratio is used for subsequent processing. When the adjustment coefficient needs to be corrected, the range of the correction coefficient is determined according to the positive or negative sign of the difference in the estimated processing time: if the difference in time is less than zero, it indicates that the processing time of the text segment is short, while the processing time of the video or audio is long. The adjustment coefficient needs to be lowered to reduce the resource allocation of the segment and make the overall load more balanced. Therefore, the correction coefficient ranges between (0.8, 1); if the difference in time is greater than zero, it indicates that the processing time of the text segment is long, while the processing time of the video or audio is short. The adjustment coefficient needs to be increased to improve the resource allocation of the text segment. Therefore, the correction coefficient ranges between (1, 1.2) and is proportional to the difference in the estimated time to achieve more accurate optimization. Because the time threshold is a natural number greater than zero, there is no case where the time difference is zero when determining to correct the adjustment coefficient.

[0092] It is understandable that, based on the first step of adjustment, the allocation ratio is further optimized through real-time monitoring and short-term prediction, making data processing more intelligent and dynamic. By predicting the expected processing time of different data fragments, the problem of uneven allocation of channel resources can be effectively reduced, overloading or inefficient operation of certain channels can be avoided, and the stability and throughput of overall data processing can be improved.

[0093] In the above embodiment, by identifying and splitting the types of files to be processed, different types of multimodal data (including text, video and text audio) are parsed and processed in parallel. By identifying and splitting the file types, the data distribution is more balanced, the parallelism of processing is improved, and resource waste and processing bottlenecks caused by uneven data block division are avoided. The data fragments are classified by cluster analysis method, and the data characterization information is accurately extracted in combination with the kernel density estimation method, so that the data characterization is more representative and the accuracy of subsequent resource allocation is improved. Through the channel allocation mechanism based on the characterization of the total data volume, the dynamic allocation of parallel channel resources is realized to ensure the load balancing of different channels, avoid processing delays due to data overload in some channels, and prevent inefficient utilization of some channel resources due to idleness. By real-time monitoring of the processing rate of each channel, based on the comparison and analysis of the data set and the historical processing data set, the allocation ratio is adjusted to ensure that the optimal resource utilization can be maintained under the dynamic load environment, thereby improving the overall processing efficiency and stability. Compared with the traditional fixed allocation or static rule adjustment method, it can adapt to the changes in data flow, making data processing more flexible and efficient, and is suitable for intelligent processing scenarios of high concurrency and large-scale multimodal data.

[0094] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.

[0095] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0096] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the above embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present invention can still be modified or replaced by equivalents. Any modification or equivalent replacement that does not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A method for monitoring a multimodal data digital processing process, characterized in that: include: Collecting files to be processed and identifying file types, and splitting the files to be processed into a number of segments to be processed according to the file types; the file types include text, video, and text, audio; Extracting each of the to-be-processed segments for clustering processing, obtaining clustering results, performing kernel density estimation on the data in each cluster according to the clustering results to determine the characterization data, and obtaining the total characterization data volume of the to-be-processed segments; Determining the allocation ratio of the parallel channels according to the total data volume representing all the fragments to be processed; Collecting the real-time processing rate of each channel in the parallel channels, and determining whether to adjust the allocation ratio according to the real-time processing rate; When it is determined that the allocation ratio needs to be adjusted, the total data volume and rate ratio representing all the fragments to be processed are taken as a data set, and the data set is compared with the historical processing data set. The adjustment coefficient is determined according to the comparison result to adjust the allocation ratio, and processing is performed with the adjusted allocation ratio.

2. The multimodal data digital processing process monitoring method according to claim 1, characterized in that: The step of extracting each of the to-be-processed segments for clustering includes: When the to-be-processed segment is text, the text is converted into word vectors based on Word2Vec, the sentence structure is extracted based on dependency syntax, and the word vectors and sentence results are used as clustering features for clustering analysis; When the to-be-processed segment is a video, a color histogram and an optical flow feature of the video are obtained, and the color histogram and the optical flow feature are used as clustering features for clustering analysis; When the segment to be processed is audio, the Mel-frequency cepstral coefficients and fundamental frequency of the audio are obtained, and the Mel-frequency cepstral coefficients and fundamental frequency are used as clustering features for clustering analysis.

3. The multimodal data digital processing process monitoring method according to claim 2, characterized in that: When extracting each of the to-be-processed segments for clustering processing, the method further includes: S1: Initialize K centroids and assign the clustering features to the nearest centroids to form K clusters; S2: Recalculate the centroid of each cluster; S3: Repeat S1 and S2 until the centroid no longer changes or the specified number of iterations is reached.

4. The multimodal data digital processing process monitoring method according to claim 3 is characterized in that: When performing kernel density estimation on the data in each cluster according to the clustering result to determine the characterization data, it includes: Gaussian mixture distribution is used to determine the characterization data in each cluster; The data corresponding to the highest frequency and the lowest frequency of the kernel density are used as the characterization data.

5. The method for monitoring the multimodal data digital processing process according to claim 4, characterized in that: When determining the allocation ratio of the parallel channels according to the total amount of representation data of all the to-be-processed fragments, it includes: The ratio of the total data amounts representing different fragments to be processed in the to-be-processed file is determined according to the total data amounts representing the fragments to be processed, and the allocation ratio of the parallel channels is determined according to the ratio of the total data amounts representing.

6. The method for monitoring the multimodal data digital processing process according to claim 5, characterized in that: When judging whether to adjust the allocation ratio according to the real-time processing rate, it includes: Obtaining the total transmission channel rate corresponding to different to-be-processed segments of the to-be-processed file according to the real-time processing rate, and obtaining a total rate ratio, and determining whether to adjust the allocation ratio according to the total rate ratio; When the total rate ratio satisfies a first condition, determining not to adjust the allocation ratio; When the total rate ratio satisfies a second condition, it is determined to adjust the allocation ratio.

7. The method for monitoring a multimodal data digital processing process according to claim 6, characterized in that: The data set is compared with the historical processing data set, and the adjustment coefficient is determined according to the comparison result to adjust the allocation ratio, including: The historical processing data set includes a plurality of historical data sets and a plurality of historical allocation ratios, and each of the historical data sets corresponds to a historical allocation ratio; When there is data in the historical processing data set whose similarity with the data set is greater than or equal to the similarity threshold, determining a similar data set and adjusting the allocation ratio by determining an adjustment coefficient according to the similar data set; When the similarity between the historical data sets in the historical processing data set and the data set is less than the similarity threshold, the historical data set corresponding to the maximum similarity is determined, and the historical allocation ratio corresponding to the historical data set is obtained, and a first adjustment coefficient is obtained according to the historical allocation ratio and the allocation ratio, and the allocation ratio is adjusted according to the first adjustment coefficient, and the adjusted allocation ratio is the product of the first adjustment coefficient and the allocation ratio.

8. The method for monitoring the multimodal data digital processing process according to claim 7, characterized in that: Determining a similar data set and determining an adjustment coefficient based on the similar data set to adjust the allocation ratio includes: Classify all data in the historical processing data set whose similarity with the data set is greater than or equal to a similarity threshold into the same type of data set; When there is only one historical data set among the data sets of the same type, obtaining a first adjustment coefficient according to the historical allocation ratio and the allocation ratio corresponding to the historical data set, and adjusting the allocation ratio according to the first adjustment coefficient; When the historical data set in the same type of data set is not unique, obtaining the average historical allocation ratio of the historical data set, classifying the data whose historical allocation ratio in the same type of data set is greater than or equal to the average historical allocation ratio into a first set, classifying the data whose historical allocation ratio in the same type of data set is less than the average historical allocation ratio into a second set, and determining the adjustment coefficient according to the first set, the second set and the average historical allocation ratio; Obtain a first variance between all historical allocation ratios in the first set and the mean of the historical allocation ratios, and obtain a second variance between all historical allocation ratios in the second set and the mean of the historical allocation ratios, sum the first variance and the second variance and take the square root to obtain the square root of the variance, and sum half of the square root of the variance with the mean of the historical allocation ratio to obtain the adjustment coefficient, the adjusted allocation ratio being the product of the allocation ratio and the adjustment coefficient.

9. The method for monitoring the multimodal data digital processing process according to claim 7, characterized in that: The method further includes obtaining a first adjustment coefficient according to the historical allocation ratio and the allocation ratio, and adjusting the allocation ratio according to the first adjustment coefficient: Collecting the second real-time processing rate of the parallel channel within a preset time period, and obtaining the estimated processing time of different to-be-processed segments of the to-be-processed file based on a short-term prediction model, and determining whether to modify the first adjustment coefficient according to the estimated processing time; When the value of the difference between the estimated processing durations of different to-be-processed segments is greater than the duration threshold, determining to modify the first adjustment coefficient; When the value of the difference between the estimated processing durations of the different to-be-processed segments is less than or equal to the duration threshold, it is determined that the first adjustment coefficient is not to be corrected, and the first adjustment coefficient is used as the adjustment coefficient to perform processing at the adjusted allocation ratio.

10. The method for monitoring the multimodal data digital processing process according to claim 9, characterized in that: When the value of the difference between the estimated processing durations of different to-be-processed segments is greater than the duration threshold, it is determined that the first adjustment coefficient is to be corrected, including: The difference in estimated processing time of different to-be-processed segments is obtained, and a correction coefficient is determined according to the difference in estimated processing time to correct the adjustment coefficient. When the difference in estimated processing time is less than zero, the correction coefficient takes a value of (0.8, 1); when the difference in estimated processing time is greater than zero, the correction coefficient takes a value of (1, 1.2), and the correction coefficient is proportional to the difference in estimated processing time.

Citation Information

Patent Citations

  • Unstructured electric power big data analysis method based on deep learning

    CN114723583A

  • Resource management system for state machine

    CN117336300A

  • Group behavior analysis method for multi-modal information fusion and dynamic updating

    CN119337197A

  • USB interface management and control method and device for information security protection

    CN119494127A

  • Video analysis task scheduling method and system based on intelligent scheduling mode

    CN119668819A

Cited By

  • Method and system for optimizing water consumption of irrigated area based on dynamic data management

    CN120611839A