Data processing method and system based on multi-modal AI model

By monitoring the feature extraction process of the multimodal AI model, the modal data that completed the feature extraction first was captured as a reference and performs de-redundancy operations, solving the problems of data redundancy and processing complexity in the multimodal AI model, and achieving efficient and accurate data processing and model performance improvement.

CN120408029AInactive Publication Date: 2025-08-01GUANGZHOU TUCSON ELECTRONIC TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510459250.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-11
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In data processing, multimodal AI models have difficulty in data alignment, complex feature extraction and fusion, high data acquisition and labeling costs, and insufficient redundancy processing, resulting in limited improvement in model performance.

Method used

By monitoring the feature extraction process of the multimodal AI model, the modal data that completes feature extraction first is captured as a reference, the feature extraction is stopped and the extraction of other modal data is analyzed, the reference features are used for de-redundant operations, the retained feature extraction strategy is performed and the features are packaged.

Benefits of technology

It realizes efficient processing and redundancy elimination of multimodal data, improves data processing efficiency and accuracy, provides a better data foundation, and improves the performance of multimodal AI models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120408029A_ABST
    Figure CN120408029A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and discloses a data processing method and system based on a multi-modal AI model, and the method comprises the steps: determining a starting moment, determining a reference condition, determining a completion condition, removing redundancy, and executing feature extraction after redundancy removal. The system corresponds to the method. According to the method, the starting moment and the multi-modal data are determined when the multi-modal AI model is started, the modal data of which feature extraction is firstly completed are captured as reference by monitoring the feature extraction process, extraction is stopped at the reference moment, and the extraction conditions of other modal data are analyzed; thirdly, performing redundancy elimination operation on the extracted features and the non-executed feature extraction strategy by taking the reference features as a benchmark, and finally executing the reserved feature extraction strategy and packaging the features; therefore, data redundancy is effectively reduced, unnecessary feature extraction calculation is avoided, a better data basis is provided for subsequent analysis and application based on multi-modal data, and the performance of the whole multi-modal data processing system is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of data processing, and specifically, to a data processing method and system based on a multi-modal AI model. Background Art

[0002] With the development of artificial intelligence technology, multi-modal AI models are increasingly widely used in various fields. Multi-modal data includes various forms such as images, texts, and voices, which can provide rich information but also bring many challenges. In terms of data processing, the heterogeneity, complementarity, and redundancy of different modal data increase the processing difficulty. Data alignment is difficult, feature extraction and fusion are complex, and the costs of data acquisition and annotation are high. These problems restrict the performance improvement of multi-modal AI models.

[0003] Chinese Patent No. CN111768367B discloses a data processing method, device, and storage medium. This invention document mainly focuses on two modalities, namely ultrasonic images and clinical texts. The utilization of data modalities is insufficient, and other potentially important modalities are not involved, resulting in its generalization ability being limited to only grading diagnoses for specific regions and being difficult to adapt to different data differences. At the same time, its consideration of real-time performance is lacking, and the complex processing flow is difficult to meet the existing data processing requirements. Moreover, its handling of data redundancy is insufficient, and it is difficult to process massive and wet data.

[0004] In view of the above problems in the prior art, there is an urgent need to propose a new technical solution for data processing based on a multi-modal AI model to solve the problems in multi-modal data processing and improve the model performance and application effects. Summary of the Invention

[0005] The purpose of this application is to provide a data processing method and system based on a multi-modal AI model to solve the technical problems raised in the above background art.

[0006] To achieve the above purpose, this application discloses the following technical solutions:

[0007] In a first aspect, this application discloses a data processing method based on a multi-modal AI model, and the method includes the following steps:

[0008] S1: Define the moment when the multi-modal AI model starts running as the starting moment, and at this starting moment, there is multi-modal data for the same object; wherein, the multi-modal data includes multiple different modal data;

[0009] S2: Monitor the process of processing the multi-modal data using the multi-modal AI model, capture the modal features of the modal data that first completes feature extraction in this process, define these modal features as reference features, and define the moment when this modal data first completes feature extraction as the reference moment;

[0010] S3: In the multimodal data, at the reference moment, stop the feature extraction of the multimodal data, and respectively analyze the completion status of the feature extraction of all modal data except the modal data corresponding to the reference feature, to obtain the extracted features of the modal data and the corresponding unexecuted feature extraction strategies;

[0011] S4: Use the reference feature to perform corresponding redundancy removal operations on the extracted features and the corresponding unexecuted feature extraction strategies, define the extracted features after performing the redundancy removal operation as retained features, and define the unexecuted feature extraction strategies after performing the redundancy removal operation as retained feature extraction strategies;

[0012] S5: Only execute the retained feature extraction strategy, and after packing the obtained features with the corresponding retained features, obtain the modal features corresponding to all modal data in the multimodal data except the modal data corresponding to the reference feature.

[0013] Preferably, in S1, at the starting moment, there is multimodal data for the same object, including:

[0014] After the multimodal AI model starts running, collect the modal data corresponding to multiple different data sources of the same object to obtain the multimodal data, and the multimodal data is related to the object in terms of time and space, and this relationship is used to represent the data set of the same object at the same moment.

[0015] Preferably, in S2, monitor the process of processing the multimodal data using the multimodal AI model, including,

[0016] Collect the log records of the operation of the multimodal AI model, and based on the log records, determine the key node information of each modal data in the feature extraction process, and the key node information at least includes the completion time of each feature extraction strategy in the feature extraction process;

[0017] For a certain modal data, when the modal data has executed all its corresponding feature extraction strategies, define the execution completion moment of its last feature extraction strategy as the feature extraction complete completion moment and output the feature extraction complete completion moment; wherein, the feature extraction complete completion moment is used to capture the modal data that first completes the feature extraction.

[0018] Preferably, in S2, defining the modal feature as the reference feature includes:

[0019] Monitor the output at the moment when the feature extraction is completely finished, capture the modal data corresponding to the first moment when the feature extraction is completely finished as the modal data that first completes the feature extraction, and obtain the modal features corresponding to this modal data. Perform a preset integrity and accuracy verification on this modal feature, and after passing this verification, define this modal feature as a reference feature and output it.

[0020] Preferably, in S2, defining the moment when the modal data first completes the feature extraction as the reference moment includes:

[0021] Based on the moment recorded by the system clock, perform a preset calibration on the moment when the feature extraction of the modal data that first completes the feature extraction is completely finished, and after completing this calibration, define this moment when the feature extraction is completely finished as the reference moment and output it.

[0022] Preferably, in S3, the extracted features are a feature set composed of modal features that have been obtained by each modal data through a preset feature extraction strategy before the reference moment; wherein, the modal features at least include data features obtained based on statistical analysis and semantic features obtained based on machine learning.

[0023] Preferably, in S3, the unexecuted feature extraction strategy is the feature extraction strategy that each modal data originally planned but has not executed after the reference moment, and this feature extraction strategy at least includes using different feature extraction algorithms and adjusting the parameters of the feature extraction.

[0024] Preferably, in S4, defining the extracted features after performing the redundancy removal operation as the retained features and defining the unexecuted feature extraction strategy after performing the redundancy removal operation as the retained feature extraction strategy includes:

[0025] Calculate the similarity between the extracted features and the reference feature, determine the extracted features with a similarity greater than or equal to the preset similarity threshold as redundant features and remove them, and define the remaining extracted features as the retained features;

[0026] Analyze the correlation between the features expected to be extracted by the unexecuted feature extraction strategy and the reference feature and the retained features; when the correlation between the features expected to be extracted and the reference feature or the retained features is greater than or equal to the preset correlation threshold, then determine that this unexecuted feature extraction strategy generates redundant features and remove them, and define the remaining unexecuted feature extraction strategy as the retained feature extraction strategy.

[0027] Preferably, in S4, it further includes:

[0028] For the redundant features or unexecuted feature extraction strategies existing in different modality data after the above redundancy removal operation, analyze the reflectivity of different modality data to the modality features, and use the modality features obtained from the modality data corresponding to the maximum value of the reflectivity as the finally retained features for describing the object; wherein, the calculation of the reflectivity is as follows:

[0029] For each modality data, respectively construct a mapping relationship model between the modality data and the redundant features; use the sample data of the modality data to train the constructed mapping relationship model to obtain a trained mapping relationship model; use the trained mapping relationship model to predict the redundant features in the modality data to obtain predicted values; calculate the error between the actual value and the predicted value of the redundant features in the modality data, and define the sum of the reciprocal of the error and a preset modality reflectivity correction value as the reflectivity of the modality data to the redundant features, and the greater the reflectivity, the stronger the reflection ability of the modality data to the redundant features, wherein the modality reflectivity correction value is used to characterize the importance of the modality data for describing the object.

[0030] In a second aspect, the present application discloses a data processing system based on a multi-modal AI model, which system is applicable to the data processing method based on a multi-modal AI model as described above, and the system includes a start time determination module, a reference situation determination module, a completion situation determination module, a redundancy removal module, and a post-redundancy execution module that are communicatively connected in sequence;

[0031] The start time determination module is configured to: define the time when the multi-modal AI model starts running as the start time, and there is multi-modal data for the same object at this start time; wherein, the multi-modal data includes multiple different modality data;

[0032] The reference situation determination module is configured to: monitor the process of processing the multi-modal data by using the multi-modal AI model, capture the modality features of the modality data that first completes feature extraction in this process, define the modality features as reference features, and define the time when the modality data first completes feature extraction as the reference time;

[0033] The completion situation determination module is configured to: in the multi-modal data, at the reference time, stop the feature extraction of the multi-modal data, and respectively analyze the completion situations of the feature extraction of all modality data except the modality data corresponding to the reference features to obtain the extracted features of the modality data and the corresponding unexecuted feature extraction strategies;

[0034] The redundancy removal module is configured to perform corresponding redundancy removal operations on the extracted features and the corresponding unexecuted feature extraction strategies by using the reference features, define the extracted features after performing the redundancy removal operations as retained features, and define the unexecuted feature extraction strategies after performing the redundancy removal operations as retained feature extraction strategies;

[0035] The post-redundancy execution module is configured to only execute the retained feature extraction strategy, and after packing the obtained features with the corresponding retained features, obtain the modal features corresponding to all modal data in the multi-modal data except the modal data corresponding to the reference features.

[0036] Beneficial effects: The data processing method and system based on a multi-modal AI model of the present application utilize the time marking and feature extraction monitoring mechanism during the processing of data by the multi-modal AI model to achieve efficient processing and redundancy elimination of multi-modal data; when the multi-modal AI model is started, the starting moment and multi-modal data are determined, and by monitoring the feature extraction process, the modal data that first completes feature extraction is captured as a reference, and the extraction of other modal data is stopped and analyzed at the reference moment; then, redundancy removal operations are performed on the extracted features and the unexecuted feature extraction strategies based on the reference features, and finally the retained feature extraction strategy is executed and the features are packed; thereby effectively reducing data redundancy, avoiding unnecessary feature extraction calculations, improving the data processing efficiency of the multi-modal AI model, enabling the finally obtained modal features to more accurately describe the object, providing a better data basis for subsequent analysis and applications based on multi-modal data, and enhancing the performance of the entire multi-modal data processing system. Description of the Drawings

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following described drawings are only some embodiments of the present application, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0038] Figure 1 It is a flowchart of the data processing method based on a multi-modal AI model provided by an embodiment of the present application;

[0039] Figure 2 It is a structural block diagram of the data processing system based on a multi-modal AI model provided by an embodiment of the present application. Detailed Embodiments

[0040] The technical solutions in the embodiments of the present application will be described clearly and completely below. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0041] In this article, the term "including" is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or also includes elements inherent in such a process, method, article or device. Without further limitation, the elements defined by the statement "including..." do not exclude the existence of additional identical elements in the process, method, article or device including the said elements.

[0042] The first aspect of this embodiment discloses a data processing method based on a multi-modal AI model as Figure 1 shown, and the method includes the following steps:

[0043] S1: Define the moment when the multi-modal AI model starts running as the starting moment, and there is multi-modal data for the same object at this starting moment; among them, the multi-modal data includes multiple different modal data;

[0044] S2: Monitor the process of processing multi-modal data by the multi-modal AI model, capture the modal features of the modal data that first completes feature extraction in this process, define the modal features as reference features, and define the moment when the modal data first completes feature extraction as the reference moment;

[0045] S3: In the multi-modal data, at the reference moment, stop the feature extraction of the multi-modal data, and respectively analyze the completion of the feature extraction of all modal data except the modal data corresponding to the reference features, to obtain the extracted features of the modal data and the corresponding unexecuted feature extraction strategies;

[0046] S4: Perform corresponding redundancy removal operations on the extracted features and the corresponding unexecuted feature extraction strategies using the reference features, define the extracted features after performing the redundancy removal operation as retained features, and define the unexecuted feature extraction strategy after performing the redundancy removal operation as the retained feature extraction strategy;

[0047] S5: Only execute the retained feature extraction strategy, and after packing the obtained features with the corresponding retained features, obtain the modal features corresponding to all modal data in the multi-modal data except the modal data corresponding to the reference features.

[0048] With the above, when the multi-modal AI model processes data, this embodiment uses the time marking and feature extraction monitoring mechanism to achieve the efficient processing and redundancy elimination of multi-modal data; when the multi-modal AI model is started, the starting time and multi-modal data are determined. By monitoring the feature extraction process, the modal data that completes feature extraction first is captured as a reference, and the extraction of other modal data is stopped and analyzed at the reference time; then, based on the reference features, redundancy elimination operations are performed on the extracted features and the unexecuted feature extraction strategies. Finally, the retained feature extraction strategy is executed and the features are packaged; thus effectively reducing data redundancy, avoiding unnecessary feature extraction calculations, improving the data processing efficiency of the multi-modal AI model, enabling the finally obtained modal features to more accurately describe the object, providing a better data basis for subsequent analysis and applications based on multi-modal data, and enhancing the performance of the entire multi-modal data processing system.

[0049] Specifically, in S1, at the starting time, there is multi-modal data for the same object, including:

[0050] After the multi-modal AI model starts running, modal data corresponding to multiple different data sources of the same object is collected to obtain multi-modal data, and the multi-modal data is related to the object in terms of time and space, and this relationship is used to represent the data set of the same object at the same moment.

[0051] With the above, this embodiment uses the method of collecting modal data related to the object in terms of time and space from multiple data sources to achieve the accurate acquisition of multi-modal data. After the multi-modal AI model is started, data is collected from different data sources to ensure that the data is closely related to the object in terms of time and space, forming a multi-modal data set for the same object at the same moment. This collection method avoids the interference of data at different times or different objects, providing an accurate data basis for subsequent feature extraction and processing. When the multi-modal AI model processes data, it can analyze based on data that accurately reflects the current state of the object, improving the pertinence and effectiveness of data processing, helping to more accurately extract information representing the object's features, and further enhancing the accuracy of the multi-modal data processing results, providing reliable data support for subsequent applications.

[0052] Specifically, in S2, the process of using the multi-modal AI model to process multi-modal data is monitored, including,

[0053] Collect the log records of the operation of the multi-modal AI model, and based on the log records, determine the key node information of each modal data in the feature extraction process, and the key node information at least includes the completion time of each feature extraction strategy in the feature extraction process;

[0054] For a certain modal data, when all its corresponding feature extraction strategies are executed, the completion time of the last feature extraction strategy is defined as the complete feature extraction completion time and the complete feature extraction completion time is output; wherein, the complete feature extraction completion time is used to capture the modal data that first completes feature extraction.

[0055] Through the above, this embodiment uses the method of collecting the running logs of the multi-modal AI model and defining the complete feature extraction completion time to achieve accurate capture of the modal data that first completes feature extraction. By collecting the log records, the key node information in the feature extraction process of each modal data is obtained, especially the completion time of the feature extraction strategy. Then, the moment when all feature extraction strategies of a certain modal data are executed is defined as the complete feature extraction completion time and output. Based on this, it is possible to accurately determine the modal data that first completes feature extraction. This method provides a reliable basis for subsequent determination of reference features and reference times, ensures the accuracy of key reference points in the data processing flow, makes the entire data processing process more logical and accurate, and helps improve the efficiency and quality of multi-modal data processing.

[0056] Specifically, in S2, defining the modal feature as the reference feature includes:

[0057] Monitoring the output of the complete feature extraction completion time, capturing the modal data corresponding to the first obtained complete feature extraction completion time as the modal data that first completes feature extraction, obtaining the modal feature corresponding to the modal data, performing preset integrity and accuracy verification on the modal feature, and after passing the verification, defining the modal feature as the reference feature and outputting it.

[0058] Through the above, this embodiment uses the means of monitoring the output of the complete feature extraction completion time and verifying the modal feature to achieve reliable definition of the reference feature. Monitoring the output of the complete feature extraction completion time, capturing the modal data and its features corresponding to the first completion time, and then performing preset integrity and accuracy verification on the modal feature. Only the modal feature that passes the verification is defined as the reference feature. Among them, the integrity and accuracy verification can be any existing method for feature verification. This process ensures the quality of the reference feature and avoids subsequent data processing deviations caused by incomplete or inaccurate features. The reliable reference feature provides an accurate reference standard for the redundancy removal operation, helps to more accurately identify and remove redundant features, improves the accuracy of multi-modal data processing, and makes the final processing result more accurately reflect the true features of the object.

[0059] Specifically, in S2, defining the moment when the modal data first completes feature extraction as the reference time includes:

[0060] Based on the moment recorded by the system clock, perform a preset calibration on the feature extraction complete moment corresponding to the modal data that first completes feature extraction among the captured data, and after completing this calibration, define the feature extraction complete moment as the reference moment and output it.

[0061] Through the above, in this embodiment, the method of using the system clock to calibrate the feature extraction complete moment of the modal data that first completes feature extraction realizes the accurate determination of the reference moment. Based on the moment recorded by the system clock, perform a preset calibration on the feature extraction complete moment of the modal data that first completes feature extraction among the captured data, and after calibration, define it as the reference moment. Among them, this calibration can be any existing time calibration method. The accurate reference moment provides a unified and accurate time benchmark for the multi-modal data processing process, enabling the stopping of feature extraction and the analysis of the extraction situations of other modal data to be based on accurate time, avoiding data processing chaos caused by time errors. This helps improve the accuracy and consistency of the entire data processing flow, ensures the reliability of the multi-modal data processing results, and provides a stable time basis for subsequent time-series-based data analysis and model training.

[0062] Specifically, in S3, the extracted features are a feature set composed of modal features obtained by each modal data through a preset feature extraction strategy before the reference moment; among them, the modal features at least include data features obtained based on statistical analysis and semantic features obtained based on machine learning.

[0063] Through the above, in this embodiment, the comprehensive integration of multi-modal data features is realized by clearly defining the extracted features. Define the set composed of modal features such as data features obtained based on statistical analysis and semantic features obtained based on machine learning, which are obtained by each modal data through a preset feature extraction strategy before the reference moment, as the extracted features. Among them, statistical analysis can be any existing statistical analysis technology, and machine learning can be any existing machine learning technology, such as deep learning technology. This comprehensive feature definition method covers information at different levels, enabling the extracted features to describe the features of the object more richly and comprehensively. It provides a rich data basis for subsequent redundancy removal operations, helps analyze and judge the redundancy of features from multiple dimensions, improves the accuracy and effectiveness of redundancy removal operations, and further improves the quality of multi-modal data processing, providing strong support for more in-depth data analysis and applications.

[0064] Specifically, in S3, the unexecuted feature extraction strategy is the feature extraction strategy that each modal data originally planned but has not executed after the reference moment, and this feature extraction strategy at least includes using different feature extraction algorithms and adjusting the parameters of feature extraction.

[0065] Through the above, in this embodiment, by clarifying the definition of the unexecuted feature extraction strategy, comprehensive management of the feature extraction strategy is achieved. It should be noted that in practical applications, the process of feature extraction from modal data is mostly a process of sequentially executing multiple preset different feature extraction strategies. The design of this embodiment is to define the strategies such as using different feature extraction algorithms and adjusting feature extraction parameters that were originally planned but not yet executed for each modal data after the reference time as unexecuted feature extraction strategies. This clear definition method helps to comprehensively grasp the situation of feature extraction during the data processing process, and enables reasonable analysis and screening of unexecuted strategies during the redundancy removal operation. It avoids blindly executing strategies that may generate redundant features, saves computing resources and time costs, and at the same time ensures the flexibility and scalability of feature extraction, enabling multi-modal data processing to select the most suitable feature extraction strategy according to the actual situation, and improving the efficiency and effect of data processing.

[0066] Specifically, in S4, the extracted features after performing the redundancy removal operation are defined as retained features, and the unexecuted feature extraction strategies after performing the redundancy removal operation are defined as retained feature extraction strategies, including:

[0067] Calculate the similarity between the extracted features and the reference features, determine the extracted features with a similarity greater than or equal to the preset similarity threshold as redundant features and remove them, and define the remaining extracted features as retained features;

[0068] Analyze the correlation between the features expected to be extracted by the unexecuted feature extraction strategies and the reference features and the retained features; when the correlation between the expected extracted features and the reference features or the retained features is greater than or equal to the preset correlation threshold, then determine that this unexecuted feature extraction strategy generates redundant features and remove it, and define the remaining unexecuted feature extraction strategies as retained feature extraction strategies.

[0069] Through the above, in this embodiment, the redundancy removal operation is performed on the extracted features and the unexecuted feature extraction strategies by using the existing similarity analysis technology and correlation analysis technology to calculate similarity and correlation, realizing the efficient reduction of multi-modal data. By calculating the similarity between the extracted features and the reference features, redundant features with a similarity higher than the threshold are removed to obtain retained features; at the same time, the correlation between the features expected to be extracted by the unexecuted feature extraction strategies and the reference features and the retained features is analyzed, and the strategies that may generate redundant features are removed to obtain retained feature extraction strategies. This redundancy removal process can effectively reduce the redundant information in the data, reduce the complexity of the data, and improve the efficiency of multi-modal data processing. The reduced data not only reduces the consumption of computing resources, but also improves the speed and accuracy of model training and data analysis, enabling the multi-modal AI model to process data more efficiently and output more valuable results.

[0070] Specifically, in S4, it further includes:

[0071] For the redundant features or unexecuted feature extraction strategies existing in different modal data after the redundancy removal operation, analyze the reflectivity of different modal data to the modal features, and use the modal features obtained from the modal data corresponding to the maximum value of the reflectivity as the finally retained features for describing the object; wherein, the calculation of the reflectivity is as follows:

[0072] For each modal data, respectively construct a mapping relationship model between the modal data and the redundant features; use the sample data of the modal data to train the constructed mapping relationship model to obtain a trained mapping relationship model; use the trained mapping relationship model to predict the redundant features in the modal data to obtain predicted values; calculate the error between the actual value and the predicted value of the redundant features in the modal data, and define the sum of the reciprocal of the error and a preset modal reflectivity correction value as the reflectivity of the modal data to the redundant features, and the greater the reflectivity, the stronger the reflectivity ability of the modal data to the redundant features. Among them, the modal reflectivity correction value is used to characterize the importance of the modal data in describing the object.

[0073] As a preferred implementation manner of this embodiment, the calculation of the reflectivity is specifically as follows:

[0074] For each modal data, respectively construct a mapping relationship model between the modal data and the redundant features. The mapping relationship model can adopt a regression model (such as polynomial regression) or a classification model (such as logistic regression classification, decision tree classification), and is specifically selected according to the data type of the redundant features and the characteristics of the modal data;

[0075] Use the sample data of the modal data to train the constructed mapping relationship model to obtain a trained mapping relationship model;

[0076] Use the trained mapping relationship model to predict the redundant features in the modal data to obtain predicted values;

[0077] Calculate the error between the actual value and the predicted value of the redundant features in the modal data. Error measurement methods such as mean square error and mean absolute error can be used;

[0078] Define the sum of the reciprocal of the error and a preset modal reflectivity correction value as the reflectivity of the modal data to the redundant features, that is wherein, Δ is the error, XZZ mtIt is the modified value of the modal reflectivity corresponding to the modal mt. This modified value of the modal reflectivity is set as an empirical value based on the common knowledge known to those skilled in the art, and is obtained by fitting the reflection degrees of different types of objects by different modal data. Exemplarily, for text-type objects, the modified value of the modal reflectivity of the corresponding text modal data is 0.3, and the modified value of the modal reflectivity of the corresponding image modal data is 0.1. At this time, the difference between the modified value of the modal reflectivity of the text modal data and the modified value of the modal reflectivity of the image modal data reflects the different reflection degrees of different modal data on the same object.

[0079] Through the above, in this embodiment, by constructing a mapping relationship model to calculate the reflectivity to select and retain modal features, in-depth processing of redundant features is achieved. For the redundant features or unexecuted feature extraction strategies in different modal data after the redundancy removal operation, a mapping relationship model between the modal data and the redundant features is constructed. After training, the predicted value of the redundant features is obtained. The reflectivity is obtained by calculating the error between the predicted value and the actual value and combining the modified value of the modal reflectivity. The features of the modal data with the largest reflectivity are selected as the retained features. This method can select the most representative features for the described object from multiple modalities, further optimizing the processing results of multi-modal data. It improves the quality and effectiveness of the data, enables the finally retained features to more accurately describe the object, and provides more reliable support for the use of multi-modal data in complex analysis and application scenarios.

[0080] It should be noted that when processing data based on a multi-modal AI model, the progress of feature extraction for different modal data may vary, and there is often redundant information among multi-modal data. If complete feature extraction is performed on all modal data, it will not only consume a large amount of computing resources and time, but may also affect the performance and efficiency of the model due to the existence of redundant features. Therefore, in combination with this embodiment, in a simple example, at the starting time t0, image modal data At0, audio modal data Bt0, and text modal data Ct0 for the same object are generated. It is monitored that at time t1, the image modal data At0 has completed all its corresponding feature extraction strategies. Then, this t1 moment is calibrated and determined as the reference time, and the modal feature TZA0 after the integrity and accuracy verification of the modal feature corresponding to the image modal data At0 is determined as the reference feature. At time t1, the feature extraction of the audio modal data Bt0 and the text modal data Ct0 is stopped, and the extracted feature tzB0 of the audio modal data Bt0 and the corresponding unexecuted feature extraction strategy WZXB0 are obtained, and the extracted feature tzC0 of the text modal data Ct0 and the corresponding unexecuted feature extraction strategy WZXC0 are obtained. Similarity analysis is performed on tzB0 and tzC0 in combination with TZA0 to complete the redundancy removal operation to obtain the retained features tzB1, tzC1, and TZA1. Correlation analysis is performed on WZXB0 and WZXC0 in combination with TZA0 to complete the redundancy removal operation to obtain the retained strategies WZXB1 and WZXC1. The retained strategies WZXB1 and WZXC1 are executed to obtain the corresponding features tzB2 and tzC2. tzB1 and tzB2 are packaged to obtain the modal state TZB of the audio modal data Bt0 and output. tzC1 and tzC2 are packaged to obtain the modal state TZC of the text modal data Ct0 and output. At this time, the modal state of the output image modal data At0 is TZA1.

[0081] The second aspect of this embodiment discloses a data processing system based on a multi-modal AI model as Figure 2 shown. This system is applicable to the data processing method based on the multi-modal AI model as described above. This system includes a starting time determination module, a reference situation determination module, a completion situation determination module, a redundancy removal module, and a post-redundancy execution module that are sequentially communicatively connected;

[0082] The starting time determination module is configured to: define the time when the multi-modal AI model starts running as the starting time, and there is multi-modal data for the same object at this starting time; wherein, the multi-modal data includes multiple different modal data;

[0083] The reference situation determination module is configured to: monitor the process of processing multimodal data using a multimodal AI model, capture the modal features of the modal data that first completes feature extraction in this process, define the modal features as reference features, and define the moment when the modal data first completes feature extraction as the reference moment;

[0084] The completion situation determination module is configured to: in the multimodal data, at the reference moment, stop the feature extraction of the multimodal data, and respectively analyze the completion situations of the feature extractions of all modal data except the modal data corresponding to the reference features, to obtain the extracted features of the modal data and the corresponding unexecuted feature extraction strategies;

[0085] The redundancy removal module is configured to: perform corresponding redundancy removal operations on the extracted features and the corresponding unexecuted feature extraction strategies using the reference features, define the extracted features after performing the redundancy removal operations as retained features, and define the unexecuted feature extraction strategies after performing the redundancy removal operations as retained feature extraction strategies;

[0086] The module executed after redundancy removal is configured to: only execute the retained feature extraction strategy, and after packing the obtained features with the corresponding retained features, obtain the modal features corresponding to all modal data except the modal data corresponding to the reference features in the multimodal data.

[0087] It should be noted that the data processing system based on the multimodal AI model in this embodiment corresponds to the aforementioned data processing method based on the multimodal AI model. Therefore, for the content not specifically described in the data processing system based on the multimodal AI model in this embodiment, such as but not limited to function definitions, working principles, and technical effects, etc., reference can be made to the records in the aforementioned data processing method based on the multimodal AI model, and this text will not elaborate here.

[0088] In summary, the data processing method and system based on the multimodal AI model in this embodiment utilize the time marking and feature extraction monitoring mechanism when processing data using the multimodal AI model to achieve efficient processing and redundancy elimination of multimodal data; determine the starting moment and multimodal data when the multimodal AI model is started, capture the modal data that first completes feature extraction as a reference by monitoring the feature extraction process, stop the extraction at the reference moment and analyze the extraction situations of other modal data; then, perform redundancy removal operations on the extracted features and the unexecuted feature extraction strategies based on the reference features, and finally execute the retained feature extraction strategy and pack the features; thereby effectively reducing data redundancy, avoiding unnecessary feature extraction calculations, improving the data processing efficiency of the multimodal AI model, enabling the finally obtained modal features to more accurately describe the object, providing a better data basis for subsequent analysis and applications based on multimodal data, and enhancing the performance of the entire multimodal data processing system.

[0089] In the embodiments provided in this application, it should be understood that the embodiments described herein can be implemented in hardware, software, firmware, middleware, code, or any suitable combination thereof. For hardware implementation, the processor can be implemented in one or more of the following units: application specific integrated circuit (ASIC), digital signal processor (DSP), digital signal processing device (DSPD), programmable logic device (PLD), field programmable gate array (FPGA), processor, controller, microcontroller, microprocessor, other electronic units designed to implement the functions described herein, or a combination thereof. For software implementation, part or all of the processes of the embodiments can be completed by instructing the relevant hardware through a computer program. When implemented, the above program can be stored in a computer-readable storage medium or transmitted as one or more instructions or codes on a computer-readable storage medium. The computer-readable storage medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transfer of a computer program from one place to another. The storage media can be any available medium accessible by a computer. The computer-readable storage medium can include, but is not limited to, RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer.

[0090] Finally, it should be noted that the above are only the preferred embodiments of this application and are not intended to limit this application. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of this application shall be included in the protection scope of this application.

Claims

1. A data processing method based on a multimodal AI model, characterized in that, The method includes the following steps: S1: Define the moment when the multi-modal AI model starts running as the starting moment, and at this starting moment, there is multi-modal data for the same object; wherein, the multi-modal data includes multiple different types of modal data; S2: Monitor the process of processing the multi-modal data using the multi-modal AI model, capture the modal features of the modal data that first completes feature extraction in this process, define this modal feature as the reference feature, and define the moment when this modal data first completes feature extraction as the reference moment; S3: In the multi-modal data, at the reference moment, stop the feature extraction of the multi-modal data, and respectively analyze the completion of feature extraction of all modal data except the modal data corresponding to the reference feature, to obtain the extracted features of the modal data and the corresponding unexecuted feature extraction strategies; S4: Use the reference feature to perform corresponding redundancy removal operations on the extracted features and the corresponding unexecuted feature extraction strategies, define the extracted features after performing the redundancy removal operation as the retained features, and define the unexecuted feature extraction strategy after performing the redundancy removal operation as the retained feature extraction strategy; S5: Only execute the retained feature extraction strategy, and after packing the obtained features with the corresponding retained features, obtain the modal features corresponding to all modal data except the modal data corresponding to the reference feature in the multi-modal data.

2. The data processing method based on a multimodal AI model according to claim 1, wherein In S1, at this starting moment, there is multi-modal data for the same object, including: After the multi-modal AI model starts running, collect the modal data corresponding to multiple different data sources of the same object to obtain the multi-modal data, and the multi-modal data is related to this object in terms of time and space, and this relationship is used to represent the data set of the same object at the same moment.

3. The data processing method based on the multimodal AI model according to claim 1, characterized in that In S2, monitor the process of processing the multi-modal data using the multi-modal AI model, including, Collect the log records of the operation of the multi-modal AI model, and based on these log records, determine the key node information of each modal data in the feature extraction process, and this key node information at least includes the completion time of each feature extraction strategy in the feature extraction process; For a certain modal data, when this modal data finishes executing all its corresponding feature extraction strategies, define the execution completion moment of its last feature extraction strategy as the feature extraction complete completion moment and output this feature extraction complete completion moment; wherein, the feature extraction complete completion moment is used to capture the modal data that first completes feature extraction.

4. The data processing method based on the multimodal AI model according to claim 3, wherein In S2, define this modal feature as the reference feature, including: Monitor the output of the feature extraction complete completion moment, capture the modal data corresponding to the first feature extraction complete completion moment obtained as the modal data that first completes feature extraction, and obtain the modal features corresponding to this modal data, perform preset integrity and accuracy verification on this modal feature, and after passing this verification, define this modal feature as the reference feature and output it.

5. The data processing method based on a multi-modal AI model according to claim 4, wherein In S2, define the moment when this modal data first completes feature extraction as the reference moment, including: Based on the moment recorded by the system clock, preset calibration is performed on the feature extraction complete moment corresponding to the modal data that first completes feature extraction among the captured ones. After completing this calibration, the feature extraction complete moment is defined as the reference moment and output.

6. The data processing method based on the multimodal AI model according to claim 1, wherein In S3, the extracted features are a set of features composed of modal features that have been obtained from each modal data through a preset feature extraction strategy before the reference moment; among them, the modal features at least include data features obtained based on statistical analysis and semantic features obtained based on machine learning.

7. The data processing method based on the multi-modal AI model according to claim 6, wherein, In S3, the unexecuted feature extraction strategy is the feature extraction strategy that each modal data originally planned but has not executed after the reference moment. This feature extraction strategy at least includes using different feature extraction algorithms and adjusting the parameters of feature extraction.

8. The data processing method based on a multimodal AI model according to claim 1, wherein In S4, defining the extracted features after performing the redundancy removal operation as the retained features and defining the unexecuted feature extraction strategy after performing the redundancy removal operation as the retained feature extraction strategy includes: Calculating the similarity between the extracted features and the reference features, determining the extracted features with a similarity greater than or equal to a preset similarity threshold as redundant features and removing them, and defining the remaining extracted features as the retained features; Analyzing the correlation between the features expected to be extracted by the unexecuted feature extraction strategy and the reference features and the retained features; when the correlation between the features expected to be extracted and the reference features or the retained features is greater than or equal to a preset correlation threshold, it is determined that the unexecuted feature extraction strategy generates redundant features and removes them, and the remaining unexecuted feature extraction strategy is defined as the retained feature extraction strategy.

9. The data processing method based on a multimodal AI model according to claim 1, wherein In S4, it also includes: For the redundant features or unexecuted feature extraction strategies existing in different modal data after the redundancy removal operation, analyzing the reflectivity of different modal data to this modal feature, and taking the modal feature obtained from the modal data corresponding to the maximum value of this reflectivity as the finally retained feature for describing the object; among them, the calculation of the reflectivity is: For each modal data, respectively constructing a mapping relationship model between this modal data and the redundant features; using the sample data of this modal data to train the constructed mapping relationship model to obtain a trained mapping relationship model; using the trained mapping relationship model to predict the redundant features in this modal data to obtain predicted values; calculating the error between the actual value and the predicted value of the redundant features in this modal data, and defining the sum of the reciprocal of this error and a preset modal reflectivity correction value as the reflectivity of this modal data to this redundant feature, and the greater the reflectivity, the stronger the ability of this modal data to reflect this redundant feature, where the modal reflectivity correction value is used to characterize the importance of the modal data for describing the object.

10. A data processing system based on a multimodal AI model, which is applicable to the data processing method based on the multimodal AI model described in any one of claims 1-9, characterized in that, The system includes a start moment determination module, a reference situation determination module, a completion situation determination module, a redundancy removal module, and a post-redundancy execution module that are sequentially communicatively connected; The starting time determination module is configured to: define the time when the multi-modal AI model starts running as the starting time, and at this starting time, there is multi-modal data for the same object; wherein, the multi-modal data includes multiple different types of modal data; The reference situation determination module is configured to: monitor the process of processing the multi-modal data using the multi-modal AI model, capture the modal features of the modal data that first completes feature extraction in this process, define these modal features as reference features, and define the time when the modal data first completes feature extraction as the reference time; The completion situation determination module is configured to: in the multi-modal data, at the reference time, stop the feature extraction of the multi-modal data, and respectively analyze the completion situations of the feature extractions of all modal data except the modal data corresponding to the reference features, to obtain the extracted features of the modal data and the corresponding unexecuted feature extraction strategies; The redundancy removal module is configured to: perform corresponding redundancy removal operations on the extracted features and the corresponding unexecuted feature extraction strategies using the reference features, define the extracted features after performing the redundancy removal operations as retained features, and define the unexecuted feature extraction strategies after performing the redundancy removal operations as retained feature extraction strategies; The post-redundancy execution module is configured to: only execute the retained feature extraction strategies, and after packing the obtained features with the corresponding retained features, obtain the modal features corresponding to all modal data in the multi-modal data except the modal data corresponding to the reference features.

Citation Information

Patent Citations

  • Data processing method, device and storage medium

    CN111768367B