Multi-mode AI interactive online learning system and method

Through the combination of feature fusion, cross-modal alignment and decision-making early warning modules, the modal differences and misalignment problems in multimodal AI interactive online learning are solved, and efficient data fusion and stability and reliability of the learning process are achieved.

CN120259832AInactive Publication Date: 2025-07-04SHANGHAI ZHIDAO KNOWLEDGE DIGITAL TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510742889.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In multimodal AI interactive online learning, modal feature representation differences, data misalignment and learning system instability problems lead to limited model performance and unstable learning process.

Method used

The feature fusion module is used for orderly progressive recognition and multimodal representation feature stitching and fusion, the cross-modal alignment module performs semantic and spatiotemporal data deviation adjustment, and the decision-making early warning module performs abnormal prediction and decision-making adjustment.

Benefits of technology

It improves the security and reliability of multimodal data fusion and the stability of learning systems, and enhances the performance and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259832A_ABST
    Figure CN120259832A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal AI interactive online learning system and method, and relates to the technical field of data processing, and the system comprises a feature fusion module, a cross-modal alignment module and a decision early warning module. The feature fusion module is used for collecting data representation features of different modalities in real time and performing splicing fusion of multi-modal representation features in real time in combination with associated information among the data; the cross-modal alignment module is used for calculating data deviations of different modals in real time according to the audio and image signals, and performing cross-modal alignment adjustment in real time according to the data deviations; and the decision early warning module is used for predicting whether the AI interactive online learning decision is abnormal or not in advance according to the data distribution, the scene and the task. According to the multi-modal AI interactive online learning system and method, data of different modalities are effectively fused, the situation that the data of different modalities are not aligned in terms of semantics and time and space is avoided, and the reliability and stability in the multi-modal AI interactive online learning process are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a multi-modal AI interactive online learning system and method. Background Art

[0002] Multi-modal AI interactive online learning is a way of using artificial intelligence technology to combine multiple data modalities including text, images, audio, and video, etc., for real-time interaction and learning. Multi-modal AI enhances the understanding ability and interactivity of intelligent systems by fusing data from different senses or interaction methods, including text, images, speech, and video, etc. This technology enables the model to simultaneously process and understand multiple types of data, thereby achieving more natural and efficient human-computer interaction.

[0003] With the rapid development of science and technology, there are still some deficiencies in multi-modal AI interactive online learning: 1. There are differences in feature representations of data in different modalities. Simple splicing or fusion methods may not be able to fully mine the correlation information between data, resulting in limited model performance and inability to effectively fuse data in different modalities in a timely manner; 2. Data in different modalities may be misaligned semantically and spatio-temporally, including the problem of deviation between audio and image signals in videos, and accurate cross-modal alignment cannot be achieved, affecting the performance of the multi-modal AI interactive online learning system; 3. When the multi-modal AI interactive online learning system faces new data distributions, scenarios, and tasks, due to the diverse combinations and changes between different modalities, the online learning system cannot quickly adapt to all changes, which may affect the reliability and stability during the learning process.

[0004] Therefore, a multi-modal AI interactive online learning system and method are proposed to solve the above problems. Summary of the Invention

[0005] The main purpose of the present invention is to provide a multi-modal AI interactive online learning system and method to solve the problems raised in the above background.

[0006] To achieve the above purpose, the technical solution adopted by the present invention is: a multi-modal AI interactive online learning system and method, including a feature fusion module, a cross-modal alignment module, and a decision warning module; The feature fusion module is used to collect real-time data representation features of different modalities, and analyze in real-time whether there are differences in the data representation features of different modalities. When analyzing, a hierarchical progressive recognition method is adopted to calculate the representation feature deviation in real-time, and the multi-modal representation features are spliced and fused in real-time in combination with the correlation information between data; The cross-modal alignment module is used to, when the splicing and fusion are abnormal, collect the audio and image signals of AI interaction data in semantics and spatio-temporal dimensions in different modalities in real time, calculate the data deviation of different modalities in real time according to the audio and image signals, and perform cross-modal alignment adjustment in real time according to the data deviation; The decision-making and early warning module is used to collect the data distribution, scenarios and tasks of the multi-modal AI interactive online learning system in real time, and predict in advance whether the AI interactive online learning decision is abnormal according to the data distribution, scenarios and tasks in real time. When it is abnormal, adjust the AI interactive online learning decision in real time according to the combinations and diversities between different modalities.

[0007] The feature fusion module includes a multi-modal AI interaction learning unit, a representation feature monitoring unit, a representation feature early warning unit and a multi-modal fusion unit; The multi-modal AI interaction learning unit is used to collect the representation features of different modalities in real time through a data collector. The different modalities include images, texts, audios, videos and sensor signals, and set the multi-modal AI interactive online learning decision, which is set in combination with the data distribution, scenarios and tasks.

[0008] The representation feature monitoring unit is used to analyze in real time whether there are differences in the data representation features of different modalities. The analysis method is as follows: Adopt a hierarchical progressive recognition method to calculate the representation feature difference in real time. Hierarchical progressive recognition is to perform multi-region hierarchical processing on the representation features of two different modalities to obtain the feature difference value between the representation features of the two different modalities , and the calculation formula is as follows: ; Among them, X and Y respectively represent the feature matrices of two different modalities, represents the kernel function, and respectively represent the th representation features of different modalities, and respectively represent the number of representation features of two different modalities. Set the feature difference threshold. If the feature difference value is greater than the feature difference threshold, it means that there are differences between the data representation features of different modalities. If not, it means that there are no differences between the data representation features of different modalities; The representation feature early warning unit is used to report a voice alarm reminder issued by the system when there are differences between the data representation features of different modalities.

[0009] The multi-modal fusion unit is used to fuse the data of different modalities. The fusion formula is as follows: ; Among them, Represents the data fusion result in different modalities, The dynamic weight of the th modality, and represents the feature vector of the

[0010] The cross-modal alignment module includes a multi-modal acquisition and monitoring unit, a multi-modal deviation unit, and a cross-modal alignment and adjustment unit; The multi-modal acquisition and monitoring unit is used to collect the audio and image signals of AI interaction data in different modalities in terms of semantics, space, and time through a data acquisition instrument.

[0011] The multi-modal deviation unit is used to calculate the representation feature deviation of different modalities in real time based on the audio and image signals, and obtain a representation feature deviation correction value. The calculation method of the representation feature deviation correction value is as follows: S1: Calculate the first-order coefficient of the regression equation, and the formula is as follows: ; where is the first-order coefficient, and represents the quadratic relationship between the first representation feature pitch angle value and the correction angle data, specifically as follows: represents the basic value of the deviation of the first representation feature, represents the basic value of the deviation of the first representation feature reaching the specified point, represents the basic value of the deviation of the second representation feature, represents the basic value of the deviation of the second representation feature reaching the specified point, represents the basic value of the deviation of the representation feature correction angle at the current moment. Here, it represents the deviation values of different time points and different representation feature correction angles, and the correction angle represents the pitch correction angle of the representation feature in different modalities and the correction angle between the corresponding representation features; S2: Calculate the second-order coefficient of the regression equation, and the formula is as follows: ; where represents the second-order coefficient, and represents the binary quadratic relationship between the first representation feature pitch angle value and the correction angle data; ; where represents the constant term coefficient, represents the average value of the correction angle deviation value at the current moment. Here, the correction angle represents the corrected pitch angle and the correction angle between the specified correction points, represents the th deviation value of the representation feature correction angle, represents the average value of all deviation deformation data in the pitch angle dataset when performing cross-modal data alignment, represents the average value of all deviation angle data in the corrected angle dataset, and n represents the correction angle between the nth corresponding representation feature and the specified correction point under different modalities; S3: Establish a regression equation based on the linear term coefficient, quadratic term coefficient, and constant term coefficient, so as to obtain the correction deviation value between the representation feature deviation correction and the corresponding pitch angle of the representation feature under different modalities; The cross-modal alignment adjustment unit is used to perform real-time alignment of cross-modal representation features according to the calculated correction deviation value. The real-time alignment adopts a hierarchical progressive method, that is, hierarchical alignment is performed in real time according to the calculated correction deviation value, and real-time tracking is carried out through a data tracker.

[0012] The decision warning module includes a multi-source data acquisition unit, an interactive decision warning unit, an abnormal decision adjustment unit, and a decision adaptive update unit; The multi-source data acquisition unit is used to collect the data distribution, scenarios, and tasks of the multi-modal AI interactive online learning system in real time through a data acquisition instrument.

[0013] The interactive decision warning unit is used to predict in advance whether the AI interactive online learning decision is abnormal according to the data distribution, scenarios, and tasks in real time. The prediction method is as follows: Step 1: Calculate the deviation value under the data distribution. The calculation formula is as follows: ; where, represents the data distribution deviation value at the current moment, and represent the mean and covariance matrix of the training data respectively. Set a deviation safety threshold. If is greater than the deviation safety threshold, it means that the data distribution is abnormal. Otherwise, it means that the data distribution is normal. represents the training data at the current moment (that is, the training data to be evaluated for predicting the AI interactive online learning decision at the current moment); Step 2: Combine the scenario and the historical scenario feature set to perform real-time prediction of the scenario, score the AI interactive online learning decision matching the scenario, and calculate the score value. The calculation formula is as follows: ; where, represents the score value, represents the historical scenario feature set, and the set includes user geographical location, device type, and timestamp. represents the current scenario feature, Represents the standard scenario features corresponding to the AI interactive online learning decision, sets the scenario adaptation score threshold. If is less than the scenario adaptation score threshold, it is predicted that the execution of the AI interactive online learning decision in the corresponding scenario is abnormal, and the reporting system issues a voice reminder alarm. Otherwise, it is predicted that the execution of the AI interactive online learning decision in the corresponding scenario is normal; Step 3: Combine the data distribution at the current moment and the AI interactive online learning decision criteria in the corresponding scenario to predict in real time whether the AI interactive online learning decision in the next cycle can be executed. Record three minutes as a cycle, and based on the data distribution and scenarios in three cycles, predict in real time the data distribution and scenarios in the corresponding area of the AI interactive online learning system. If the average value of the deviation values of the data distributions in the three cycles is greater than the average value of the deviation safety thresholds, and the average value of the scenario score values is less than the average value of the scenario adaptation score thresholds, it is determined that the execution of the AI interactive online learning decision in the corresponding scenario is abnormal. Otherwise, it is determined that the execution of the AI interactive online learning decision in the corresponding scenario is normal.

[0014] The abnormal decision adjustment unit is used to receive in real time the predicted judgment result of the execution of the AI interactive online learning decision through a data receiver, and adjust the AI interactive online learning decision in real time according to the combination and variety among different modalities; The decision adaptive update unit is used to adjust the AI interactive online learning decision in real time according to the task requirements among different modalities. The AI interactive online learning decision is determined by combining the data distribution, scenarios, and tasks in different modalities. The tasks include the items for online learning and the problems solved during interaction.

[0015] A usage method of a multi-modal AI interactive online learning method system, including the following steps: Step 1: Enter the feature fusion module, configure the multi-modal AI interactive online learning control terminal, represent features by collecting data in different modalities in real time, analyze the differences in data representation features in different modalities in real time, mine the correlation information between data in real time for multi-modal representation feature splicing and fusion, effectively fuse the data in different modalities in a timely manner, and formulate an interaction decision after fusion; Step 2: Enter the cross-modal alignment module, collect the audio and image signals in the semantics and space-time of the AI interaction data in different modalities in real time, calculate the data deviation in different modalities in real time, and perform cross-modal alignment adjustment according to the data deviation in real time; Step 3: Enter the decision warning module, collect the data distribution, scenarios, and tasks of the multi-modal AI interactive online learning system in real time, predict in advance whether the interaction decision is abnormal according to the data distribution, scenarios, and tasks in real time, and adjust the interaction decision in real time according to the combination and variety among different modalities when it is abnormal.

[0016] The present invention has the following beneficial effects: 1. In the present invention, by setting a feature fusion module, when performing multi-modal AI interactive online learning operations, by adopting a hierarchical progressive recognition method to calculate the representation feature deviation in real time, and combining the correlation information between data to splice and fuse the multi-modal representation features in real time, the system can effectively fuse different modal data in a timely manner. The fused multi-modal data is used to make AI interactive online learning decisions, so that when there are differences in feature representation of different modal data, different modal data can be effectively fused in a timely manner, increasing the safety and reliability of multi-modal data fusion.

[0017] 2. In the present invention, by setting a cross-modal alignment module, when performing multi-modal AI interactive online learning operations, by calculating the data deviation of different modalities in real time according to audio and image signals, it is possible to avoid the situation where different modal data is misaligned in semantics and space-time, so that the audio and image signals in the video during AI interactive learning can reduce the deviation, further increasing the cross-modal alignment effect and improving the performance of the multi-modal AI interactive online learning system.

[0018] 3. In the present invention, by setting a decision warning module, when performing multi-modal AI interactive online learning operations, by predicting in advance whether the AI interactive online learning decision is abnormal according to the data distribution, scene and task in real time, and adjusting the AI interactive online learning decision in real time, during the multi-modal AI interactive online learning process, it is possible to make decision adjustments in real time according to the data distribution, scene and task that appear during the interaction process, improving the reliability and stability during the multi-modal AI interactive online learning process, and having good practicality. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 is a schematic diagram of the architecture of a multi-modal AI interactive online learning system of the present invention; Figure 2 is an overall flow schematic diagram of the usage method of a multi-modal AI interactive online learning method system of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific embodiments.

[0021] Embodiment 1 Please refer to Figures 1 to 2 as shown: A multi-modal AI interactive online learning system and method includes a feature fusion module, a cross-modal alignment module and a decision warning module; The feature fusion module is used to represent features by collecting data of different modalities in real time, and analyze in real time whether there are differences in the data representation features of different modalities. When analyzing, a hierarchical progressive recognition method is adopted to calculate the deviation of the representation features in real time, and the concatenation and fusion of multi-modal representation features are carried out in real time in combination with the correlation information between the data; The cross-modal alignment module is used when the concatenation and fusion are abnormal. By collecting the audio and image signals of the AI interaction data in different modalities in terms of semantics and space-time in real time, the data deviation of different modalities is calculated in real time, and the cross-modal alignment adjustment is carried out in real time according to the data deviation; The decision warning module is used to collect the data distribution, scenarios and tasks of the multi-modal AI interactive online learning system in real time, and predict in advance whether the AI interactive online learning decision is abnormal according to the data distribution, scenarios and tasks in real time. When it is abnormal, the AI interactive online learning decision is adjusted in real time according to the combination and variety between different modalities.

[0022] The feature fusion module includes a multi-modal AI interaction learning unit, a representation feature monitoring unit, a representation feature warning unit and a multi-modal fusion unit; The multi-modal AI interaction learning unit is used to collect the representation features of different modalities in real time through a data collector. The different modalities include images, texts, audios, videos and sensor signals, and set the multi-modal AI interactive online learning decision, which is set in combination with the data distribution, scenarios and tasks.

[0023] The representation feature monitoring unit is used to analyze in real time whether there are differences in the data representation features of different modalities. The analysis method is as follows:

[0024] The hierarchical progressive recognition method is adopted to calculate the representation feature difference in real time. The hierarchical progressive recognition is to perform multi-region hierarchical processing on the representation features of two different modalities to obtain the feature difference value between the representation features of the two different modalities The calculation formula is as follows: ; Among them, X and Y respectively represent the feature matrices of two different modalities, represents the kernel function, and respectively represent the th representation features of different modalities, and respectively represent the number of representation features of the two different modalities. Set the feature difference threshold. If the feature difference value is greater than the feature difference threshold, it means that there are differences between the data representation features of different modalities. Otherwise, it means that there are no differences between the data representation features of different modalities; The feature warning unit is used to report to the system to issue a voice alarm reminder when there are differences between the data representation features of different modalities.

[0025] The multimodal fusion unit is used to fuse data of different modalities, and the fusion formula is as follows: ; Wherein, represents the data fusion result under different modalities, the dynamic weight of the th modality, represents the th feature vector of the modality. Through combining the correlation information between data, the splicing fusion of multimodal representation features is carried out in real time. The correlation information includes the cohesion of text, the cohesion of video frame sets, and the connection of image representation features, enabling the system to effectively fuse data of different modalities in a timely manner. The fused multimodal data formulates AI interactive online learning decisions, enabling real-time deviation splicing and fusion when there are differences in feature representation of data of different modalities, fully mining the correlation information between data to increase the accuracy of fusion, and effectively fusing data of different modalities in a timely manner.

[0026] Embodiment 2 Please refer to Figures 1 to 2 as shown: Based on Embodiment 1, the cross-modal alignment module includes a multimodal acquisition and monitoring unit, a multimodal deviation unit, and a cross-modal alignment and adjustment unit; The multimodal acquisition and monitoring unit is used to collect audio and image signals in semantics and spatio-temporal of AI interaction data under different modalities in real time through a data collector.

[0027] The multimodal deviation unit is used to calculate the representation feature deviation of different modalities in real time according to the audio and image signals, and obtain a representation feature deviation correction value. The calculation method of the representation feature deviation correction value is as follows: S1: Calculate the first-order term coefficient of the regression equation, and the formula is as follows: ; Wherein, is the first-order term coefficient, and represents the unary quadratic relationship formed between the first representation feature pitch angle value and the correction angle data, specifically as follows: represents the basic value of the deviation of the first representation feature, represents the basic value of the deviation of the first representation feature reaching the specified point, represents the basic value of the deviation of the second representation feature, represents the basic value of the deviation of the second representation feature reaching the specified point, Represents the basic value indicating the deviation of the feature correction angle at the current moment, which represents the deviation values of different time points and different feature correction angles, and the correction angle represents the pitch correction angle of the representation feature in different modalities and the correction angle between the corresponding representation features; S2: Calculate the quadratic term coefficient of the regression equation, and the formula is as follows: ; Among them, Represents the quadratic term coefficient, and represents the binary quadratic relationship formed between the pitch angle value of the first representation feature and the correction angle data; ; Among them, Represents the constant term coefficient, Represents the average value of the correction angle deviation value at the current moment. Here, the correction angle represents the corrected pitch angle and the correction angle between the specified correction points, Represents the Deviation value of the correction angle of the th representation feature; Represents the average value of all deviation deformation data in the pitch angle dataset during cross-modal data alignment, Represents the average value of all deviation angle data in the correction angle dataset. n represents the correction angle between the nth corresponding representation feature and the specified correction point in different modalities; S3: Based on the linear term coefficient, quadratic term coefficient, and constant term coefficient, establish a regression equation to obtain the correction deviation value between the representation feature deviation correction and the pitch angle of the corresponding representation feature in different modalities; The cross-modal alignment adjustment unit is used to perform real-time alignment of cross-modal representation features according to the calculated correction deviation value. The real-time alignment adopts a hierarchical progressive method, that is, real-time hierarchical alignment is performed according to the calculated correction deviation value, and real-time tracking is performed through a data tracker. The data deviation of different modalities is calculated in real-time according to the audio and image signals, and cross-modal alignment adjustment is performed in real-time according to the data deviation. The cross-modal alignment adjustment adopts a hierarchical progressive method to avoid misalignment of data in different modalities in terms of semantics and space-time.

[0028] Embodiment III Please refer to Figures 1 to 2 As shown: Based on the basis of Embodiment I, the decision warning module includes a multi-source data acquisition unit, an interactive decision warning unit, an abnormal decision adjustment unit, and a decision adaptive update unit; The multi-source data acquisition unit is used to collect the data distribution, scenarios, and tasks of the multi-modal AI interactive online learning system in real-time through a data acquisition instrument.

[0029] The interactive decision-making early warning unit is used to predict in advance whether the AI interactive online learning decision is abnormal in real time according to the data distribution, scenarios, and tasks. The prediction method is as follows: Step 1: Calculate the deviation value under the data distribution. The calculation formula is as follows: ; Among them, represents the data distribution deviation value at the current moment, and represent the mean and covariance matrix of the training data respectively. Set the deviation safety threshold. If is greater than the deviation safety threshold, it means that the data distribution is abnormal. Otherwise, it means that the data distribution is normal. represents the training data at the current moment; Step 2: Combine the scenario and the historical scenario feature set to perform real-time prediction of the scenario, score the AI interactive online learning decision that matches the scenario, and calculate the score value. The calculation formula is as follows: ; Among them, represents the score value, represents the historical scenario feature set, and the set includes user geographical location, device type, and timestamp. represents the current scenario feature, represents the standard scenario feature corresponding to the AI interactive online learning decision. Set the scenario adaptation score threshold. If is less than the scenario adaptation score threshold, it is predicted that the AI interactive online learning decision execution in the corresponding scenario is abnormal, and the system is reported to issue a voice reminder alarm. Otherwise, it is predicted that the AI interactive online learning decision execution in the corresponding scenario is normal; Step 3: Combine the data distribution at the current moment and the AI interactive online learning decision standard in the corresponding scenario to predict in real time whether the AI interactive online learning decision in the next cycle can be executed. Record three minutes as a cycle. According to the data distribution and scenarios in three cycles, predict the data distribution and scenarios in the corresponding area of the AI interactive online learning system in real time. If the average value of the data distribution deviation values in the three cycles is greater than the average value of the deviation safety threshold, and the average value of the scenario score values is less than the average value of the scenario adaptation score threshold, it is determined that the AI interactive online learning decision execution in the corresponding scenario is abnormal. Otherwise, it is determined that the AI interactive online learning decision execution in the corresponding scenario is normal.

[0030] The abnormal decision adjustment unit is used to receive the predicted AI interactive online learning decision execution judgment result in real time through the data receiver, and adjust the AI interactive online learning decision in real time according to the combination and variety among different modalities; The decision-making adaptive update unit is used to adjust the AI interactive online learning decision in real time according to the task requirements between different modalities. The AI interactive online learning decision is determined by combining the data distribution, scenarios, and tasks in different modalities. The tasks include the items of online learning and the problems solved during interaction. By predicting in advance whether the AI interactive online learning decision is abnormal according to the data distribution, scenarios, and tasks in real time, and adjusting the AI interactive online learning decision in real time according to the combination and variety between different modalities when it is abnormal, during the multi-modal AI interactive online learning process, the decision can be adjusted in real time according to the data distribution, scenarios, and tasks that appear during the interaction, enabling the system to adapt to the combination and variety between different modalities.

[0031] In the present invention, there is provided a multi-modal AI interactive online learning system and method. When the system operates, a multi-modal AI interactive online learning control terminal is configured to enter the feature fusion module. The multi-modal AI interactive online learning control terminal is configured to represent features by collecting data of different modalities in real time, analyze the differences in the data representation features of different modalities in real time, and mine the correlation information between data in real time for splicing and fusing multi-modal representation features, so as to effectively fuse the data of different modalities in a timely manner. After fusion, an interaction decision is made. By collecting the data representation features of different modalities in real time and analyzing whether there are differences in the data representation features of different modalities in real time, a hierarchical progressive recognition method is adopted to calculate the deviation of the representation features in real time during the analysis, and the splicing and fusion of multi-modal representation features are carried out in real time in combination with the correlation information between data, so that the system can effectively fuse the data of different modalities in a timely manner. The fused multi-modal data makes an AI interactive online learning decision, so that when there are differences in the feature representation of data of different modalities, deviation splicing can be carried out in real time to achieve fusion, fully mining the correlation information between data to increase the accuracy of fusion, effectively fusing the data of different modalities in a timely manner, and increasing the safety and reliability of multi-modal data fusion; Enter the cross-modal alignment module. By collecting the audio and image signals in terms of semantics and space-time of AI interaction data under different modalities in real time, calculate the data deviation of different modalities in real time, and perform cross-modal alignment adjustment according to the data deviation. By collecting the data representation features of different modalities in real time and analyzing whether there are differences in the data representation features of different modalities in real time, a hierarchical progressive recognition method is adopted to calculate the deviation of the representation features in real time during the analysis, and the splicing and fusion of multi-modal representation features are carried out in real time in combination with the correlation information between data, so that the system can effectively fuse the data of different modalities in a timely manner. The fused multi-modal data makes an AI interactive online learning decision, so that when there are differences in the feature representation of data of different modalities, deviation splicing can be carried out in real time to achieve fusion, fully mining the correlation information between data to increase the accuracy of fusion, effectively fusing the data of different modalities in a timely manner, and increasing the safety and reliability of multi-modal data fusion;Enter the decision warning module. By collecting the data distribution, scenarios, and tasks of the multi-modal AI interactive online learning system in real time, it predicts in advance whether the interaction decision is abnormal according to the data distribution, scenarios, and tasks in real time. When an abnormality occurs, it adjusts the interaction decision in real time according to the combinations and diversities among different modalities. By collecting the data distribution, scenarios, and tasks of the multi-modal AI interactive online learning system in real time, it predicts in advance whether the AI interactive online learning decision is abnormal according to the data distribution, scenarios, and tasks in real time. When an abnormality occurs, it adjusts the AI interactive online learning decision in real time according to the combinations and diversities among different modalities. During the multi-modal AI interactive online learning process, it can make decision adjustments in real time according to the data distribution, scenarios, and tasks that appear during the interaction, enabling the system to adapt to the combinations and diversities among different modalities, further improving the reliability, stability, and practicability during the multi-modal AI interactive online learning process.

[0032] The above shows and describes the basic principles, main features, and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification only illustrates the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A multimodal AI interactive online learning system, characterized in that, The system includes a feature fusion module, a cross-modal alignment module, and a decision warning module; The feature fusion module is used to represent features by collecting data of different modalities in real time, and analyze in real time whether there are differences in the data representation features of different modalities. When analyzing, a hierarchical progressive recognition method is used to calculate the representation feature deviation in real time, and the multi-modal representation features are spliced and fused in real time in combination with the correlation information between the data; The cross-modal alignment module is used when the splicing and fusion is abnormal. By collecting the audio and image signals of the AI interaction data in different modalities in terms of semantics, space, and time in real time, calculate the data deviation of different modalities in real time according to the audio and image signals, and perform cross-modal alignment adjustment in real time according to the data deviation; The decision warning module is used to collect the data distribution, scenarios, and tasks of the multi-modal AI interactive online learning system in real time, and predict in advance whether the AI interactive online learning decision is abnormal according to the data distribution, scenarios, and tasks in real time. When it is abnormal, adjust the AI interactive online learning decision in real time according to the combinations and diversities between different modalities; The cross-modal alignment module includes a multi-modal acquisition monitoring unit, a multi-modal deviation unit, and a cross-modal alignment adjustment unit; The multi-modal deviation unit is used to calculate the representation feature deviation of different modalities in real time according to the audio and image signals to obtain a representation feature deviation correction value. The calculation method of the representation feature deviation correction value is as follows: S1: Calculate the first-order term coefficient of the regression equation. The formula is as follows: ; Among them, is the coefficient of the first-order term, and represents the quadratic relationship formed between the first characteristic pitch angle value and the correction angle data, specifically as follows: represents the basic value of the deviation of the first characteristic, represents the basic value of the deviation of the first characteristic reaching the specified point, represents the basic value of the deviation of the second characteristic, represents the basic value of the deviation of the second characteristic reaching the specified point, represents the basic value of the deviation of the characteristic correction angle at the current moment. Here, it represents the deviation values at different time points and different characteristic correction angles, and the correction angle represents the pitch correction angle of the characteristic in different modes and the correction angle between the corresponding characteristics; S2: Calculate the second-order term coefficient of the regression equation. The formula is as follows: ; Among them, B represents the second-order term coefficient, and represents the binary quadratic relationship formed between the first representation feature pitch angle value and the correction angle data; ; Among them, represents the coefficient of the constant term, represents the average value of the deviation of the correction angle at the current moment. Here, the correction angle represents the corrected pitch angle and the correction angle between the specified correction point, represents the th deviation value of the characteristic correction angle; represents the average value of all deviation deformation data in the pitch angle dataset during cross-modal data alignment, represents the average value of all deviation angle data in the correction angle dataset. n represents the correction angle between the nth corresponding representation feature and the specified correction point under different modalities; S3: Establish a regression equation according to the first-order term coefficient, the second-order term coefficient, and the constant term coefficient, so as to obtain the correction deviation value of the representation feature deviation when different modalities are corrected and the corresponding representation feature pitch angle.

2. The system according to claim 1, wherein: The feature fusion module includes a multi-modal AI interaction learning unit, a representation feature monitoring unit, a representation feature warning unit, and a multi-modal fusion unit; The multi-modal AI interaction learning unit is used to collect the representation features of different modalities in real time through a data collector. The different modalities include images, texts, audios, videos, and sensor signals, and set the multi-modal AI interactive online learning decision, which is set in combination with the data distribution, scenarios, and tasks.

3. The system according to claim 2, characterized in that: The representation feature monitoring unit is used to analyze in real time whether there are differences in the data representation features of different modalities. The analysis method is as follows: The feature difference calculation of the representation features is performed in real time by using a hierarchical progressive recognition method. The hierarchical progressive recognition is to perform hierarchical processing on the representation features of two different modalities in multiple regions to obtain the feature difference value between the representation features of the two different modalities , and the calculation formula is as follows: ; Among them, X and Y respectively represent the feature matrices of two different modalities. represents the kernel function. and respectively represent the th representation features of different modalities. and respectively represent the numbers of representation features of two different modalities. Set a feature difference threshold. If the feature difference value is greater than the feature difference threshold, it means that there are differences between the data representation features of different modalities. Otherwise, it means that there are no differences between the data representation features of different modalities. represents the spatial norms of two different modalities, which are used to measure the distance between the centers of the data representation feature distributions of different modalities. The representation feature warning unit is used to report a voice alarm reminder issued by the system when there are differences between the data representation features of different modalities.

4. The system according to claim 3, characterized in that: The multi-modal fusion unit is used to fuse the data of different modalities. The fusion formula is as follows: ; Among them, represents the data fusion result in different modalities, the dynamic weight of the th modality, represents the feature vector of the th modality. FFN represents a feed-forward neural network used for non-linear transformation of the input data representation features, represents a normalization layer, represents an attention function used to calculate the interaction relationship between and H. H represents the global feature set in different modalities.

5. The system according to claim 1, characterized in that: The multi-modal acquisition monitoring unit is used to collect the audio and image signals of the AI interaction data in different modalities in terms of semantics, space, and time through a data collector.

6. The system according to claim 1, characterized in that: The cross-modal alignment adjustment unit is used to perform real-time alignment of cross-modal representation features according to the calculated correction deviation value. The real-time alignment adopts a hierarchical progressive method, that is, hierarchical alignment is performed in real time according to the calculated correction deviation value, and real-time tracking is carried out through a data tracker.

7. The system according to claim 1, characterized in that: The decision-making and early warning module includes a multi-source data acquisition unit, an interactive decision-making and early warning unit, an abnormal decision-making adjustment unit, and a decision-making adaptive update unit; The multi-source data acquisition unit is used to collect the data distribution, scenarios, and tasks of the multi-modal AI interactive online learning system in real time through a data collector.

8. The system according to claim 7, wherein: The interactive decision-making and early warning unit is used to predict in advance whether the AI interactive online learning decision is abnormal according to the data distribution, scenarios, and tasks in real time. The prediction method is as follows: Step 1: Calculate the deviation value under the data distribution. The calculation formula is as follows: ; Among them, represents the data distribution deviation value at the current moment, and respectively represent the mean value and covariance matrix of the training data. Set a deviation safety threshold. If is greater than the deviation safety threshold, it indicates that the data distribution is abnormal. Otherwise, it indicates that the data distribution is normal. represents the training data at the current moment; Step 2: Combine the scenario and the historical scenario feature set to perform real-time prediction of the scenario, score the AI interactive online learning decision that matches the scenario, and calculate the score value. The calculation formula is as follows: ; Among them, represents the scoring value, represents the set of historical scenario features, which includes user geographical location, device type, and timestamp, represents the current scenario feature, represents the standard scenario feature corresponding to the AI interactive online learning decision, and sets the scenario adaptation scoring threshold. If is less than the scenario adaptation scoring threshold, it is predicted that the execution of the AI interactive online learning decision in the corresponding scenario is abnormal, and the system is reported to issue a voice reminder alarm. Otherwise, it is predicted that the execution of the AI interactive online learning decision in the corresponding scenario is normal; Step 3: Combine the data distribution at the current moment and the AI interactive online learning decision criteria under the corresponding scenario to predict in real time whether the AI interactive online learning decision in the next cycle is executable. Record three minutes as a cycle, and predict the data distribution and scenarios in the corresponding area of the AI interactive online learning system in real time according to the data distribution and scenarios in three cycles. If the average value of the deviation values of the data distribution in three cycles is greater than the average value of the deviation safety threshold, and the average value of the scenario score values is less than the average value of the scenario adaptation score threshold, it is determined that the execution of the AI interactive online learning decision under the corresponding scenario is abnormal; otherwise, it is determined that the execution of the AI interactive online learning decision under the corresponding scenario is normal.

9. The system according to claim 8, wherein: The abnormal decision-making adjustment unit is used to receive the predicted judgment result of the execution of the AI interactive online learning decision in real time through a data receiver, and adjust the AI interactive online learning decision in real time according to the combination and variety among different modalities; The decision-making adaptive update unit is used to adjust the AI interactive online learning decision in real time according to the task requirements among different modalities. The AI interactive online learning decision is determined by combining the data distribution, scenarios, and tasks under different modalities. The tasks include the items of online learning and the problems solved during interaction.

10. A method for using a multi-modal AI interactive online learning system according to any one of claims 1-9, characterized in that, It includes the following steps: Step 1: Enter the feature fusion module, configure the multi-modal AI interactive online learning control terminal, collect the data representation features of different modalities in real time, analyze the differences in the data representation features of different modalities in real time, mine the correlation information between the data in real time for splicing and fusing the multi-modal representation features, effectively fuse the data of different modalities in a timely manner, and formulate an interactive decision after fusion; Step 2: Enter the cross-modal alignment module, collect the audio and image signals in semantics and space-time of the AI interaction data under different modalities in real time, calculate the data deviation of different modalities in real time, and perform cross-modal alignment adjustment according to the data deviation in real time. Step 3: Enter the decision warning module. By collecting the data distribution, scenarios, and tasks of the multi-modal AI interactive online learning system in real time, predict in advance whether the interaction decision is abnormal according to the data distribution, scenarios, and tasks in real time, and adjust the interaction decision in real time according to the combination and variety among different modalities when an abnormality occurs.

Citation Information

Cited By

  • Online service interaction anomaly analysis method based on AI server and big data

    CN121000773A