Multi-modal big data analysis processing method and device based on artificial intelligence
By employing scene awareness and dynamic weight fusion mechanisms, the problems of scene adaptability, emotion-semantic association, and cross-modal error correction in multimodal data processing are solved, achieving higher accuracy and reliability in speech input recognition and conversion.
Patent Information
- Application Number
- CN202511739114.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-25
- Publication Date
- 2026-02-27
- Estimated Expiration
- 2045-11-25
AI Technical Summary
Existing multimodal data processing technologies suffer from insufficient scene adaptability, weak emotion-semantic association, fixed modal fusion methods, and lack of cross-modal error correction in the recognition and conversion of voice input, resulting in insufficient recognition accuracy and reliability.
By employing an AI-based multimodal big data analysis method, through scene perception, sentiment-semantic joint analysis, dynamic weight fusion, and cross-modal error correction mechanisms, the weights of modal data are dynamically adjusted and consistency verification and correction are performed, thereby improving the accuracy and reliability of multimodal data processing.
It improves the accuracy and scene adaptability of multimodal data processing, enhances sentiment-semantic association, and ensures the reliability and accuracy of fusion results.
Smart Images

Figure CN121580302A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of artificial intelligence and multi-modal data processing, and relates to a multi-modal big data analysis processing method and device based on artificial intelligence. BACKGROUND
[0002] In the current rapid development of information technology, multi-modal data processing technology has been widely applied in the field of speech input recognition and conversion. It aims to improve the accuracy and comprehensiveness of recognition by fusing various modal data such as speech and facial images, and provides strong support for intelligent interaction, voice assistants, real-time translation and many other scenarios.
[0003] However, the existing recognition and conversion technology for speech input still has many limitations, which restricts its efficient application in a wider range of scenarios: Firstly, the scene adaptability is insufficient. Most existing technologies only fuse speech and facial images for general scenarios, without fully considering the differences in modal data requirements for different application scenarios. For example, in a noisy industrial production scene, the anti-interference processing requirement for speech signals is higher, while in a quiet office scene, more attention may be paid to the cooperative analysis of facial micro-expressions and speech. This universal processing method leads to a significant decrease in recognition accuracy in specific scenarios.
[0004] Secondly, the emotional-semantic correlation is weak. Although existing technologies can recognize user emotions through facial features, they fail to effectively combine them with speech semantic content to verify the consistency of emotions and semantics. When there are contradictions such as the user's speech expressing "understanding" but the facial features showing "confusion", the technology cannot correct it, resulting in ambiguous converted information and affecting the subsequent judgment of the user's true intention.
[0005] Thirdly, the modal fusion method is fixed. Existing technologies mostly use static weight fusion strategies when fusing multi-modal data, i.e., pre-setting the fusion weights of speech, facial images and other modal data, without dynamically adjusting according to real-time data quality. In actual applications, the quality of modal data is often unstable, such as speech with noise, blurred images, etc. Fixed weight distribution cannot adapt to real-time changes in data quality, affecting the reliability of the fusion results.
[0006] Fourthly, cross-modal error correction is missing. When a single modal recognition error occurs, such as a speech recognition error, existing technologies cannot use other modal data (such as facial lip feature) for auxiliary correction. Due to the lack of mutual checking and correction mechanism between cross-modal, the error rate of converted information is high, which is difficult to meet the application requirements of high-precision recognition.
[0007] In summary, the existing multi-modal data processing technology has the problems of insufficient scene adaptability, weak emotion-semantic correlation, fixed modal fusion method and lack of cross-modal error correction in speech input recognition and conversion, and an improved technical solution is needed to solve the problems to improve the accuracy, reliability and scene adaptability of speech input recognition and conversion. SUMMARY
[0008] To solve the problems in the background art, the present application provides a multi-modal big data analysis and processing method and device based on artificial intelligence, aiming to improve the accuracy of multi-modal data processing, scene adaptability and reliability of information conversion by introducing scene perception, emotion-semantic joint analysis, dynamic weight fusion and cross-modal error correction mechanism.
[0009] The first aspect of the present application provides a multi-modal big data analysis and processing method based on artificial intelligence, comprising: Based on the modal configuration parameters, the multi-modal data is collected and preprocessed; Obtain deep semantic information, facial emotion sequence and speech emotion sequence, and perform consistency verification on the correlation of the deep semantic information, the facial emotion sequence and the speech emotion sequence under the same timestamp; Determine the dynamic weight based on the voice modal confidence, image modal confidence and auxiliary data confidence, and fuse the deep semantic information, facial emotion sequence and speech emotion sequence using the dynamic weight; According to the cross-modal verification, the semantic ambiguity or error part in the fusion result is corrected, and based on the corrected fusion result, the conversion recognition information is generated.
[0010] Optionally, the process of generating the deep semantic information comprises: performing preliminary recognition on the noise-reduced speech data based on a pre-trained speech recognition model to generate an initial text, combining a domain knowledge base corresponding to the scene type to correct and complete the initial text to obtain the deep semantic information.
[0011] Optionally, the process of generating the facial emotion sequence comprises: extracting the facial feature point time sequence motion trajectory in the image data based on the optical flow method, and matching the time sequence motion trajectory with a preset emotion template through a dynamic time warping algorithm to obtain the facial emotion sequence.
[0012] Optionally, the process of generating the speech emotion sequence comprises: extracting the fundamental frequency and syllable interval time of the speech data, and inputting a sentiment classification model based on a gated recurrent unit network to obtain the speech emotion sequence.
[0013] Optionally, the consistency verification includes: if there is a contradiction between facial emotion, voice emotion and deep semantic information at the same timestamp, then the scene rule base is invoked for correction.
[0014] Optionally, the formula for calculating the confidence level of the speech modality is: The formula for calculating the image modality confidence is as follows: The confidence level of the auxiliary data The similarity between the text draft and the speech semantics is calculated based on the cosine distance between the bidirectional encoder representation converter vectors; in, For speech modal confidence, The scene coefficient is the ratio of the number of matched domain terms to the total number of words, and the speech rate stability is the ratio of 1 minus the standard deviation of speech rate to the average speech rate; Image modality confidence is defined as follows: clear frame percentage is the ratio of the number of unblurred frames to the total number of frames; and emotional fluctuation anomaly is the ratio of the absolute value of the difference between the current emotion and the historical emotion to a preset threshold.
[0015] Optionally, the dynamic weights include speech weights. Image weights and auxiliary data weights , ; ; ;in, For speech modal confidence, For image modal confidence, To assist in data confidence; The cross-modal verification includes: if the words recognized by speech recognition do not match the facial lip features, then the domain knowledge base is invoked to perform word replacement; if there is semantic ambiguity, then a unique interpretation is determined by combining the scene type.
[0016] A second aspect of this application provides a multimodal big data analysis and processing device based on artificial intelligence, comprising: The processing module is used to collect multimodal data and preprocess the multimodal data based on modal configuration parameters; The analysis module is used to acquire deep semantic information, facial emotion sequences, and voice emotion sequences, and to perform consistency verification on the correlation between the deep semantic information, the facial emotion sequences, and the voice emotion sequences at the same timestamp. The fusion module is used to determine dynamic weights based on the speech modality confidence, image modality confidence, and auxiliary data confidence, and to fuse deep semantic information, facial emotion sequences, and speech emotion sequences using the dynamic weights. The generating module is configured to correct the ambiguous or incorrect part in the fusion result according to the cross-modal verification, and generate the conversion recognition information based on the corrected fusion result.
[0017] Optionally, the fusion module comprises: The calculation formula of the image modality confidence is: The auxiliary data confidence is a similarity between the text draft and the voice semantics, and the similarity is calculated based on a bidirectional encoder representation transformer vector cosine distance; wherein, is the voice modality confidence, is a scene coefficient, a word matching rate is a ratio of the number of matched field terms to the total number of words, and a speech speed stability is 1 minus a ratio of a speech speed standard deviation to an average speech speed; and the is the image modality confidence, a clear frame proportion is a ratio of the number of non-blurred frames to the total number of frames, and an emotion fluctuation abnormal value is an absolute value of a difference between a current emotion and a historical emotion divided by a preset threshold.
[0018] Optionally, the fusion module comprises: the dynamic weight comprises a voice weight , an image weight and an auxiliary data weight , ; ; wherein, is the voice modality confidence, is the image modality confidence, is the auxiliary data confidence; The cross-modal verification comprises: if the words of the voice recognition do not match the facial lip feature, the field knowledge base is called to replace the words; and if the semantics are ambiguous, the unique explanation is determined in combination with the scene type.
[0019] Compared with the prior art, the present application has the following beneficial effects: The present application provides a multi-modal big data analysis processing method and device based on artificial intelligence, a dynamic weight fusion algorithm based on modality confidence, which replaces fixed weight and improves the reliability of the fusion result; the modality confidence and the weight are quantified through a mathematical model, so that the multi-modal fusion is upgraded from experience-driven to data-driven, and the processing precision is greatly improved. BRIEF DESCRIPTION OF DRAWINGS
[0020] Figure 1 is a flowchart of the multi-modal big data analysis processing method based on artificial intelligence in an embodiment of the present application; Figure 2is a schematic view of an artificial intelligence-based multi-modal big data analysis processing device in an embodiment of the present application. DETAILED DESCRIPTION
[0021] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of the present application.
[0022] In an embodiment, as shown in Figure 1 , an artificial intelligence-based multi-modal big data analysis processing method is provided. Taking the method applied in Figure 1 as an example, the following specific steps are described. S10: Based on the modal configuration parameters, multi-modal data is collected and preprocessed.
[0023] Specifically, the modal configuration parameters in the present application are configured by collecting the scene feature data of the current environment, inputting the scene classification model to determine the scene type, and calling the preset modal configuration parameters according to the scene type. The process is as follows: The scene feature data includes environmental audio, background image and user identity label, wherein the environmental audio collects the sound signal in the current scene through the microphone of the terminal device, the background image collects the picture information in the current scene through the camera of the terminal device, and the user identity label is the identity information provided by the user when logging into the system. The collected environmental audio, background image and user identity label are input into the scene classification model, which is a hybrid architecture of convolutional neural network and long short-term memory network. The scene type is output by jointly analyzing the spectral features of the environmental audio, the visual features of the background image and the attribute features of the user identity label. The scene type includes but is not limited to teaching scene, medical consultation scene and daily conversation scene. After determining the scene type, the system calls the preset modal configuration parameters corresponding to the scene type. The modal configuration parameters include speech recognition weight and face feature capture frame rate. For example, in the modal configuration parameters in the teaching scene, the speech recognition weight is set to a high value for subject-specific terms, and the face feature capture frame rate is set to a high frame rate for students' micro-expression. In the modal configuration parameters in the medical consultation scene, the speech recognition weight is set to a high value for symptom description terms, and the face feature capture frame rate is set to a high frame rate for patient's pain expression features. In the modal configuration parameters in the daily conversation scene, the speech recognition weight and the face feature capture frame rate are balanced.
[0024] Based on the modal configuration parameters, collect voice data, image data and auxiliary data, perform noise reduction processing on the voice data, and feature point normalization processing on the image data to obtain a multi-modal data process as follows: According to the called modal configuration parameters, the voice data of the user is collected by the directional microphone, the image data of the user is collected by the high-definition camera, and the auxiliary data is obtained, the auxiliary data includes the text draft input by the user and the historical interaction record of the user stored by the system; The text draft input by the user includes: "pre-input draft" before voice input, that is, the user inputs the text content related to this voice theme in advance before starting voice description, for example, in the teaching scene, the user inputs the draft before voice explanation "mathematics problem solving method", and clearly defines the theme category of voice input. Secondly, the synchronous supplementary draft in the voice input, that is, the user synchronously inputs the text segment when finding that the voice expression is not accurate enough or there is information omission in the voice input process, for example, in the medical consultation scene, the user synchronously inputs the draft to supplement the symptoms when describing the physical discomfort in the voice, "night fever, no runny nose", and supplement the details not clearly stated in the voice.
[0025] The adaptive filtering algorithm is used to process the collected voice data, the noise in the environment is filtered, and the clear user voice signal is retained; The feature point normalization processing is performed on the collected image data, the face feature points of the user in the image are extracted, and the coordinates of these face feature points are uniformly mapped to the preset face coordinate system, so that the positions of the same face feature points in the images collected at different times and different angles have consistency; After the above processing, the multi-modal data containing the noise-reduced voice data, the feature point normalized image data and the auxiliary data are obtained, and the multi-modal data can be directly used for subsequent analysis and processing steps.
[0026] In this embodiment, the scene classification model is a hybrid architecture of convolutional neural network and long short-term memory network, which jointly analyzes the spectral features of environmental audio, the visual features of background image and the attribute features of user identity label, and specifically includes the following contents: Specifically, the hybrid architecture of convolutional neural network and long short-term memory network is composed of feature extraction layer, feature fusion layer and classification layer. Among them, the feature extraction layer contains two parallel convolutional neural network branches, which are used to process the spectral features of environmental audio and the visual features of background image respectively; The feature fusion layer accesses the long short-term memory network for processing the fused time sequence features; The classification layer is a full connection layer for outputting the scene type.
[0027] The spectral feature extraction process of the environmental audio is: the collected environmental audio signal is preprocessed, for example, framing, windowing, the time-domain audio signal is converted into a two-dimensional mel spectrum graph through mel spectrum conversion, for example, the horizontal axis is time and the vertical axis is mel frequency, to obtain the spectral features of the environmental audio; the spectral features of the environmental audio are input into the first convolutional neural network branch, and the high-frequency energy distribution, spectral envelope and other key spatial features in the spectral features of the environmental audio are extracted through the multi-layer stacking of the convolutional layer and the pooling layer, and an audio feature vector is output.
[0028] The visual feature extraction process of the background image is: the collected background image is standardized, for example, size normalization, pixel value normalization, to obtain the visual features of the background image; the visual features are input into the second convolutional neural network branch, and the texture, color distribution, object contour and other spatial features in the image are captured through convolution operation, and after dimension reduction by the pooling layer, an image feature vector is output.
[0029] The attribute feature processing process of the user identity label is: the user identity label is a text form of identity label, such as "teacher", "patient" and "ordinary user", which is converted into a fixed-dimensional attribute feature vector through an embedding layer, and the attribute feature vector contains the category information of the user identity, such as occupation and role.
[0030] The joint analysis process is: the audio feature vector, the image feature vector and the attribute feature vector of the user identity label are spliced to obtain a fusion feature vector; the fusion feature vector is input into a long short-term memory network, which has the ability to capture time sequence dependency, to analyze the correlation rules of audio, image and identity features at different times, such as in a teaching scene, the environmental audio often contains board writing sound and question and answer sound, the background image often contains blackboard and desks and chairs, and the user identity label is often "teacher" or "student", and the three are correlated in time sequence. The time sequence feature vector output by the long short-term memory network is processed by a classification layer to obtain the classification result of the scene type, for example, a teaching scene, a medical consultation scene or a daily conversation scene.
[0031] Through the joint analysis of the three features by the above-mentioned mixed architecture of convolutional neural network and long short-term memory network, the scene classification model can comprehensively utilize the multi-dimensional information of environmental audio, background image and user identity to improve the accuracy of scene type recognition.
[0032] S20: Obtain the deep semantic information, the facial emotion sequence, and the speech emotion sequence, and perform consistency verification on the association of the deep semantic information, the facial emotion sequence, and the speech emotion sequence at the same timestamp.
[0033] Specifically, the process of recognizing the speech data in the multi-modal data and generating deep semantic information in combination with the domain knowledge base is as follows: based on a pre-trained speech recognition model, the noise-processed speech data is preliminarily recognized, the speech signal is converted into a corresponding text form, and an initial text is generated; according to the determined scene type, the corresponding domain knowledge base is called, for example, the teaching scene corresponds to the subject term library, and for example, the medical consultation scene corresponds to the symptom library; the initial text is corrected and completed through the domain knowledge base, for example, in the teaching scene, if "single original" appears in the initial text, it is corrected to "single substance" in combination with the subject term library, and if the subject is omitted, such as "explain the problem solving steps", it is completed to "teacher explains the problem solving steps" in combination with the context and user identity label, and finally the complete and accurate deep semantic information is obtained.
[0034] The process of analyzing the image data in the multi-modal data to generate a facial emotion sequence is as follows: based on the optical flow method, the image data processed by the feature point normalization is processed, the motion trajectories of the facial feature points such as eyebrows, corners of the mouth, eyeballs, etc. in the time dimension are tracked and extracted, and the time sequence motion trajectories of the facial feature points are obtained; the time sequence motion trajectories are input into the dynamic time warping algorithm, and are matched with the preset emotion template to determine the emotion type at each timestamp, the preset emotion template, for example, the trajectory feature of "confusion" corresponds to the raised eyebrows and drooping corners of the mouth, and the trajectory feature of "understanding" corresponds to the relaxed eyebrows and slightly raised corners of the mouth. The emotion types of each timestamp are integrated in time sequence to generate a facial emotion sequence in the form of "timestamp 1: emotion type 1; timestamp 2: emotion type 2……".
[0035] The process of analyzing the tone and speed features of the speech data to generate a speech emotion sequence is as follows: the tone features and speed features are extracted from the noise-processed speech data, wherein the tone features are obtained by calculating the fundamental frequency of the speech signal, and the speed features are obtained by calculating the syllable interval time; the extracted fundamental frequency and syllable interval time are input into an emotion classification model based on a gated recurrent unit network, the model learns the association rules of tone, speed and emotion, and outputs the emotion type at each timestamp; the emotion types of each timestamp are integrated in time sequence to generate a speech emotion sequence, which is consistent with the facial emotion sequence, i.e. "timestamp 1: emotion type 1; timestamp 2: emotion type 2……".
[0036] The consistency verification process of the association between the facial emotion, the voice emotion and the deep semantic information under the same timestamp is specifically as follows: for each timestamp, the facial emotion corresponding to the timestamp from the facial emotion sequence, the voice emotion from the voice emotion sequence and the semantic content in the deep semantic information are associated and analyzed to determine whether the three are matched, for example, if the deep semantic information is "this question is very simple", the corresponding facial emotion and voice emotion should tend to be "calm" or "confident"; if there is an association contradiction among the three, such as the deep semantic information is "I am very happy", but the facial emotion is "angry" and the voice emotion is "irritated", the scene rule library is called for correction, wherein the facial emotion is preferred in the teaching scene, such as the student may hide the true emotion due to shyness, and finally the verified and corrected emotion and semantic association result is obtained.
[0037] S30: determining a dynamic weight based on the voice modality confidence, the image modality confidence and the auxiliary data confidence, and fusing the deep semantic information, the facial emotion sequence and the voice emotion sequence by using the dynamic weight.
[0038] Specifically, when calculating the voice modality confidence, the formula is used, wherein, is the voice modality confidence; is a scene coefficient; the word matching rate is the ratio of the number of matched field terms in the deep semantic information to the total number of words, for example, the proportion of the number of recognized terms such as "single substance" and "function" in the teaching scene to the total number of words matched with the discipline term library; the speech speed stability is 1 minus the ratio of the speech speed standard deviation to the average speech speed, wherein the speech speed standard deviation is the dispersion degree of the syllable interval time in the voice data, and the average speech speed is the average value of the syllable interval time in the voice data.
[0039] The syllable interval time is the core basic data for calculating the average speech speed and the speech speed standard deviation, and the extraction process needs to be performed on the voice data after noise reduction by the adaptive filtering algorithm: first, the voice endpoint detection is performed on the noise-reduced voice data, the valid speech segment and the silent segment are distinguished by the joint determination of the short-time energy and the zero-crossing rate, and the interference of the silent segment on the syllable interval calculation is excluded; then the syllable boundary in the valid speech segment is recognized by using the method based on the pre-trained syllable segmentation model, the model can accurately locate the start time and end time of each syllable by learning the acoustic features of the syllables in a large amount of labeled voice data; finally, the start time difference of the adjacent two syllables is calculated to obtain the single syllable interval time, denoted as wherein i is the serial number of the syllable interval, taking values of 1, 2, …, n, and n is the total number of syllable interval times in the current valid speech segment. For example, if the start time of the first syllable is 0.2 seconds and the start time of the second syllable is 0.5 seconds, the corresponding syllable interval time is = 0.8 - 0.5 = 0.3 seconds, and so on to obtain all = 0.8 - 0.5 = 0.3 seconds, and so on to obtain all .
[0040] The average speech speed is the average value of the syllable interval time in the speech data, which is used to quantify the overall speed of the current speech. The calculation process includes screening valid data, summation, and average calculation. The specific process is as follows: first, screen the valid syllable interval time, exclude abnormal values caused by user pauses or speech recognition errors, and retain the valid syllable interval time within the normal speed range ; then calculate the sum of all valid syllable interval times after screening, denoted as Σ = + +…+ , where n is the number of valid syllable interval times; finally, calculate the average speech speed by the arithmetic average formula, which is: where represents the average speech speed, with the unit of "seconds per syllable", i.e., the average interval time of each syllable, the smaller the better. For example: if 5 syllable interval times are screened out in the current valid speech segment, they are = 0.2 seconds, = 0.3 seconds, = 0.25 seconds, = 0.3 seconds, = 0.25 seconds, then Σ = 0.2 + 0.3 + 0.25 + 0.3 + 0.25 = 1.3 seconds, n = 5, and substituting the formula gives seconds per syllable, i.e., the average interval of each syllable is 0.26 seconds, reflecting that the overall speech speed of the current speech is relatively stable.
[0041] The speech speed standard deviation is a quantitative indicator of the dispersion degree of the syllable interval time in the speech data. The higher the dispersion degree, the greater the fluctuation of the speech speed, i.e., the speech speed is fast and slow; the lower the dispersion degree, the more stable the speech speed. Its calculation needs to take the average speech speed as the benchmark, following the steps of difference calculation, square calculation, average calculation, and square root calculation. The specific formula and process are as follows: calculate the difference between each valid syllable interval time and the average speech speed , denoted as ; square each difference to obtain to eliminate the influence of positive and negative differences canceling each other out; calculate the average value of all , i.e., the variance S 2 , which is: ; and Square root operation is performed to obtain the speech speed standard deviation , the formula is: , the unit is consistent with the syllable interval time.
[0042] When calculating the image modality confidence, the formula is used, wherein, is the image modality confidence; the clear frame proportion is the ratio of the number of frames without blur to the total number of frames in the image data, and the frame without blur refers to the frame in which the facial feature points are clear and identifiable; the emotion fluctuation abnormal value is the ratio of the absolute value of the difference between the current emotion and the historical emotion to the preset threshold value, the current emotion is the emotion type corresponding to a timestamp in the image data, the historical emotion is the average of the emotion types of the previous N consecutive frames, and the preset threshold value is the upper limit of the normal fluctuation value obtained based on the historical emotion data statistics.
[0043] First, the current emotion is the emotion type corresponding to a timestamp in the image data, and the emotion type comes from the complete analysis process of the image data in the standardized multi-modal data. The time sequence motion trajectory of the facial feature points such as eyebrows, corners of the mouth, eyeballs, etc. in the image data is extracted by the optical flow method, and then the time sequence motion trajectory is input into the dynamic time warping algorithm and matched with the preset emotion template. For example, the trajectory feature of "confusion" corresponds to the upward eyebrow + drooping mouth, and the trajectory feature of "understanding" corresponds to the relaxed eyebrow + slightly raised mouth, and finally the unique emotion type at the timestamp is determined, such as "confusion", "calm", "irritated" and "understanding". In order to realize numerical calculation, a fixed emotion quantization value needs to be preset for each emotion type, and the quantization value needs to reflect the gradient difference of the emotion, such as low negative emotion quantization value and high positive emotion quantization value. For example, the preset "irritated" corresponds to the quantization value 0, "confusion" corresponds to the quantization value 1, "calm" corresponds to the quantization value 3, and "understanding" corresponds to the quantization value 5. The quantization rule is pre-stored in the emotion coding library of the system to ensure that the quantization value of the same emotion type is completely unified in different calculation scenarios, and to provide a basis for subsequent difference calculation.
[0044] Secondly, the calculation of historical emotion, which is the average of the emotion types of the previous N frames before the current emotion corresponding timestamp, where the value of "continuous N frames" needs to be combined with the scene type preset, and the value logic is adapted to the characteristics of emotion change in the scene: in the teaching scene, the user's emotion change is usually relatively gentle, and N is valued at 10 frames to fully reflect the stable trend of emotion; in the medical consultation scene, the user may have rapid fluctuations in emotion due to disease description, treatment scheme communication, etc., and N is valued at 5 frames to accurately capture recent changes in emotion; in the daily conversation scene, the user's emotional fluctuations are between the two, and N is valued at 8 frames. In specific calculation, first, the emotion types of the previous N frames from the current timestamp are extracted from the facial emotion sequence, then each frame of emotion type is converted to the corresponding emotion quantization value through the emotion coding library, and finally the arithmetic average formula is used to calculate the average of the historical emotion, the formula is: historical emotion average = (1st frame emotion quantization value + 2nd frame emotion quantization value + … + Nth frame emotion quantization value) / N. For example, the current emotion corresponding timestamp is t10 (N=10 in the teaching scene), then the emotion types of t0 to t9 frames are extracted, and if their corresponding quantization values are 1, 1, 2, 2, 3, 2, 2, 1, 1, 2, then the historical emotion average = (1+1+2+2+3+2+2+1+1+2) / 10 = 1.7.
[0045] Finally, the determination of the preset threshold and the calculation of the emotional fluctuation abnormal value, the preset threshold is the upper limit value of the normal emotional fluctuation in each scene based on a large amount of historical emotional data statistics, and needs to be set by scene to adapt to the emotional fluctuation law of different scenes: in the teaching scene, 100,000 pieces of emotional fluctuation data of teachers and students are counted, and the normal emotional fluctuation amplitude is obtained through data analysis, that is, the absolute value of the difference between the single frame emotion and the historical mean value, so the preset threshold of the teaching scene is set to 2.0; in the medical consultation scene, 80,000 pieces of emotional fluctuation data of patients are counted, and the 95% normal emotional fluctuation amplitude is not more than 3.5, so the preset threshold is set to 3.5; in the daily conversation scene, 150,000 pieces of emotional fluctuation data of ordinary users are counted, and the 95% normal emotional fluctuation amplitude is not more than 2.5, so the preset threshold is set to 2.5, which is pre-stored in the scene rule library and is called synchronously with the determination of the scene type. After obtaining the current emotional quantization value, the historical emotional mean value and the preset threshold, the abnormal value is calculated through the formula: emotional fluctuation abnormal value = absolute value of (current emotional quantization value - historical emotional mean value) / preset threshold, for example, in the teaching scene, the current emotional quantization value is 0, which corresponds to "frustration", the historical emotional mean value is 1.7, and the preset threshold is 2.0, so the emotional fluctuation abnormal value = |0-1.7| / 2.0 = 0.85; if the current emotional quantization value is 5, it corresponds to "understanding", and the historical emotional mean value is 1.7, then the emotional fluctuation abnormal value = |5-1.7| / 2.0 = 1.65. The abnormal value can be used to judge whether the emotional fluctuation is beyond the normal range, wherein the abnormal value ≤1 is determined as normal emotional fluctuation, and the abnormal value >1 is determined as emotional fluctuation abnormality.
[0046] When calculating the auxiliary data confidence, if the auxiliary data contains a text draft input by the user, is the similarity between the text draft and the voice semantics, which is calculated based on the cosine distance of the bidirectional encoder representation transformer vector, that is, the text draft and the deep semantic information are respectively converted into bidirectional encoder representation transformer vectors, and then the cosine distance of the two vectors is calculated; if the auxiliary data does not contain a text draft, =0.
[0047] The process of determining the dynamic weight based on the voice modality confidence, the image modality confidence and the auxiliary data confidence is as follows: the dynamic weight includes the voice weight , the image weight and the auxiliary data weight , and the calculation formulas of the three are respectively ; ; ; the , , calculated by the above formulas satisfy + + = 1, and dynamically adjusts with real-time changes of , , For example, when the voice data is clear and the image data is fuzzy, that is, the value is high and the value is low, the increases and the decreases to highlight the contribution of the voice modal data.
[0048] The process of fusing the deep semantic information, the facial emotion sequence and the voice emotion sequence according to the dynamic weight is as follows: the deep semantic information, the facial emotion sequence and the voice emotion sequence are respectively converted into vector forms, wherein the deep semantic information is converted into a semantic vector through a word embedding model, and the facial emotion sequence and the voice emotion sequence are converted into emotion vectors through an emotion coding model; the three vectors are weighted and summed according to the dynamic weight, that is, the fusion result vector = W1 x semantic vector + W2 x facial emotion vector + W3 x voice emotion vector; the fusion result vector integrates the effective information of each modal data, and the weight dynamically adjusts with the quality of each modal data, so that the fusion result is more in line with the reliability characteristics of the current data.
[0049] S40: correcting the semantic ambiguity or error part in the fusion result according to the cross-modal verification, and generating the conversion recognition information based on the corrected fusion result.
[0050] Specifically, the process of correcting the semantic ambiguity or error part in the fusion result through cross-modal verification is as follows: for the semantic ambiguity in the fusion result, such as a word corresponding to multiple interpretations or an error, such as a part where the word recognized by voice recognition does not match the actual intention, a multi-dimensional verification is performed by calling a cross-modal verification mechanism. If the word recognized by voice recognition does not match the facial lip feature in the image data, for example, the voice recognition is "four" but the facial lip feature shows the lip shape of "ten", the word matching the facial lip feature is replaced from the domain knowledge base corresponding to the scene type to ensure that the word is consistent with the lip action; if the semantics is ambiguous, such as "dried yam" in the text may refer to agricultural products or medicinal materials, a unique interpretation is determined in combination with the determined scene type, for example, "medicinal material" is determined in the medical scene, and "agricultural product" is determined in the commercial scene; after the above verification and correction, an accurate and unambiguous fusion result is obtained.
[0051] Based on the corrected fusion result, the conversion recognition information including the basic text, the emotional label and the scene note is generated. The process is as follows: the basic text is the semantic content corrected by cross-modal verification, that is, the complete text formed by integrating deep semantic information and cross-modal correction result, covering the core semantics expressed by the user, such as "the key to solving this problem is the oxidizing property of single substance"; the emotional label is composed of a timestamp and a corresponding emotional type, the timestamp is the specific time when the emotion occurs, which is "t1" and "t3", the emotional type is the emotion determined after consistency verification and fusion, which is "confusion" and "understanding", the label form is "[timestamp: emotional type]", for example, "[t1: confusion][t3: understanding]"; the scene note is the explanation of the domain term related to the scene type, that is, for the professional terms appearing in the basic text, the corresponding definition or explanation is retrieved from the domain knowledge base, for example, the note of "single substance" in the teaching scene is "definition: pure substance composed of the same element", and the note of "fever" in the medical consultation scene is "definition: physiological state of human body temperature exceeding the normal range of 36.3-37.2℃"; the basic text, emotional label and scene note are integrated to form the final conversion recognition information, which not only contains accurate semantic content, but also reflects the emotional changes of the user, and at the same time, professional explanation is added to adapt to the scene demand.
[0052] In an embodiment, as shown in Figure 2 , a multi-modal big data analysis processing device based on artificial intelligence is provided, which corresponds to the multi-modal big data analysis processing method based on artificial intelligence in the above embodiment. The multi-modal big data analysis processing device based on artificial intelligence comprises a processing module, an analysis module, a fusion module and a generation module, and the functions of each module are described as follows: The processing module is configured to collect multi-modal data based on the modal configuration parameters and pre-process the multi-modal data. The analysis module is configured to obtain deep semantic information, facial emotion sequence and speech emotion sequence, and perform consistency verification on the association of the deep semantic information, the facial emotion sequence and the speech emotion sequence under the same timestamp. The fusion module is configured to determine a dynamic weight based on the speech modal confidence, the image modal confidence and the auxiliary data confidence, and fuse the deep semantic information, the facial emotion sequence and the speech emotion sequence by using the dynamic weight. The generation module is configured to correct the semantic ambiguity or error part in the fusion result according to cross-modal verification, and generate conversion recognition information based on the corrected fusion result.
[0053] Optionally, the fusion module comprises: the calculation formula of the speech modal confidence is: The calculation formula of the image modal confidence is: ; the auxiliary data confidence is a similarity of the text draft and the voice semantics, and the similarity is calculated based on a bidirectional encoder representation transformer vector cosine distance; wherein, is a voice modality confidence, is a scene coefficient, a word matching rate is a ratio of a number of matched field terms to a total number of words, and a speech speed stability is 1 minus a ratio of a speech speed standard deviation to an average speech speed; the is an image modality confidence, a clear frame proportion is a ratio of a number of unblurred frames to a total number of frames, and an emotion fluctuation abnormal value is a ratio of an absolute value of a difference between a current emotion and a historical emotion to a preset threshold.
[0054] Optionally, the fusion module comprises: a dynamic weight comprising a voice weight , an image weight , and an auxiliary data weight , ; ; ; wherein, is a voice modality confidence, is an image modality confidence, is an auxiliary data confidence; The cross-modality verification comprises: if words recognized by voice do not match facial lip feature, calling a field knowledge base to replace the words; and if semantics are ambiguous, determining a unique explanation in combination with a scene type.
[0055] The specific limitations of the multi-modal big data analysis processing device based on artificial intelligence can be referred to the limitations of the multi-modal big data analysis processing method based on artificial intelligence in the foregoing, which will not be repeated here. Each module in the multi-modal big data analysis processing device based on artificial intelligence can be realized by software, hardware, and a combination thereof, in whole or in part. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a memory in a computer device in software form, so as to be called and executed by a processor to perform the operations corresponding to each module.
[0056] Although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some of the technical features, and any modification, equivalent replacement, improvement, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.
Claims
1. An artificial intelligence-based multi-modal big data analysis processing method, characterized by, The method comprises the steps of: Based on the modal configuration parameters, collect multi-modal data and preprocess the multi-modal data; Obtain deep semantic information, facial emotion sequence, and speech emotion sequence, and perform consistency verification on the correlation of the deep semantic information, the facial emotion sequence, and the speech emotion sequence at the same timestamp; Determine a dynamic weight based on the speech modal confidence, the image modal confidence, and the auxiliary data confidence, and fuse the deep semantic information, the facial emotion sequence, and the speech emotion sequence by using the dynamic weight; According to the cross-modal verification, the semantic ambiguous or incorrect part in the fusion result is corrected, and based on the corrected fusion result, conversion recognition information is generated.
2. The method of claim 1, wherein the method is based on artificial intelligence. The process of generating the deep semantic information comprises the steps of: performing preliminary identification on the noise-reduced speech data based on a pre-trained speech recognition model to generate an initial text, and combining a domain knowledge base corresponding to a scene type to correct and complete the initial text to obtain deep semantic information. 3.The method of claim 1, wherein, The process of generating the facial emotion sequence comprises the steps of: extracting a facial feature point time sequence motion trajectory in the image data based on an optical flow method, and matching the time sequence motion trajectory with a preset emotion template by a dynamic time warping algorithm to obtain a facial emotion sequence. 4.The method of claim 1, wherein, The process of generating the speech emotion sequence comprises the steps of: extracting a fundamental frequency and an inter-syllable interval time of the speech data, and inputting a sentiment classification model based on a gated recurrent unit network to obtain a speech emotion sequence. 5.The method of claim 1, wherein, The consistency verification comprises: if the facial emotion, the speech emotion, and the deep semantic information at the same timestamp exist correlation contradictions, a scene rule base is called for correction. 6.The method of claim 1, wherein, The calculation formula of the voice modality confidence is: The calculation formula of the image modality confidence is: The auxiliary data confidence is the similarity of the text draft and the voice semantics, and the similarity is calculated based on a bidirectional encoder representation transformer vector cosine distance. wherein, is a voice modality confidence, is a scene coefficient, a word matching rate is a ratio of a number of matched field terms to a total number of words, and a speech speed stability is 1 minus a ratio of a standard deviation of speech speed to an average speech speed; the is an image modality confidence, a clear frame proportion is a ratio of a number of frames without blur to a total number of frames, and an emotion fluctuation abnormal value is a ratio of an absolute value of a difference between a current emotion and a historical emotion to a preset threshold. 7.The method of claim 1, wherein, The dynamic weights include a speech weight , an image weight , and an auxiliary data weight , ; ; ; wherein is a speech modality confidence, is an image modality confidence, is an auxiliary data confidence; The cross-modal verification comprises: if the words of speech recognition and the facial lip shape features do not match, the domain knowledge base is called for word replacement; if the semantics are ambiguous, a unique explanation is determined in combination with the scene type.
8. An artificial intelligence-based multi-modal big data analysis processing apparatus, characterized by, The method comprises the steps of: The processing module is configured to collect multi-modal data and preprocess the multi-modal data based on modal configuration parameters; The analysis module is configured to obtain deep semantic information, facial emotion sequence, and speech emotion sequence, and perform consistency verification on the correlation of the deep semantic information, the facial emotion sequence, and the speech emotion sequence at the same timestamp; The fusion module is configured to determine a dynamic weight based on the speech modal confidence, the image modal confidence, and the auxiliary data confidence, and fuse the deep semantic information, the facial emotion sequence, and the speech emotion sequence by using the dynamic weight; The generation module is configured to correct the semantic ambiguous or incorrect part in the fusion result according to cross-modal verification, and generate conversion recognition information based on the corrected fusion result. 9.The artificial intelligence-based multi-modal big data analysis processing apparatus of claim 8, wherein The fusion module includes: a calculation formula of the voice modality confidence is ; a calculation formula of the image modality confidence is ; and the auxiliary data confidence is a similarity of a text draft and voice semantics, and the similarity is calculated based on a bidirectional encoder representation transformer vector cosine distance. wherein, is a voice modality confidence, is a scene coefficient, a word matching rate is a ratio of a number of matched field terms to a total number of words, and a speech speed stability is 1 minus a ratio of a standard deviation of speech speed to an average speech speed; the is an image modality confidence, a clear frame proportion is a ratio of a number of frames without blur to a total number of frames, and an emotion fluctuation abnormal value is a ratio of an absolute value of a difference between a current emotion and a historical emotion to a preset threshold. 10.The artificial intelligence-based multi-modal big data analysis processing apparatus of claim 8, wherein The fusion module includes: dynamic weights including speech weights , image weights , and auxiliary data weights , ; ; ; wherein is a speech modality confidence, is an image modality confidence, is an auxiliary data confidence; The cross-modal verification comprises: if the words of speech recognition and the facial lip shape features do not match, the domain knowledge base is called for word replacement; if the semantics are ambiguous, a unique explanation is determined in combination with the scene type.
Citation Information
Patent Citations
Multi-modal emotion recognition system and method
CN118260638A
Emotion recognition method based on large model and related device
CN119904901A
Voice data extraction method and system adopting artificial intelligence
CN120319224A
Voice emotion recognition method and device based on context information, equipment and medium
CN120636474A
Multi-dimensional identity information identification method and apparatus, computer device and storage medium
WO2021000829A1